The gap between signing and verifying
Signature checks fail far more often from re-serialisation than from cryptography: the sender signed one byte string and the receiver re-encoded it into another before checking.
Almost every signature verification failure I have debugged had nothing to do with cryptography. The keys were right, the algorithm was right, the library was right. The sender had signed one byte string, and by the time the receiver checked the signature it was checking a different byte string, because something between the network and the verification code had parsed the message and written it back out. A JSON body became an object and became JSON again with the keys in a different order. A framework stripped a trailing newline. A proxy normalised a unicode escape.
The signature was over the bytes the sender sent; the check was over the bytes the receiver's stack reconstructed; and the two differed by a character nobody could see. I call the set of those transformations the re-serialisation gap, and a signed-message design is rated by whether the gap is empty.
Where the bytes diverge
The pipeline on each side has more stages than the diagram in the documentation shows. The sender builds an object, serialises it to bytes, computes a signature over those bytes, and sends both. The receiver's web server reads the bytes, a middleware may decompress or decode them, the framework parses them into an object for the handler's convenience, and then the verification code, if it is written naively, serialises that object back to bytes and computes the signature over the result. Every stage after the network is a chance for the bytes to change without the content changing, and the signature does not care about content. It cares about bytes.
The gap has a known membership. Key reordering, because most serialisers write object keys in whatever order the language's map iterates them. Whitespace, because pretty-printed and compact JSON are the same object and different bytes.
Unicode, because a serialiser may escape a character the sender wrote literally, or normalise a composed character into a decomposed one. Trailing newlines, which some HTTP clients add and some frameworks strip. Number formatting, where 1.0 and 1 and 1e0 are one value and three byte strings. And, the one that catches people who think they are safe, the framework's decoding of the body into a string at all, which can replace bytes it considers invalid with a replacement character before the handler ever sees them.
What the providers do
The public webhook schemes show the two design choices side by side, and the well-designed ones share a rule. Stripe's verification guide signs a payload built from a timestamp, a dot, and the raw request body, with a header that carries both the timestamp and the signature; the libraries reject anything older than five minutes by default; and the documentation says, in a highlighted box, that any manipulation of the raw body causes verification to fail. Slack's scheme is the same shape: a version string, the timestamp and the raw body joined by colons, an HMAC, and a check that the timestamp is within five minutes. GitHub's deliveries sign the raw body alone, with a secret and no timestamp, and the guidance is to compare in constant time.
Not one of them signs a JSON object. They all sign bytes, and the ones that bind a timestamp sign it as part of the same byte string, so that the replay check and the integrity check are one operation. That is the design rule in its entirety: sign and verify the exact bytes on the wire, bind a timestamp into them, and never let a parsed object stand in for the bytes it came from.
The practical symptom of the gap is a verification that fails intermittently, which is the worst kind of failure to debug. It works in the test suite, where the test constructs the request and the framework's parser round-trips it cleanly. It fails in production for the one sender whose serialiser emits a unicode escape, or for the one message that contains a number with a trailing zero, and the failure is logged as a bad signature, which sends the investigation toward the keys. The tell is that the same key verifies most messages. A key that is wrong fails all of them.
When the bytes cannot be kept
There are designs where the receiver genuinely cannot get at the original bytes, because the message has passed through a system that reconstructs it, a message queue that stores objects, a database that holds the payload as a document, a second service that forwards a parsed copy. For those, the gap cannot be made empty by discipline, and it has to be closed by specification: both sides agree on a canonical serialisation, so that the bytes are a deterministic function of the content, and the signature is computed over the canonical form on both ends. RFC 8785, the JSON Canonicalization Scheme, is the standard way to do that for JSON: keys sorted, whitespace removed, numbers and strings serialised one way. A design that signs objects without a canonical form is a design that will fail on the first serialiser upgrade, and the failure will look like a broken key.
The two approaches are not equal, and the order of preference is clear. Keep the bytes if you can, because a byte-for-byte check has no serialiser in it and nothing to disagree about. Canonicalise if you must, and then show that both sides actually use the canonical form, which means a test that signs on one side and verifies on the other through the real pipeline rather than a unit test of the canonicaliser alone.
Three gates, in order
The verification code itself is three gates, and the order matters. The first is the signature, computed over the raw bytes and compared in constant time, because a comparison that returns early on the first mismatched byte leaks how many bytes matched. The second is the timestamp, checked against the receiver's clock with a tolerance measured in minutes, which turns a captured message into one that expires. The third is an idempotency key, the message's own identifier, checked against a store of recently processed ones, which turns a replay inside the window into a no-op rather than a duplicate action.
The order is not aesthetic. Checking the timestamp before the signature lets an attacker learn your clock tolerance from an unsigned message; checking the idempotency key before the signature lets an attacker fill the store with forged ids. The signature goes first because it is the only gate that costs nothing to fail, and everything after it runs only on messages that are provably from the sender.
The same gap in every signed message
Webhooks are where the gap is most visible, because the two sides are different companies and the bytes cross a network, but the gap is the same in every signed structure. A JWS token is a signature over base64 segments, and a library that decodes, reformats and re-encodes the payload before verifying has opened the gap in a place nobody looks. A signed configuration file, a signed manifest, a signed audit record stored in a database: each is a case where the bytes the signer hashed and the bytes the verifier hashes are separated by a store or a parser that may not round-trip. The rule is the same. Find the exact bytes that were signed, keep them, and verify those. If they cannot be kept, specify a canonical form and prove both sides produce it. The cryptography is the easy part. The bytes are where it fails.
Get new posts by email
Occasional essays on engineering, AI, and building for the people technology leaves behind.
Subscribe with RSS