Every AES-GCM key has a collision horizon
With a random 96-bit nonce, AES-GCM's safety is a countdown, not a property. Compute the countdown for every key, and design so that it can never be reached.
AES-GCM is the default authenticated cipher in almost every library, and it has one rule that libraries cannot enforce for you: never encrypt two messages under the same key with the same nonce. Break the rule once and the keystream repeats, the XOR of the two plaintexts is exposed, and the authentication key falls out of the tags. That is not a weakness people argue about. It is the design. What people under-appreciate is that with random nonces the rule is not a property you have; it is a countdown you are running, and the countdown has a number.
The number
GCM's standard nonce is 96 bits. If nonces are drawn at random, the chance that two of them collide grows with the square of the number of messages, the birthday bound, and NIST SP 800-38D draws the line at 2 to the 32 invocations per key for random nonce constructions, so that the collision probability stays below 2 to the negative 32. Four billion messages sounds like a lot until you are a server encrypting records, tokens or log lines, at which point it is a few weeks.
I think of it as the collision horizon: the number of messages a single key may encrypt before the nonce-collision probability crosses a tolerance you have written down. NIST's horizon is 2 to the 32 at a tolerance of 2 to the negative 32. A stricter tolerance gives a nearer horizon; a looser one, further. What matters is that every key in a design has one, that the number is known, and that a mechanism exists to stop the key before it gets there.
What a collision costs
The reason the horizon deserves a number rather than a shrug is what happens on the far side of it. GCM encrypts by generating a keystream from the key and the nonce and XORing it with the plaintext. Two messages under the same key and nonce get the same keystream, so XORing the two ciphertexts together cancels the keystream and leaves the XOR of the two plaintexts, which for structured data such as JSON, headers or file formats is usually enough to recover both. Worse, the authentication tag is computed with a hash key derived from the key alone, and two tags over the same nonce let an attacker solve for that hash key, after which they can forge tags for any message under that key. One collision does not leak one message. It breaks the key.
Where the horizon is near and where it is far
It helps to place a few systems against the number, because the horizon is far for some designs and uncomfortably close for others. A vault on a phone that encrypts files the user adds by hand will never approach four billion files, and for that design the birthday bound is not the risk. A server that encrypts every session token, every audit log line or every row written to a table under one long-lived key can reach 2 to the 32 in weeks, and a busy one in days. A messaging protocol that derives a fresh key per conversation and rotates it on a schedule sits in between, with the horizon reset at every rotation.
The rule that falls out is about key scope rather than message count. The wider the scope of a key, the more messages share it and the nearer its horizon; the narrower the scope, the further away. A key per file, per session or per record has a horizon nobody can reach. A key per deployment has one that a successful product will reach on schedule, and the day it does is not marked on any calendar unless someone put it there.
Making the horizon unreachable
There are two honest ways to deal with a countdown, and both are design decisions rather than parameter choices.
The first is to give each key so little to encrypt that the horizon cannot be reached. Derive a fresh key per file, per session or per record from a master key and a unique identifier, using a key derivation function, and the collision horizon of each derived key is one message: there is nothing to collide with. This is the design I use for the encrypted media vault in the Minimalist Calculator. The master key lives in the Android Keystore and never encrypts a file directly; each file gets its own derived key, and the nonce question disappears with the reuse it was protecting against.
The second is to make nonces that cannot repeat: a counter rather than a random draw. A counter has no birthday bound; it collides only when it is reset. But resets happen in exactly the places engineers forget to look. A counter kept in memory resets when the process restarts. A counter kept on disk resets when the disk is restored from a backup taken before the last messages were sent, and a vault whose files are backed up and restored is the textbook case. Whoever restores the backup also restores the counter, and the next file encrypts under a nonce that was already used. A counter is safe only when its persistence is more reliable than the data it protects, which is a high bar and a strange one.
There is a third option, which is to change the cipher. AES-GCM-SIV is built so that a repeated nonce leaks only whether two messages were identical, not their contents and not the key, which turns the countdown into a graceful degradation. XChaCha20-Poly1305 uses a 192-bit nonce, so that random nonces have a horizon nobody will reach. Both are good answers when the library offers them. Neither removes the need to know the horizon of the keys you already have.
The mechanism, not the policy
A written policy that says "rotate keys every quarter" is not a mechanism, and this is where many designs quietly fail the test. The horizon is counted in messages, not months, and a quarter of messages is a different number in a slow month and a busy one. The mechanism has to count. A per-key message counter that refuses to encrypt past a threshold well below the horizon, and forces a rotation, is the honest implementation, and it is a few lines. A calendar reminder is a hope.
The counter has the same persistence problem as a counter nonce, but with a gentler failure mode: if it is lost and restarts from zero, the key lives somewhat longer than intended rather than immediately reusing a nonce. That is why I would rather keep a message counter for rotation than a counter nonce for uniqueness, and rely on derived per-message keys or a misuse-resistant mode for the uniqueness itself.
Write the horizon down
The practice I follow is short. For every key in the design, write three things next to it: what nonce construction it uses, what its collision horizon is at a tolerance stated in the same line, and what mechanism stops the key before the horizon. A key with random nonces and no counter of messages has a horizon it cannot see and no mechanism, and it does not pass. A per-file derived key has a horizon of one and needs no mechanism. A counter nonce is allowed when the sentence describing what happens on restore from backup is written and true.
None of that is cryptography. The cryptography was finished when the standard was published. What remains is the engineering of a number, and the number is the difference between a system that is safe and a system that is safe so far.
Get new posts by email
Occasional essays on engineering, AI, and building for the people technology leaves behind.
Subscribe with RSS