Stateless tokens still need a kill switch
Every token system has a revocation lag, and a JWT sets it to the token lifetime. Write the allowed lag down per endpoint class and the stateless debate turns into arithmetic.
The stateless-versus-stateful token argument has been running for a decade and it is the wrong argument. Every token system has a revocation lag: the longest time a principal whose access has been revoked can keep acting. A session table sets that lag to one database round trip. A signed, self-contained token sets it to the token's lifetime, because nothing checks anything until the token expires. Neither number is right or wrong in itself. What is wrong is not knowing the number, and most systems do not, because the lifetime was set once by whoever configured the identity provider and never connected to what the endpoints behind it actually do. The fix is to write the acceptable lag down per endpoint class, as a number of seconds, next to the route, and to choose mechanisms that fit under it.
What the providers set for you
The lifetime most teams inherit is whatever their identity provider's default is, and those defaults span two orders of magnitude. Auth0's API access tokens default to 86,400 seconds, a full day, and can be raised to thirty. Microsoft Entra issues access tokens with a random lifetime between 60 and 90 minutes. Amazon Cognito lets an access token live anywhere from five minutes to a day, with an hour as the usual setting. Firebase ID tokens last an hour, and the same page is candid that revocation is not detected until the token is next refreshed unless the server explicitly asks. A GitHub App's installation token expires after an hour.
None of those defaults is a decision about your endpoints. An hour is a reasonable lag for reading a dashboard and an unreasonable one for moving money, and a single token lifetime cannot be both. The provider does not know which routes you have. You do.
The revocation budget
So the number goes on the route. For each class of endpoint, declare a revocation budget: the number of seconds a revoked identity may still be accepted there. It sits next to the rate limit and the required scopes, and it is enforced by a constraint that is easy to state: the token lifetime, plus any cache the resource server keeps of introspection results or key sets, plus the interval of whatever online check the route performs, must sum to less than the budget. The token lifetime alone is the lag when there is no online check; the online check's interval is the lag when there is one; caches add to whichever applies.
For the CRM I built, the classes fall out quickly. Reads of a user's own pipeline and contacts can tolerate an hour: the harm of a revoked user reading their own data for another hour is small and the cost of checking every read is not. Writes and reads of shared records get five minutes, met by a five-minute access token and refresh rotation, which means nothing on those routes ever calls the identity provider. The routes that move money or change permissions, exporting a whole tenant's contacts, changing a role, connecting a payment account, get ten seconds, which no token lifetime meets, so those routes perform an online check on every call: an introspection request, or a lookup in a denylist of revoked subjects that is updated synchronously when a revocation happens. And the issuance path itself has a budget of zero, because a stolen refresh token is a stolen identity for as long as it works.
The mechanism most routes need
The mechanism that covers the middle rows is refresh token rotation with reuse detection, and it is worth drawing because it is where the stateless system keeps its state. Access tokens are short and self-contained. Refresh tokens are opaque, stored server-side in families, and rotated on every use: the client presents a refresh token, receives a new access token and a new refresh token, and the old refresh token is marked used. If a used refresh token is ever presented again, someone has a copy, and the whole family is revoked, which logs out both the thief and the victim and forces a fresh login.
Rotation bounds the damage from a stolen refresh token to the time until either party next refreshes, and the reuse detection is what makes the budget on the issuance row zero rather than "until the refresh token expires". It does require the authorisation server to keep state, a table of token families, and that is the honest answer to the stateless purists: the system is stateless where the budget allows and stateful where it does not, and the budget is what says which is which.
Where the online check goes
The ten-second row needs a check on every call, and there are two honest ways to build it. The first is token introspection as RFC 7662 defines it: the resource server sends the token to the authorisation server and asks whether it is still active. It is simple, it is a network round trip on every protected call, and its lag is the round trip plus whatever the authorisation server caches internally, which has to be known and added to the row.
The second is a revocation list held close to the resource servers. When a principal is revoked, the authorisation server writes the subject and a revoked-at timestamp to a small shared store, replicated to every resource server within a second or two. A resource server checking a token compares the token's issued-at claim with the subject's revoked-at time, and rejects any token issued before the revocation. The list stays small because an entry only needs to live for the longest access token lifetime in the system, after which every token it could affect has expired on its own. The lag is the replication delay, and the check is a local lookup rather than a round trip.
Both mechanisms have one dependency that is easy to forget: the key set. A resource server verifies signatures against the authorisation server's published keys, and it caches them. If a signing key is compromised and rotated, the cache's time-to-live is a revocation lag of its own, on every row, and it belongs in the sum. A day-long key cache under a ten-second budget is the kind of mismatch the budget is there to catch.
Arithmetic, not debate
The reason to write the budget down rather than argue about architecture is that it turns every later decision into a check. Someone proposes raising the access token lifetime to a day to reduce load on the identity provider: which routes have a budget under a day, and what online check will they gain to compensate? Someone adds a cache of introspection results with a five-minute TTL in front of the money-moving routes: the sum on that row is now five minutes plus the token lifetime, which is over the ten-second budget, so the cache is wrong for that row and fine for the writes row. Someone proposes dropping the family table because it is state: the issuance row's budget is zero, and nothing stateless meets zero.
Each of those is a two-line calculation instead of a meeting, and the calculation has the same shape as a rate limit, which is why the budget belongs where the rate limit lives, in the route's declaration, where the next engineer will see it. Revocation is not a feature the identity provider gives you. It is a number you owe each endpoint, and once the number is written down, the mechanism chooses itself.
Get new posts by email
Occasional essays on engineering, AI, and building for the people technology leaves behind.
Subscribe with RSS