15 September 2026 · 5 min read

Two dials on the rate limiter

A per-tenant limit stops tenants hurting each other; a per-principal limit stops one bug hurting you. They come from different numbers, and one dial with one number gives neither.

Most rate limiters in multi-tenant software are one middleware with one number, and the number was chosen by someone who had to pick something. It is usually described as protecting the platform, which it does, badly, and as being fair to tenants, which it is not. The reason is that a rate limit is two protections wearing one hat. A per-tenant limit exists so that tenants cannot hurt each other. A per-principal limit, per user or per API key, exists so that one runaway client cannot hurt you. Those are set from different numbers, derived from different facts about the system, and putting both behind a single dial means the dial is wrong for at least one of them at every setting.

What the public limits are made of

Every large API publishes its limits, and reading them side by side shows the two dials in the open. GitHub's REST API allows an authenticated user 5,000 requests an hour, 60 unauthenticated, and 15,000 an hour for Enterprise Cloud, with a separate secondary limit of no more than 100 concurrent requests, 900 points a minute per endpoint, and 90 seconds of CPU per 60 seconds of real time. Stripe's global limit is 100 requests a second per account in live mode, 25 in a sandbox, with 25 a second on individual endpoints and separate concurrency limits. Shopify's GraphQL Admin API meters query cost in points from a leaky bucket that refills at 100 points a second on standard plans and 1,000 on Plus. Slack's Web API tiers its methods at one, twenty, fifty or a hundred-plus calls a minute, with a special rule that a message may be posted about once a second per channel.

Published rate limits of four public APIs, converted to requests per second Horizontal bars on a logarithmic scale of requests per second: Stripe live mode, 100; Stripe per endpoint, 25; GitHub Enterprise Cloud, 15,000 an hour, about 4.2; Slack tier 4, 100 a minute, about 1.7; GitHub authenticated, 5,000 an hour, about 1.4; Slack tier 1, one a minute, about 0.017. Shopify is noted separately because it meters points rather than requests. Four vendors, four different answers to a different question each Published limit converted to requests per second, log scale from 0.01 to 100 Stripe, live, per account 100 a second Stripe, per endpoint 25 GitHub, Enterprise Cloud 4.2, from 15k an hour Slack, tier 4 1.7, from 100 a minute GitHub, authenticated user 1.4, from 5,000 an hour Slack, tier 1 0.017, from one a minute keyed by account: the tenant dial keyed by user, token or method: the principal dial
Source: the rate-limit documentation of Stripe, GitHub and Slack, September 2026; Shopify's point-based limits are omitted because they do not convert to requests.

Read closely, the vendors are not disagreeing about the right number. They are answering different questions. Stripe's 100 a second is keyed by account, which is the tenant: it is a share of Stripe's capacity that one customer may consume. GitHub's 5,000 an hour is keyed by user or token, which is the principal: it is the rate at which one client may reasonably act, and the secondary limits on concurrency and CPU are the platform protecting itself from one badly written integration regardless of who owns it. Slack's tiers are per method, a third axis, set by how expensive each call is to serve. None of these vendors has one number, and the reason is that no single number answers both questions.

The two numbers, and where they come from

The per-tenant number comes from capacity. If the platform can serve a certain number of requests a second at the latency it promises, and a certain number of tenants are active at once, then capacity divided by active tenants is the rate every tenant can be guaranteed when everyone is busy at the same time. I call that the fairness floor. It is the number that goes in the contract, because it is the one the platform can honour under full saturation, and it is computed from two facts the platform knows about itself: what it can serve and how many are being served.

The per-tenant ceiling, the limit actually enforced, sits above the floor by exactly the burst the platform can absorb. Most tenants are idle most of the time, so the capacity not being used by idle tenants is available to busy ones, and the ceiling lets a busy tenant use it. But the ceiling is not a promise; it is an opportunity, and the difference between ceiling and floor is what shrinks when many tenants are busy at once. Contracts that quote the ceiling are contracts that break on the busiest day of the year.

An edge limiter keyed by tenant in front of an application limiter keyed by principal, with the source of each number A flow from left to right. Requests enter an edge limiter keyed by tenant, whose number comes from capacity divided by active tenants plus absorbable burst. They then reach an application limiter keyed by user or API key, whose number comes from the fastest legitimate client. Then the application. Two callouts show the number sources. Two limiters, two keys, two sources for the number Requests all tenants Edge limiter keyed by tenant tenant against tenant Application limiter keyed by user or key contains one client's bug App capacity divided by active tenants plus the burst you can absorb the floor is the contract the fastest legitimate client measured, with headroom a bug outruns any user One middleware with one number is one of these two, mislabelled as both.
Illustrative: the layering as I build it; the edge limiter is the one the contract references, the application limiter is the one that catches the retry loop.

The per-principal number comes from a different fact entirely: how fast the fastest legitimate client actually goes. A person clicking through a dashboard makes a request every few seconds. A well-written integration syncing records makes a handful a second. A retry loop with no backoff makes hundreds. The per-principal limit is set just above the legitimate maximum, measured rather than guessed, and its job is to stop the retry loop, the runaway script, the integration that fetches a list on every keystroke, before the damage reaches the tenant limit and gets billed to every other user of the same tenant. It has nothing to do with capacity. A platform with infinite capacity would still want it, because the bug is also costing the tenant money and filling their audit log.

Why one dial fails both ways

Set the single number from capacity and it is far above any legitimate client's rate, so the retry loop runs freely until it exhausts the tenant's share, at which point every user in that tenant is locked out by one colleague's bug. Set it from the fastest client and it is far below what a large tenant with many users legitimately needs in aggregate, so the tenant hits it during normal operation and the platform's answer is to raise it, which moves the number away from the client rate and back toward the first failure. Teams that have one dial oscillate between these two settings for years, and every incident review proposes moving it.

The algorithm is the smaller decision

The choice between a token bucket and a sliding window matters less than the choice of keys and numbers, but it is worth one figure, because the two respond differently to the same burst. A token bucket lets a client spend saved-up capacity in a burst up to the bucket size and then holds it to the refill rate; a sliding window counts requests over a trailing interval and rejects once the count is reached, with no notion of saved capacity. For the tenant dial the bucket is right, because burst is precisely what the ceiling above the floor is for. For the principal dial the window is often better, because a client that has been idle for an hour and then fires a thousand requests in a second is more likely to be a bug than a legitimate burst, and the window rejects it while a large bucket would wave it through.

Token bucket and sliding window responding to the same burst Two lines over time showing requests admitted per second during a burst. The token bucket line rises to the burst rate for a short interval, draining saved tokens, then falls to the refill rate. The sliding window line rises only to the window's rate and stays flat, rejecting the excess from the start. A shaded band marks the burst. Same burst, two shapes of admission Requests admitted per second during a burst; illustrative token bucket sliding window burst limit 0 burst begins and ends at the band bucket: spends, then holds window: holds throughout
Illustrative: the two algorithms' admission curves for one burst; the shapes are the point, not the values.

The floor goes in the contract

On the platform I run, the practical output of this is two limiters and one document. The edge limiter is keyed by tenant, with a ceiling set from capacity and a floor computed from capacity divided by active tenants, and the floor is the number written into the service description, because it is the one that holds when everyone is busy. The application limiter is keyed by user and by API key, set from measured client rates with headroom, and its rejections are logged with the principal, because a principal hitting it is a bug report with a name attached. The two are tuned separately, by different people, from different dashboards, and neither one is ever described as "the rate limit".

The document is the part most teams skip. It states the floor, the ceiling, and the per-principal limit, and it says which is a promise and which is a courtesy. A tenant reading it knows what they are guaranteed and what they might get; an engineer reading it knows which dial an incident is about. That is what two dials buy: not a better number, but the ability to say which question a number is answering.

Rate LimitingAPI DesignMulti-tenant SaaS
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS