Timeouts run in two directions
Through a proxy chain, idle timeouts must get longer as you go inward and deadlines must get shorter. Most timeout bugs are one of those staircases built the wrong way round.
The random 502 and the hung connection pool are the two timeout bugs every backend engineer meets, and they are usually treated as unrelated. They are the same bug in two directions. A request passes through a chain of hops, a load balancer, a reverse proxy, an application server, a database, and every hop has two timeouts on it: how long it will keep an idle connection open, and how long it will wait for a response. Those two numbers have to move in opposite directions along the chain, and when either one does not, the failure has a name.
Two staircases
Draw every hop from the client to the database and write two numbers on each. The first is the idle timeout: how long this hop keeps a connection open with nothing happening on it. The second is the deadline: how long this hop waits for the hop behind it to answer before giving up.
The idle staircase must ascend inward. Each hop's idle timeout must be longer than the timeout of the hop in front of it, so that the outer hop is always the one that closes an idle connection. If an inner hop closes first, the outer hop may already have written the next request onto a socket the inner hop just shut, and the outer hop reports that as a bad gateway.
The deadline staircase must descend inward. Each hop's deadline must be shorter than the deadline of the hop in front of it, so that the innermost hop is always the one that gives up first and the failure propagates outward as a clean error. If an inner hop waits longer than an outer one, the outer hop has already sent an error to the client while the inner hop is still working, holding a worker, a connection, and a slot in the pool for a response nobody will read.
The chart is the defaults, and both staircases are broken out of the box. That is not a criticism of any of the projects; each default is sensible for the component on its own. The bugs are in the combination, and nobody ships the combination.
The 502 is the idle staircase reversed
Node's HTTP server closes an idle keep-alive connection after five seconds by default, per its server.keepAliveTimeout documentation. nginx keeps an idle upstream connection for longer than that, and an application load balancer in front of nginx holds client connections for sixty, its documented default. So a request arrives at nginx, nginx picks an upstream connection that has been idle for five seconds and a few milliseconds, writes the request onto it, and Node has already closed its end. nginx sees a connection reset with no response, and the client gets a 502. It happens on a small fraction of requests, at random, more often under light load than heavy, which is the signature that sends people looking at the application code for a bug that is not there.
The fix is one line, and it is the line the staircase tells you to write: set Node's keep-alive timeout above nginx's, and nginx's above the load balancer's. Node also has a headers timeout that must sit above its keep-alive timeout, which the Node documentation notes, because a keep-alive connection that has just been reused is waiting for headers, and the two timers interact. Once the idle staircase ascends inward, the outermost hop always closes first, and a socket is never reused after the far end has left.
The hung pool is the deadline staircase reversed
The other direction is quieter and worse. PostgreSQL's statement_timeout is disabled by default, so a query that hits a lock or a bad plan runs until it finishes, however long that is. nginx's read timeout is sixty seconds, so after a minute nginx returns a 504 to the client and forgets the request. But the request is not over. Node is still waiting on the database, the database is still running the query, and the connection is still checked out of the pool. Repeat that a dozen times in a bad minute and the pool is empty, every new request queues behind a connection that will not come back for a while, and the service is down for a reason that no log line will state directly.
The fix is again what the staircase says: a statement timeout at the database shorter than the application's request deadline, which is shorter than the proxy's read timeout, which is shorter than the load balancer's. Then the database is always the first to give up, the application receives a clean error and releases the connection, the proxy receives a clean error from the application, and the client receives a real status instead of a silence that turned into a 504.
Why the deadline staircase is the harder one
The idle staircase is a configuration problem: four numbers in four config files, checked once. The deadline staircase is a design problem, because a deadline the innermost hop enforces has to be a deadline the application can live with. Setting a statement timeout of fifteen seconds means every query in the system must finish in fifteen seconds, including the reporting query somebody wrote last spring. That is the correct constraint, and enforcing it surfaces the queries that were quietly taking a minute. But it surfaces them as errors, and a team that is not ready for that will raise the timeout back to nothing and lose the staircase.
The way through is to set the innermost deadline first, at the value the product needs, and to treat every query that breaks it as a defect to fix rather than a reason to loosen the number. The QR ticketing system's gate check is the sharpest version of this I have built: the deadline is however long a visitor will stand at a gate, the database has to answer inside it, and a query that cannot is not a slow query. It is a broken gate.
One more row belongs on the deadline staircase, and it is the one that hides. Health checks are requests too, and the load balancer's health check has its own timeout, usually a few seconds. If the application answers its health endpoint by touching the database, and the database's statement timeout is longer than the health check's, a slow database makes the health check time out, the load balancer marks the instance unhealthy, and traffic moves to the other instances, which are talking to the same slow database. The staircase says the health check's deadline is the shortest of all, and a health endpoint that cannot answer inside it should not be doing the work that made it slow.
The check on paper
Both bugs can be found before they fire, without load tests, by drawing the chain and writing the two numbers on each hop. Read the idle column from the outside in: every number must be larger than the one before it. Read the deadline column from the outside in: every number must be smaller. Any step where a column flattens or reverses is a bug, and it is a bug you found in five minutes with a pen. Most timeout bugs I have seen were one staircase built the wrong way round, and every one of them was visible on paper the whole time.
Get new posts by email
Occasional essays on engineering, AI, and building for the people technology leaves behind.
Subscribe with RSS