22 August 2025 · 5 min read

Timeouts run in two directions

Through a proxy chain, idle timeouts must get longer as you go inward and deadlines must get shorter. Most timeout bugs are one of those staircases built the wrong way round.

The random 502 and the hung connection pool are the two timeout bugs every backend engineer meets, and they are usually treated as unrelated. They are the same bug in two directions. A request passes through a chain of hops, a load balancer, a reverse proxy, an application server, a database, and every hop has two timeouts on it: how long it will keep an idle connection open, and how long it will wait for a response. Those two numbers have to move in opposite directions along the chain, and when either one does not, the failure has a name.

Two staircases

Draw every hop from the client to the database and write two numbers on each. The first is the idle timeout: how long this hop keeps a connection open with nothing happening on it. The second is the deadline: how long this hop waits for the hop behind it to answer before giving up.

The idle staircase must ascend inward. Each hop's idle timeout must be longer than the timeout of the hop in front of it, so that the outer hop is always the one that closes an idle connection. If an inner hop closes first, the outer hop may already have written the next request onto a socket the inner hop just shut, and the outer hop reports that as a bad gateway.

The deadline staircase must descend inward. Each hop's deadline must be shorter than the deadline of the hop in front of it, so that the innermost hop is always the one that gives up first and the failure propagates outward as a clean error. If an inner hop waits longer than an outer one, the outer hop has already sent an error to the client while the inner hop is still working, holding a worker, a connection, and a slot in the pool for a response nobody will read.

Documented default timeouts along a common chain Four hops on the horizontal axis: an AWS application load balancer, nginx, a Node.js HTTP server and PostgreSQL. Two series on a logarithmic seconds axis. Idle timeouts: 60 at the load balancer, 75 at nginx, 5 at Node, which breaks the ascending rule and is highlighted. Deadlines: 60 at the load balancer, 60 at nginx, 300 at Node, and no limit at PostgreSQL, which breaks the descending rule and is highlighted. Defaults, before anyone sets anything Seconds, log scale; the two series should step opposite ways idle timeout deadline 1,000 100 10 1 ALB nginx Node.js PostgreSQL Hops, from the client inward 5 s: closes before nginx 300 s: outlives the proxy no limit 60 75 / 60
Source: default values from the AWS ALB documentation (idle timeout 60 s), the nginx keepalive_timeout (75 s) and proxy_read_timeout (60 s) directives, the Node.js HTTP server (keepAliveTimeout 5 s, requestTimeout 300 s), and PostgreSQL's statement_timeout (0, disabled).

The chart is the defaults, and both staircases are broken out of the box. That is not a criticism of any of the projects; each default is sensible for the component on its own. The bugs are in the combination, and nobody ships the combination.

The 502 is the idle staircase reversed

Node's HTTP server closes an idle keep-alive connection after five seconds by default, per its server.keepAliveTimeout documentation. nginx keeps an idle upstream connection for longer than that, and an application load balancer in front of nginx holds client connections for sixty, its documented default. So a request arrives at nginx, nginx picks an upstream connection that has been idle for five seconds and a few milliseconds, writes the request onto it, and Node has already closed its end. nginx sees a connection reset with no response, and the client gets a 502. It happens on a small fraction of requests, at random, more often under light load than heavy, which is the signature that sends people looking at the application code for a bug that is not there.

The race that produces a random 502 Three lanes: nginx, the idle upstream socket, and Node. Node closes the socket at five seconds of idleness. A few milliseconds later nginx writes a new request onto the same socket, receives a reset, and returns 502 to the client. Setting Node's keep-alive timeout above nginx's fixes the order. nginx Idle upstream socket Node.js last response, t = 0; socket goes idle t = 5,000 ms: keepAliveTimeout, FIN t = 5,004 ms: new request written reset, no response 502 Bad Gateway to the client
Illustrative: the ordering of events that the default values allow; the fix is to make Node the last to close.

The fix is one line, and it is the line the staircase tells you to write: set Node's keep-alive timeout above nginx's, and nginx's above the load balancer's. Node also has a headers timeout that must sit above its keep-alive timeout, which the Node documentation notes, because a keep-alive connection that has just been reused is waiting for headers, and the two timers interact. Once the idle staircase ascends inward, the outermost hop always closes first, and a socket is never reused after the far end has left.

The hung pool is the deadline staircase reversed

The other direction is quieter and worse. PostgreSQL's statement_timeout is disabled by default, so a query that hits a lock or a bad plan runs until it finishes, however long that is. nginx's read timeout is sixty seconds, so after a minute nginx returns a 504 to the client and forgets the request. But the request is not over. Node is still waiting on the database, the database is still running the query, and the connection is still checked out of the pool. Repeat that a dozen times in a bad minute and the pool is empty, every new request queues behind a connection that will not come back for a while, and the service is down for a reason that no log line will state directly.

The fix is again what the staircase says: a statement timeout at the database shorter than the application's request deadline, which is shorter than the proxy's read timeout, which is shorter than the load balancer's. Then the database is always the first to give up, the application receives a clean error and releases the connection, the proxy receives a clean error from the application, and the client receives a real status instead of a silence that turned into a 504.

The two staircases, built the right way round A table with four hops as rows, load balancer, nginx, Node, PostgreSQL, and two columns. The idle column rises inward: 60, 75, 80 seconds, and not applicable for the database. The deadline column falls inward: 60, 30, 20, 15 seconds. Example values, chosen to satisfy the rules. Idle ascends inward, deadline descends inward Example values in seconds that satisfy both rules HOP IDLE TIMEOUT DEADLINE Load balancer 60 60 nginx 75 30 Node.js 80 20 PostgreSQL not a hop it holds 15 Outermost closes idle sockets first; innermost gives up on slow work first.
Illustrative: one consistent assignment; the specific numbers matter less than their order.

Why the deadline staircase is the harder one

The idle staircase is a configuration problem: four numbers in four config files, checked once. The deadline staircase is a design problem, because a deadline the innermost hop enforces has to be a deadline the application can live with. Setting a statement timeout of fifteen seconds means every query in the system must finish in fifteen seconds, including the reporting query somebody wrote last spring. That is the correct constraint, and enforcing it surfaces the queries that were quietly taking a minute. But it surfaces them as errors, and a team that is not ready for that will raise the timeout back to nothing and lose the staircase.

The way through is to set the innermost deadline first, at the value the product needs, and to treat every query that breaks it as a defect to fix rather than a reason to loosen the number. The QR ticketing system's gate check is the sharpest version of this I have built: the deadline is however long a visitor will stand at a gate, the database has to answer inside it, and a query that cannot is not a slow query. It is a broken gate.

One more row belongs on the deadline staircase, and it is the one that hides. Health checks are requests too, and the load balancer's health check has its own timeout, usually a few seconds. If the application answers its health endpoint by touching the database, and the database's statement timeout is longer than the health check's, a slow database makes the health check time out, the load balancer marks the instance unhealthy, and traffic moves to the other instances, which are talking to the same slow database. The staircase says the health check's deadline is the shortest of all, and a health endpoint that cannot answer inside it should not be doing the work that made it slow.

The check on paper

Both bugs can be found before they fire, without load tests, by drawing the chain and writing the two numbers on each hop. Read the idle column from the outside in: every number must be larger than the one before it. Read the deadline column from the outside in: every number must be smaller. Any step where a column flattens or reverses is a bug, and it is a bug you found in five minutes with a pen. Most timeout bugs I have seen were one staircase built the wrong way round, and every one of them was visible on paper the whole time.

Reverse ProxiesNode.jsReliability
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS