Postgres counts processes, not requests
Connection exhaustion under serverless is a topology problem. Postgres runs one backend per connection, every instance carries a pool, and the platform sets the instance count.
The error arrives as "too many connections", and the first instinct is to raise the limit. The limit is not the problem. Postgres runs one operating-system process per connection, and the number of connections is not something your code decides; it is the product of two numbers, only one of which you control. Every instance of your application carries a pool, and the number of instances is set by whoever scales your compute. Under a serverless platform that is the platform, and it can decide to run fifty copies of your function in the same second. Fifty pools of ten is five hundred backends asking for a database that was provisioned for a hundred.
The pool product
I keep this as an equation because it is the only way the conversation stays honest. The pool product is the number of instances that can run at once, times the maximum connections each instance's pool will open. It has to stay below the database's connection ceiling, minus the connections reserved for administration, replication and monitoring, or the database will refuse connections at exactly the moment traffic is highest. Nobody should set a pool size until they can write the product down, and most teams cannot, because they do not know the first factor.
Postgres itself ships with a max_connections default of 100, and a managed provider sizes the ceiling to the compute it sells. Supabase's pooling and limits documentation gives its smallest compute about 60 direct connections and 200 client connections through its transaction-mode pooler, with larger tiers scaling both. Those are the ceilings the product has to fit under.
The multiplier changed again
The classic version of this problem was the one-request-per-instance model of early serverless functions: each invocation got its own instance, each instance opened its own connection, and a burst of a thousand requests was a thousand connections. The classic advice, a pool size of one per function and a pooler in front of the database, came from that world. Vercel's Fluid compute changed the multiplier, and its knowledge base on connection pools explains how: an instance now serves concurrent requests and stays warm between them, so a pool initialised at module scope survives across invocations and is shared by the requests an instance handles. The platform also provides an attachDatabasePool helper that closes idle connections before an instance is suspended, so that suspended instances do not hold connections the database still counts.
That is better, and it means the 2022 arithmetic is wrong in both directions. A pool size of one is now too small, because one instance serves many requests at once and they would queue on a single connection. A pool size of ten is too large if the instance count can still climb into the dozens. The right P is a function of the concurrency per instance, and the right N is whatever the platform's scaling limits say, and the product still has to be written down.
Why raising the limit makes it worse
The instinct to raise max_connections deserves a paragraph of its own, because it is not merely ineffective; it is harmful in a way that hides for a while. A Postgres backend is a process with its own memory, and a server sized for a hundred of them will run five hundred only by swapping or by starving the shared buffers that make queries fast. The connections are accepted, the error goes away, and the database gets slower for everyone, including the instances that were behaving. Then the slow queries hold their connections longer, the pools open more, and the new limit is reached too, at a higher level of pain.
The reason Postgres counts processes rather than requests is that a process is the unit that costs memory and scheduler time, and the count is a statement about what the machine can do. A limit set from the machine is a limit that means something. A limit raised to make an error message stop is a limit that has been disconnected from the thing it measured.
The only lever that changes the product
Every setting in the application changes one factor of the product. The pooler changes the product itself. A transaction-mode pooler such as PgBouncer or Supavisor accepts a large number of client connections, hundreds or thousands, and multiplexes their transactions onto a small fixed number of server connections to Postgres, so that the database sees the server pool's size regardless of how many instances are connected to the pooler. The topology becomes N times P client connections to the pooler, which is cheap, and a fixed number of backends behind it, which is what the database can actually afford.
That is why the pooler is not optional under serverless, and why it was never really about performance. It is a fixed point in a topology whose other numbers move. Behind it, the pool product still matters, because the pooler has its own client ceiling and a burst of instances can exhaust that too, but the ceiling is an order of magnitude higher and the failure is a queue rather than a refused connection.
Two things the pooler cannot do are worth knowing before relying on it. Transaction mode breaks session state: prepared statements, session-level settings and advisory locks held across transactions do not survive the multiplexing, and an ORM that uses them needs to be told. And the pooler does not fix a pool that never returns connections; an instance that opens a connection and holds it through an idle period is still holding one of the pooler's server slots, which is precisely the case the platform's idle-close helper exists for.
Reading the failure from the symptoms
Connection exhaustion presents in two ways, and the topology tells you which factor of the product is to blame. A sharp cliff at the moment of a traffic burst, "too many connections" errors arriving in a cluster and then clearing, is the instance count: the platform scaled out, N jumped, and the product crossed the ceiling. A slow climb over hours or days, with the count of open connections rising while traffic stays flat, is the pool: instances are opening connections and not returning them, usually because something holds a connection across an await, or because an idle instance was suspended with its pool open. The first is fixed with a pooler and a cap on N. The second is fixed in the application, and no pooler will save a pool that leaks.
How I set it now
SocialSure's platform runs Prisma and Drizzle against Postgres from containers, which is a kinder topology than serverless because the instance count is mine, but the discipline is the same. First, find N: the maximum instances the platform will run, from its scaling configuration or its documented limit, and if it is unbounded, treat that as the first bug. Second, find the ceiling: max_connections minus the reserved slots, from the database's documentation, not from memory. Third, put a transaction-mode pooler in front of the database and point every instance at it. Fourth, set P from the concurrency each instance actually serves, with a ceiling such that N times P stays under the pooler's client limit, and write the product in the deployment configuration next to the pool size, so that the next person who raises P sees the multiplication they are doing.
The error was never that the limit was too low. It was that the product was never written down, and a number nobody wrote down is a number the platform will choose for you.
Get new posts by email
Occasional essays on engineering, AI, and building for the people technology leaves behind.
Subscribe with RSS