27 February 2026 · 5 min read

The queue that joins your transaction

The reason to run background jobs from a Postgres table is not that it is cheaper than Redis. It is that the enqueue can share a transaction with the row that caused it.

The argument about where background jobs should live is usually conducted in the wrong units. One side says Redis is faster and built for it; the other says Postgres is already there and one fewer thing to run. Both are true and neither is the reason to choose. The reason to run jobs from a Postgres table is that the enqueue can sit in the same transaction as the row that caused it: the order is inserted and the job to email its confirmation is inserted in one commit, and if the commit fails, neither exists. No external broker can offer that, however fast it is, because it is not in the transaction. The folklore about throughput, meanwhile, is wrong by two orders of magnitude in both directions, so it should not be what decides anything.

The same-transaction test

Every job type answers one question: must this job not exist if the write that caused it rolled back? For a confirmation email, yes: an email about an order that was never saved is a bug. For a search-index update, yes: indexing a document that does not exist produces a phantom result. For a nightly report, no: it is scheduled by the clock, not by a write. For a cache warm, no: a warm for a row that did not commit is wasted work and nothing worse.

The same-transaction test as a decision flow, with a vacuum tiebreak A flow. For each job type: must the job not exist if the causing write rolled back? If yes, the job goes in a Postgres table, enqueued in the same transaction. If no, either store works, and the tiebreak is whether the jobs table's dead-tuple rate would outrun autovacuum; if it would, use the external broker; if not, the table is fine. Atomicity first, vacuum second, speed not at all Must not exist if the causing write rolled back? yes Postgres jobs table same transaction no Would dead tuples outrun autovacuum? yes External broker Redis, or a queue service no Either store the table, one fewer thing Throughput is absent from the flow: at any realistic rate both stores have plenty.
Illustrative: the decision as I apply it per job type; the vacuum tiebreak is discussed below.

For the yes answers, the table is not a preference; it is the only design that is correct without a second mechanism. Any external broker requires the transactional outbox pattern to get the same guarantee: write the intent to a table in the transaction, then have a relay copy it to the broker afterwards. That is a jobs table with an extra hop. If you are going to have the table anyway, the question is whether the hop buys anything, and for most workloads it does not.

The folklore is wrong both ways

The reason people reach for the broker is a number that circulates in blog posts and conference talks: that a Postgres queue tops out somewhere in the low hundreds of jobs per second. The number is not from any measurement I can find, and the measurements that exist say something else. Graphile Worker, a mature Postgres-backed job runner, publishes performance figures: about 202,000 jobs queued per second from a single batched add call, about 183,000 jobs processed per second with batching enabled and 24 concurrent jobs per worker, and around 15,600 per second without batching, on a desktop processor with the database local, with an average latency from enqueue to execution start of about four milliseconds.

Measured throughput of a Postgres-backed queue against the folk number, jobs per second on a log scale Horizontal bars on a log scale: the folk number repeated in blog posts, about 200 jobs per second; Graphile Worker without batching, about 15,600 processed per second; with batching, about 183,000 processed per second; queued per second with a batched add, about 202,000. The folk number is three orders of magnitude below the measured figures. The folk number is off by three orders of magnitude Jobs per second, log scale from 100 to 1,000,000 repeated claim measured The folk number about 200, unsourced Graphile, unbatched 15,600 processed Graphile, batched 183,000 Graphile, queued 202,000 A hundred pixels per decade; measured on one desktop with a local database. a networked production database will be slower, and still nowhere near the folk number
Source: Graphile Worker's performance page for the measured bars; the folk number is the figure commonly repeated without measurement, shown for scale.

That is the first direction the folklore is wrong: a Postgres queue is not slow. The second direction is the one that matters more for small systems. Nobody running a product with a few hundred users needs a hundred thousand jobs a second, or a thousand, or often a hundred. The workload that the throughput argument is conducted over does not exist for most of the people conducting it, and the argument distracts from the property that does matter at every scale, which is whether the job and the row commit together.

How the table works

The mechanism that makes a table a queue is one clause. Workers select the next available job with a row lock, skipping rows other workers have already locked, using SELECT FOR UPDATE SKIP LOCKED, so that two workers polling at the same instant claim disjoint rows without blocking each other. The worker runs the job, marks the row done or failed inside the same lock, and commits. A worker that crashes mid-job releases its lock when its connection dies, and the row becomes available again, which is the retry.

Two workers claiming disjoint rows with SELECT FOR UPDATE SKIP LOCKED A jobs table with five rows. Worker A locks row 1; worker B, polling at the same time, skips the locked row 1 and locks row 2. Rows 3 to 5 remain available. Each worker processes its row and marks it done in the same transaction, then commits. One clause, no coordinator JOBS TABLE row 1: send confirmation, pending row 2: update search index, pending row 3: resize upload, pending row 4: send confirmation, pending row 5: warm cache, pending Worker A FOR UPDATE SKIP LOCKED LIMIT 1: row 1 Worker B, same instant skips locked row 1, claims row 2 Each marks its row done and commits; a crashed worker's row returns to pending.
Illustrative: the claiming pattern from the PostgreSQL SELECT documentation, drawn for two workers.

Two details make the table behave. The first is the index: the claim query filters on status and orders by a run-at time, and without an index on exactly that pair it scans, which is where the slowness stories come from at surprisingly small sizes. The second is the poll. Workers that poll every few hundred milliseconds are fine at small scale and wasteful at large; Postgres's notify mechanism lets a worker sleep until a row is inserted and wake immediately, which turns the poll into a fallback and drops the enqueue-to-start latency to the few milliseconds the measurements show. Neither detail is exotic. Both are the difference between the folk number and the measured one.

The real cost is vacuum

If speed is not the cost of a Postgres queue, what is? Dead tuples. Every job row is inserted, updated at least once when claimed, updated again when finished, and usually deleted, and each update leaves a dead version of the row behind for autovacuum to reclaim. A jobs table at a steady rate produces dead tuples at a multiple of that rate, and autovacuum's default trigger fires when dead tuples exceed a threshold of fifty rows plus twenty percent of the table, which on a small, hot table means it fires constantly, and on a table with a large backlog of old completed jobs means it fires rarely relative to the churn on the live rows.

Dead tuples against job rate under default autovacuum thresholds, from a simple model Two lines against jobs per second. With completed jobs deleted promptly and the table kept small, dead tuples stay bounded, because autovacuum triggers often on a small table. With completed jobs retained, the table grows, the twenty percent threshold grows with it, vacuum runs less often relative to churn, and dead tuples climb steadily with the job rate until bloat dominates. Keep the table small and vacuum keeps up Dead tuples between vacuums by job rate; two policies completed rows deleted retained high 0 1 a second 100 10,000 bounded: vacuum fires often on a small table the threshold grows with the table
Illustrative: a model of dead-tuple accumulation under the default autovacuum threshold, fifty rows plus twenty percent of the table, for two retention policies; the shapes follow from the threshold rule and the values are not measured.

So the tiebreak in the decision flow is not speed but vacuum. A jobs table that deletes completed rows promptly, or partitions them by day and drops old partitions, stays small, and autovacuum on a small table is cheap and frequent. A jobs table that keeps every completed job for audit grows without bound, its vacuum threshold grows with it, and the live rows at the head of the table sit among an increasing pile of dead versions. The second design is where the "Postgres queues are slow" folklore probably came from: not from Postgres, but from a table nobody ever cleaned.

The retention question also settles the audit argument. A jobs table that doubles as a history of every job ever run is doing two jobs, and the second one is what bloats it. The history belongs in a separate table, written once per completed job by the worker in the same transaction that deletes the live row, so that the live table stays small and the history table is append-only, which is the shape autovacuum handles best. Partitioning the history by month and dropping old partitions is a single statement, and it is the entire retention policy.

What I actually do

In the Node.js and PostgreSQL backend I run, the jobs table holds the job types that pass the same-transaction test, which is most of them: anything caused by a write. The enqueue is a row insert in the same transaction as the write, workers claim with SKIP LOCKED, and completed rows are moved to a history table by a nightly job so that the live table stays at a few hundred rows and vacuum never notices it. Jobs that are scheduled by the clock rather than by a write go through the same table for the sake of having one place to look, since the vacuum load at those rates is nothing. The broker is the thing I would add for a workload that fails the vacuum tiebreak, a job rate high enough that even a small table's churn outruns autovacuum, and I have not needed it. The decision was never about speed. It was about what commits together.

PostgreSQLBackground JobsRedis
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS