3 September 2025 · 6 min read

There are two gaps, not one

Every write-then-publish has two places to crash. The outbox pattern closes only the first, so a team that stops there has traded lost events for duplicated ones.

The bug that teaches most teams about event-driven systems is the lost event: the database commit succeeded, the publish to the broker failed, and a downstream service never learned that an order existed. The fix everybody reaches for is the outbox pattern, and it works. Then a second bug arrives, quieter than the first, and it turns out the fix had moved the problem rather than removed it. There are two gaps in every write-then-publish pipeline, not one, and the outbox closes exactly one of them.

Gap one: between the commit and the publish

The first gap is the one the outbox is for. A service writes a row and then publishes an event, and between those two actions the process can die, the network can fail, or the broker can refuse. If the write committed and the publish did not, the event is lost. If the publish happened and the write then rolled back, the event describes something that never occurred. The two operations are on different systems and cannot share a transaction.

The outbox closes the gap by putting the event in the same transaction as the row. The service writes the order and, in the same commit, writes an outbox row describing the event. A relay reads the outbox and publishes to the broker, marking rows as sent. Now the event cannot be lost, because it was committed with the data, and it cannot describe a rollback, because it rolled back too. The relay can crash and resume, and it will publish anything it had not marked.

Gap two: between the side effect and the acknowledgement

Which brings the second gap. The relay published the event, the consumer received it, the consumer sent the confirmation email, and then the consumer crashed before acknowledging the message. The broker, not having seen the acknowledgement, delivers the message again. The consumer sends the email again. The outbox made the event exactly-once in the sense that it exists once in the log. It did nothing for the consumer, whose side effect and whose acknowledgement are, again, on different systems and cannot share a transaction.

The write-then-publish pipeline with both gaps marked Five stages left to right: the service writes the row and an outbox row in one transaction; the relay publishes to the broker; the broker delivers to the consumer; the consumer performs its side effect; the consumer acknowledges. Gap one is marked between commit and publish, closed by the outbox. Gap two is marked between the side effect and the acknowledgement, and is not closed by the outbox. Two places to crash The outbox closes the first; only an idempotent consumer closes the second Write row and outbox row Relay publishes Broker delivers Side effect email, charge Ack to broker gap 1 closed by the outbox gap 2 still open A crash in gap 2 means the side effect happens again when the message is redelivered. draw both gaps on the diagram before choosing any tooling
Illustrative: the pipeline as it exists in any service that writes to a database and publishes to a broker.

A team that stops at the outbox has made a trade it did not intend: lost events for duplicated ones. Duplicates are usually less damaging than losses, which is why the trade feels like progress, but a duplicated charge or a duplicated dispatch is still a bug, and it is a bug that appears only under failure, which is the hardest kind to reproduce.

The two-gap rule

So the rule I apply to every pipeline is short. It is correct only when gap one is closed by a same-transaction record and gap two is closed by a same-transaction dedupe. The first is the outbox row. The second is a processed-message table on the consumer's side: before performing the side effect, the consumer inserts the message id into a table with a unique constraint, in the same transaction as whatever database change the side effect involves. A redelivered message hits the constraint and is acknowledged without being processed. Where the side effect is not a database change, an email, a call to a payment provider, the dedupe key travels with it, as an idempotency key the downstream system honours, and the processed-message row is written when the downstream confirms.

Both gaps get drawn on the sequence diagram before any tooling is chosen, because the tooling conversation is where teams get lost. Which broker, which delivery guarantee, which client library: none of it closes either gap. The gaps are closed by transactions and constraints in the database the service already has, and the broker's job is only to carry messages between them reliably enough that the dedupe table is rarely consulted.

What the broker's retention buys

The broker still matters, in one specific way: how long it keeps a message the consumer has not acknowledged, which is the safety margin for gap two. A consumer that crashes and is not restarted for a day needs the message to still be there when it comes back. The defaults differ widely: Kafka's log.retention.hours is 168, a week; Amazon SQS keeps a message for four days unless told otherwise; Google Pub/Sub's subscription default is seven days; and Redis Pub/Sub keeps nothing at all.

Default message retention across common brokers Horizontal bars in hours: Redis Pub/Sub, zero, messages are dropped if no subscriber is connected; Amazon SQS, 96 hours, four days; Kafka, 168 hours, seven days; Google Pub/Sub, 168 hours, seven days. Redis Streams retain until trimmed and are shown as unbounded. How long the safety margin lasts by default Default retention of an unacknowledged message, hours a week no margin Redis Pub/Sub 0: dropped if nobody is listening Amazon SQS 96 Kafka 168 Google Pub/Sub 168 Redis Streams until trimmed
Source: Redis Pub/Sub and Streams documentation, the SQS developer guide, Kafka's log.retention.hours default, and Pub/Sub subscription properties.

The first bar is the one to notice. Redis Pub/Sub delivers to whoever is connected at the instant of publication and keeps nothing, so for a consumer that was down at that instant the safety margin is zero and gap two is not a gap; it is a hole. That is the right behaviour for a dashboard tick and the wrong one for an order event, and choosing between Pub/Sub and Streams is exactly the question of whether a missed message still costs something after the next one arrives. The other brokers give days, which is long enough for a consumer to be repaired, provided the consumer is idempotent when it comes back.

The dedupe key is the hard part

The processed-message table is simple to describe and has one design decision that decides whether it works: what the key is. The obvious choice, the broker's message id, is wrong in the case that matters. If the relay crashes after publishing but before marking the outbox row as sent, it will publish the same event again with a new message id, and a consumer keyed on message id will process both. The key has to be the event's own identity, minted when the outbox row was written and carried through every hop unchanged: the order id plus the event type plus a sequence, or a UUID generated at the point of the original commit. Then two publishes of one event collide on the same key, wherever in the pipeline the duplication happened.

That is the same lesson as idempotency keys in an API: the key identifies the intent, not the delivery attempt. A consumer that keys on the attempt is idempotent against the broker's retries and nothing else.

Four outcomes

Closing neither gap, one, or both produces four systems, and each has a failure with a name.

The four outcomes of closing neither, one or both gaps A two by two grid. Horizontal axis: gap one closed, no on the left and yes on the right. Vertical axis: gap two closed, no at the bottom and yes at the top. Bottom left: lost and duplicated events. Bottom right: events never lost but duplicated on redelivery; the outbox alone. Top left: no duplicates but events can be lost; a careful consumer behind a leaky publisher. Top right: correct. Lost, never duplicated careful consumer, leaky publisher Correct outbox and idempotent consumer Lost and duplicated the starting point Duplicated, never lost the outbox alone: where most teams stop Gap one closed by an outbox: no to yes Gap two closed by a same-transaction dedupe: no to yes
Illustrative: the grey dot is where the first bug's fix leaves a system; the gold dot is where the rule requires it to be.

SocialSure's services are event-driven, and the honest description of how they got to the top-right cell is that they visited the bottom-right one first. The outbox went in after a lost event; the processed-message table went in after the duplicate that the outbox made possible. The second fix was smaller than the first and later, which is the wrong order, and the two-gap rule exists so that the next pipeline gets both on the first day.

What the rule does not say

It does not say exactly-once. Exactly-once delivery between independent systems is not a property a broker can provide, and the vendors that claim it are describing their own internal semantics, not the edge between your database and your email provider. What the rule provides is effectively-once: the event exists once in the log, and the side effect happens once because repeats are recognised and discarded. That is the property a business needs, and it is built from two transactions and two constraints, not from a delivery guarantee.

It also does not say the dedupe table can be forgotten. The processed-message rows have to live at least as long as the broker's retention, because a message can be redelivered for as long as the broker keeps it, and a dedupe table pruned before the broker's window closes reopens gap two for the tail. The retention chart above is the minimum age of the rows, and it is one more reason the broker's default matters.

Draw both gaps. Close both with the database. Then pick the broker, and set its retention to the length of the repair you might one day need.

Event-DrivenPostgreSQLReliability
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS