There are two gaps, not one
Every write-then-publish has two places to crash. The outbox pattern closes only the first, so a team that stops there has traded lost events for duplicated ones.
The bug that teaches most teams about event-driven systems is the lost event: the database commit succeeded, the publish to the broker failed, and a downstream service never learned that an order existed. The fix everybody reaches for is the outbox pattern, and it works. Then a second bug arrives, quieter than the first, and it turns out the fix had moved the problem rather than removed it. There are two gaps in every write-then-publish pipeline, not one, and the outbox closes exactly one of them.
Gap one: between the commit and the publish
The first gap is the one the outbox is for. A service writes a row and then publishes an event, and between those two actions the process can die, the network can fail, or the broker can refuse. If the write committed and the publish did not, the event is lost. If the publish happened and the write then rolled back, the event describes something that never occurred. The two operations are on different systems and cannot share a transaction.
The outbox closes the gap by putting the event in the same transaction as the row. The service writes the order and, in the same commit, writes an outbox row describing the event. A relay reads the outbox and publishes to the broker, marking rows as sent. Now the event cannot be lost, because it was committed with the data, and it cannot describe a rollback, because it rolled back too. The relay can crash and resume, and it will publish anything it had not marked.
Gap two: between the side effect and the acknowledgement
Which brings the second gap. The relay published the event, the consumer received it, the consumer sent the confirmation email, and then the consumer crashed before acknowledging the message. The broker, not having seen the acknowledgement, delivers the message again. The consumer sends the email again. The outbox made the event exactly-once in the sense that it exists once in the log. It did nothing for the consumer, whose side effect and whose acknowledgement are, again, on different systems and cannot share a transaction.
A team that stops at the outbox has made a trade it did not intend: lost events for duplicated ones. Duplicates are usually less damaging than losses, which is why the trade feels like progress, but a duplicated charge or a duplicated dispatch is still a bug, and it is a bug that appears only under failure, which is the hardest kind to reproduce.
The two-gap rule
So the rule I apply to every pipeline is short. It is correct only when gap one is closed by a same-transaction record and gap two is closed by a same-transaction dedupe. The first is the outbox row. The second is a processed-message table on the consumer's side: before performing the side effect, the consumer inserts the message id into a table with a unique constraint, in the same transaction as whatever database change the side effect involves. A redelivered message hits the constraint and is acknowledged without being processed. Where the side effect is not a database change, an email, a call to a payment provider, the dedupe key travels with it, as an idempotency key the downstream system honours, and the processed-message row is written when the downstream confirms.
Both gaps get drawn on the sequence diagram before any tooling is chosen, because the tooling conversation is where teams get lost. Which broker, which delivery guarantee, which client library: none of it closes either gap. The gaps are closed by transactions and constraints in the database the service already has, and the broker's job is only to carry messages between them reliably enough that the dedupe table is rarely consulted.
What the broker's retention buys
The broker still matters, in one specific way: how long it keeps a message the consumer has not acknowledged, which is the safety margin for gap two. A consumer that crashes and is not restarted for a day needs the message to still be there when it comes back. The defaults differ widely: Kafka's log.retention.hours is 168, a week; Amazon SQS keeps a message for four days unless told otherwise; Google Pub/Sub's subscription default is seven days; and Redis Pub/Sub keeps nothing at all.
The first bar is the one to notice. Redis Pub/Sub delivers to whoever is connected at the instant of publication and keeps nothing, so for a consumer that was down at that instant the safety margin is zero and gap two is not a gap; it is a hole. That is the right behaviour for a dashboard tick and the wrong one for an order event, and choosing between Pub/Sub and Streams is exactly the question of whether a missed message still costs something after the next one arrives. The other brokers give days, which is long enough for a consumer to be repaired, provided the consumer is idempotent when it comes back.
The dedupe key is the hard part
The processed-message table is simple to describe and has one design decision that decides whether it works: what the key is. The obvious choice, the broker's message id, is wrong in the case that matters. If the relay crashes after publishing but before marking the outbox row as sent, it will publish the same event again with a new message id, and a consumer keyed on message id will process both. The key has to be the event's own identity, minted when the outbox row was written and carried through every hop unchanged: the order id plus the event type plus a sequence, or a UUID generated at the point of the original commit. Then two publishes of one event collide on the same key, wherever in the pipeline the duplication happened.
That is the same lesson as idempotency keys in an API: the key identifies the intent, not the delivery attempt. A consumer that keys on the attempt is idempotent against the broker's retries and nothing else.
Four outcomes
Closing neither gap, one, or both produces four systems, and each has a failure with a name.
SocialSure's services are event-driven, and the honest description of how they got to the top-right cell is that they visited the bottom-right one first. The outbox went in after a lost event; the processed-message table went in after the duplicate that the outbox made possible. The second fix was smaller than the first and later, which is the wrong order, and the two-gap rule exists so that the next pipeline gets both on the first day.
What the rule does not say
It does not say exactly-once. Exactly-once delivery between independent systems is not a property a broker can provide, and the vendors that claim it are describing their own internal semantics, not the edge between your database and your email provider. What the rule provides is effectively-once: the event exists once in the log, and the side effect happens once because repeats are recognised and discarded. That is the property a business needs, and it is built from two transactions and two constraints, not from a delivery guarantee.
It also does not say the dedupe table can be forgotten. The processed-message rows have to live at least as long as the broker's retention, because a message can be redelivered for as long as the broker keeps it, and a dedupe table pruned before the broker's window closes reopens gap two for the tail. The retention chart above is the minimum age of the rows, and it is one more reason the broker's default matters.
Draw both gaps. Close both with the database. Then pick the broker, and set its retention to the length of the repair you might one day need.
Get new posts by email
Occasional essays on engineering, AI, and building for the people technology leaves behind.
Subscribe with RSS