You pay for the context every turn
In a retrieval chat the bill is input tokens: every turn resends the prompt, passages and history while the model writes a few hundred back. One ratio shows where the cost is.
The intuition about language model bills is that you pay for what the model writes. Output tokens are the expensive ones, five or six times the price of input at every provider, so the answer must be where the money goes. In a retrieval chat, it is not. Every turn resends the system prompt, the retrieved passages and the entire conversation so far, and the model writes a few hundred tokens back. Over a ten-turn conversation the input tokens billed outrun the output tokens by a factor that starts around twelve and, depending on one design decision, ends anywhere between twenty and a hundred. That factor has a name in my cost reviews, and it is the first number I compute for any workload that talks to a model more than once.
The resend ratio
The resend ratio is input tokens billed divided by output tokens generated, over a whole session. It is one number, it comes straight from the usage fields every API returns, and it answers the only question that matters before optimising: is this workload's cost in generation, which is rare, or in context, which is usual? A ratio near one is a generation workload, where the model writes about as much as it reads, and the levers are shorter answers and cheaper models. A ratio in the tens is a context workload, and the levers are the ones that follow: what goes in the prompt, what stays in the history, and what the cache can reuse. Most teams measure the second kind of workload with the instincts of the first, which is how a chat product ends up choosing a cheaper model to fix a bill that the model's price had almost nothing to do with.
The chart's two lines are the same conversation under two history policies, and the difference between them is the single largest cost decision in a retrieval chat. If the passages retrieved for each turn are kept in the stored history, so that turn ten's prompt contains the passages from turns one through nine as well as its own, the per-turn ratio reaches 96 by turn ten and the session averages 54 input tokens per output token. If each turn's passages are used and then dropped, with only the question and the answer retained, turn ten's ratio is 24 and the session averages 18. Same model, same answers, three times the bill.
Why the price tables make it worse than it looks
Output tokens are more expensive, and that fact is what hides the ratio. At the current list prices, the output price is five times the input price across the Claude line, Fable 5.1 at 10 and 50 dollars per million tokens, Opus 5 at 5 and 25, Sonnet 5 at 2 and 10, Haiku 4.5 at 1 and 5, per the published pricing; OpenAI's current table runs from five times for its largest models to six for most of the mid and small ones and eight for one small model. So a workload with a resend ratio of five spends equal money on input and output. A ratio of 18 spends about three and a half times as much on input as on output. A ratio of 54 spends eleven times as much.
What prompt caching actually attacks
Prompt caching is the provider's answer to the resend ratio, and it works on exactly one thing: a prefix of the prompt that is byte-for-byte identical to a prefix the provider has seen recently. Both major providers document the same shape with different numbers. OpenAI's prompt caching guide discounts cached input by up to 90 percent, requires a prefix of at least 1,024 tokens on its current models, keeps entries alive for at least thirty minutes after the last use on those models, and states plainly that reuse requires the entire rendered prefix to match. Anthropic's documentation prices cache reads at a tenth of the input price, cache writes at 1.25 times for a five-minute lifetime or twice for an hour, sets minimum cacheable lengths from 512 to 4,096 tokens depending on the model, and fixes the prefix order as tools, then system, then messages, with explicit breakpoints that cache everything up to that point.
The word that matters in both is prefix. A cache hit reuses everything up to the first byte that differs from the last request, and nothing after it. So the layout of the prompt decides whether the cache ever hits, and the natural layout for a retrieval chat, system prompt, then this turn's passages, then the history, then the question, is the worst one: the passages change every turn, they sit right after the system prompt, and everything after them is uncacheable.
The recommended layout inverts the order: the stable prefix first, then the history, then this turn's passages, then the question. The history grows by one exchange per turn, and because it grows at the end, each turn's history is the previous turn's history plus a suffix, so the previous turn's cache entry is a prefix of this turn's prompt and it hits. The passages, which change every turn, go after the breakpoint and are billed in full, and they are the only thing that is. On the ten-turn conversation above, with passages dropped from history, that turns most of the 18-to-1 into cache reads at a tenth of the price, and the effective ratio, in money rather than tokens, comes down to the low single digits.
There is one more subtlety in the numbers. A cache write costs more than a plain input token at one provider, 1.25 times for the short lifetime, so a prefix that is written and never read again costs more than not caching it. The layout above makes each turn's write the next turn's read, which is the case caching was built for, and a conversation that ends after one turn pays a small premium for nothing. The resend ratio, measured per session, is how you tell those apart.
Where the ratio comes from in practice
The workload I have described is the shape of the helpdesk chat I worked on: a system prompt, a handful of retrieved passages per turn, a history, and short answers with citations. The ratios above are computed from stated assumptions and not from that system's traffic, but the shape is the shape, and it is the shape of most retrieval chats and most agent loops. The AI automation work I do for clients as a retained service has the same structure with a different label on the box: a stable instruction set, a changing context, a growing history, and a bill that is dominated by the parts that are resent.
So the review starts the same way each time. Pull the usage counts, divide input by output over a session, and look at the number. If it is near five, the model's answers are the cost and the conversation is about which model. If it is in the tens, the context is the cost, and the conversation is about three things in order: what is in the history, where the passages sit, and whether the prefix is stable enough to cache. The provider's discount is real, and it applies only to the part of the prompt you were disciplined enough to keep the same.
Get new posts by email
Occasional essays on engineering, AI, and building for the people technology leaves behind.
Subscribe with RSS