18 September 2026 · 5 min read

You pay for the context every turn

In a retrieval chat the bill is input tokens: every turn resends the prompt, passages and history while the model writes a few hundred back. One ratio shows where the cost is.

The intuition about language model bills is that you pay for what the model writes. Output tokens are the expensive ones, five or six times the price of input at every provider, so the answer must be where the money goes. In a retrieval chat, it is not. Every turn resends the system prompt, the retrieved passages and the entire conversation so far, and the model writes a few hundred tokens back. Over a ten-turn conversation the input tokens billed outrun the output tokens by a factor that starts around twelve and, depending on one design decision, ends anywhere between twenty and a hundred. That factor has a name in my cost reviews, and it is the first number I compute for any workload that talks to a model more than once.

The resend ratio

The resend ratio is input tokens billed divided by output tokens generated, over a whole session. It is one number, it comes straight from the usage fields every API returns, and it answers the only question that matters before optimising: is this workload's cost in generation, which is rare, or in context, which is usual? A ratio near one is a generation workload, where the model writes about as much as it reads, and the levers are shorter answers and cheaper models. A ratio in the tens is a context workload, and the levers are the ones that follow: what goes in the prompt, what stays in the history, and what the cache can reuse. Most teams measure the second kind of workload with the instincts of the first, which is how a chat product ends up choosing a cheaper model to fix a bill that the model's price had almost nothing to do with.

Resend ratio per turn across a ten-turn retrieval chat, under two history policies Two rising lines against turn number one to ten. With passages kept out of the stored history, the per-turn ratio rises from 12 at turn one to 24 at turn ten. With passages kept in the history, it rises from 12 to 96. Both use a 600-token system prompt, four 300-token passages per turn, 40-token questions and 150-token answers. Twelve to one at the first turn, and it only goes up Input per output token, per turn; see caption for inputs passages kept in history passages dropped 100 75 50 25 0 1 4 7 10 Turn number 12 96 at turn ten 24 Cumulative over the session: 54 to 1 with passages kept, 18 to 1 with them dropped.
Illustrative: computed from the stated assumptions, a 600-token system prompt, four 300-token passages retrieved fresh each turn, 40-token questions and 150-token answers; no measured traffic is involved.

The chart's two lines are the same conversation under two history policies, and the difference between them is the single largest cost decision in a retrieval chat. If the passages retrieved for each turn are kept in the stored history, so that turn ten's prompt contains the passages from turns one through nine as well as its own, the per-turn ratio reaches 96 by turn ten and the session averages 54 input tokens per output token. If each turn's passages are used and then dropped, with only the question and the answer retained, turn ten's ratio is 24 and the session averages 18. Same model, same answers, three times the bill.

Why the price tables make it worse than it looks

Output tokens are more expensive, and that fact is what hides the ratio. At the current list prices, the output price is five times the input price across the Claude line, Fable 5.1 at 10 and 50 dollars per million tokens, Opus 5 at 5 and 25, Sonnet 5 at 2 and 10, Haiku 4.5 at 1 and 5, per the published pricing; OpenAI's current table runs from five times for its largest models to six for most of the mid and small ones and eight for one small model. So a workload with a resend ratio of five spends equal money on input and output. A ratio of 18 spends about three and a half times as much on input as on output. A ratio of 54 spends eleven times as much.

Ratio of output to input token price across current models from two providers Horizontal bars showing the output price divided by the input price: Claude Fable 5.1, Opus 5, Sonnet 5 and Haiku 4.5 all at 5; gpt-6-astra and gpt-5.6-sol at 5; gpt-5.6-terra, gpt-5.5, gpt-5.4 and gpt-5.4-mini at 6; gpt-5-mini at 8. A marker at the ratio shows where a workload's resend ratio would make input and output cost the same. Output costs five to eight times input; the ratio decides which dominates Output price divided by input price, list prices per million tokens, September 2026 Claude Fable 5.15 Claude Opus 55 Claude Sonnet 55 Claude Haiku 4.55 gpt-6-astra5 gpt-5.6-terra6 gpt-5.4-mini6 gpt-5-mini8 Anthropic OpenAI Fifty pixels per unit; a resend ratio above the bar means input dominates.
Source: computed from the list prices on the Claude pricing page and the OpenAI pricing page as read on 18 September 2026; ratios only, since the prices themselves change more often than the ratios do.

What prompt caching actually attacks

Prompt caching is the provider's answer to the resend ratio, and it works on exactly one thing: a prefix of the prompt that is byte-for-byte identical to a prefix the provider has seen recently. Both major providers document the same shape with different numbers. OpenAI's prompt caching guide discounts cached input by up to 90 percent, requires a prefix of at least 1,024 tokens on its current models, keeps entries alive for at least thirty minutes after the last use on those models, and states plainly that reuse requires the entire rendered prefix to match. Anthropic's documentation prices cache reads at a tenth of the input price, cache writes at 1.25 times for a five-minute lifetime or twice for an hour, sets minimum cacheable lengths from 512 to 4,096 tokens depending on the model, and fixes the prefix order as tools, then system, then messages, with explicit breakpoints that cache everything up to that point.

The word that matters in both is prefix. A cache hit reuses everything up to the first byte that differs from the last request, and nothing after it. So the layout of the prompt decides whether the cache ever hits, and the natural layout for a retrieval chat, system prompt, then this turn's passages, then the history, then the question, is the worst one: the passages change every turn, they sit right after the system prompt, and everything after them is uncacheable.

Prompt layout with the cache breakpoint: stable prefix first, retrieved passages after it, history last Two stacked prompt layouts. The naive layout: system prompt, then this turn's passages, then history, then question; the cache breakpoint falls after the system prompt and only the system prompt is reused. The recommended layout: tools and system prompt, then the history so far, then this turn's passages, then the question; the breakpoint falls after the history, so the system prompt and the whole history are reused and only the passages and question are billed at full price. Put what changes last NAIVE: PASSAGES RIGHT AFTER THE SYSTEM PROMPT system this turn's passages history, growing every turn question breakpoint: only the system prompt is ever reused RECOMMENDED: STABLE PREFIX, THEN HISTORY, THEN WHAT CHANGED tools, system history so far this turn's passages question breakpoint: system and history reused at a tenth of the price cacheable, identical to the last turn changes every turn, billed in full History grows one exchange a turn, extending the cached prefix instead of breaking it.
Illustrative: the two layouts as the caching rules in both providers' documentation imply; the breakpoint is where the cached prefix ends.

The recommended layout inverts the order: the stable prefix first, then the history, then this turn's passages, then the question. The history grows by one exchange per turn, and because it grows at the end, each turn's history is the previous turn's history plus a suffix, so the previous turn's cache entry is a prefix of this turn's prompt and it hits. The passages, which change every turn, go after the breakpoint and are billed in full, and they are the only thing that is. On the ten-turn conversation above, with passages dropped from history, that turns most of the 18-to-1 into cache reads at a tenth of the price, and the effective ratio, in money rather than tokens, comes down to the low single digits.

There is one more subtlety in the numbers. A cache write costs more than a plain input token at one provider, 1.25 times for the short lifetime, so a prefix that is written and never read again costs more than not caching it. The layout above makes each turn's write the next turn's read, which is the case caching was built for, and a conversation that ends after one turn pays a small premium for nothing. The resend ratio, measured per session, is how you tell those apart.

Where the ratio comes from in practice

The workload I have described is the shape of the helpdesk chat I worked on: a system prompt, a handful of retrieved passages per turn, a history, and short answers with citations. The ratios above are computed from stated assumptions and not from that system's traffic, but the shape is the shape, and it is the shape of most retrieval chats and most agent loops. The AI automation work I do for clients as a retained service has the same structure with a different label on the box: a stable instruction set, a changing context, a growing history, and a bill that is dominated by the parts that are resent.

So the review starts the same way each time. Pull the usage counts, divide input by output over a session, and look at the number. If it is near five, the model's answers are the cost and the conversation is about which model. If it is in the tens, the context is the cost, and the conversation is about three things in order: what is in the history, where the passages sit, and whether the prefix is stable enough to cache. The provider's discount is real, and it applies only to the part of the prompt you were disciplined enough to keep the same.

CostPrompt CachingRAG
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS