Treat the model like a game client
An agent's tool calls are intents from a client you do not control. Games solved that years ago: validate every intent on the host, and send each client only what it may see.
Every online card game has one rule that nothing else works without: the client does not get to say what happened. It says what it wants to happen, and the host decides. When I built Turup Chaal, that rule was the whole anti-cheat design, and I wrote about it in the previous post. I have since come to think it is also the whole design for an agent that calls tools.
An agent's tool call looks like an action. It is not one. It is a message from a client you do not control, produced by a process you cannot inspect, asking for something to happen. The system that receives that message is the host, and the host owes it exactly the treatment a card game gives a tapped card.
An intent is not an action
The tau-bench paper measured how reliably function-calling agents finish realistic tasks in a retail domain and an airline domain, with a simulated user and a written policy the agent must follow. State-of-the-art agents succeeded on fewer than half of the tasks. The authors then added a second metric, pass^k, the chance that all k independent runs of the same task succeed. In the retail domain, pass^8 was below 25%. Run the same agent on the same task eight times and fewer than one task in four comes out right every time.
In a game you would describe that as a client that sends a bad move a third of the time. Nobody ships a card game that applies such moves and reconciles later. The OWASP Top 10 for LLM applications names the equivalent failure in agents directly: LLM06:2025 Excessive Agency, damaging actions performed in response to unexpected, ambiguous or manipulated model output, whatever caused the output. OWASP splits the cause three ways: functionality the agent did not need, permissions wider than the task, and autonomy over actions that should have waited for a person. The remedy it gives is the netcode remedy: enforce authorisation outside the model, in the system that executes.
Three checks a host already knows how to write
In Turup Chaal a PlayCardIntent passes three gates before anything changes: is it this seat's turn, does this hand hold the card, does the led suit allow it. Rename the nouns and you have the validation layer for a tool server.
- Is it this actor's turn. The workflow state must allow the action now. A refund exists only for an order in a refundable state; a "cancel shipment" intent on a delivered order is rejected, not interpreted.
- Does the actor hold the card. The principal the agent acts for must own the object. The agent is not the customer; it is the customer's client. If the customer cannot see order 4471, neither can the agent, whatever the model says about it.
- Does the rule allow it. Arguments are checked against live state, never against the model's description of state. The refund amount is compared with what was paid, not with what the model reports was paid.
if (!allowed(order.state, "refund")) return reject("Order is not refundable in this state.");
if (order.customerId !== actor.id) return reject("Actor does not own this order.");
if (amount > order.paidMinusRefunded) return reject("Refund exceeds what was paid.");
What the host returns on rejection matters as much as the rejection. Turup Chaal sends a rule violation event and changes nothing. A tool server should do the same: a structured refusal with the reason, so the model can try a legal move, and no partial state left behind. "Trust the client and reconcile later" was never a design in games. It should not become one here.
Autonomy is a turn order
OWASP's third cause, excessive autonomy, has a game shape too. Some actions in a card game belong to a seat that no client occupies: the host deals, the host scores the trick, the host ends the round. In an agent system, the high-impact actions belong to a seat only a person can play. "Close the account" is not an intent the agent's seat can make; it is a proposal that moves the workflow into a state where the human seat is on turn. The agent has not been denied anything. It has been given a turn order, and the order is data the host enforces, not a policy the model is asked to remember.
The half everyone skips
Validation stops an agent from doing something illegal. It does nothing to stop the agent from reading what it should not, and reading is where the newer failures live. A tool that returns the full customer record hands the model every field, including the ones that carry other people's data and the ones an attacker filled with instructions. A prompt injection sitting in a note field cannot fire if the note is never sent.
The game answer is redaction at the source. After each accepted move, the host builds a snapshot for each player that contains only that player's hand, the table, and the counts, and sends it only to that player. A tool result is a snapshot. Build it for this agent and this task: the fields the next step needs, nothing else. The model cannot leak, act on, or be steered by a value that never reached it.
Why the numbers make this a host problem
If eight runs of a task were independent coin flips with success probability p, pass^8 would be p to the eighth. At p equal to 0.7, that is under 6%. The measured pass^8 in tau-bench sits well above that curve, which sounds like good news and is not. It means failures cluster by task: when an agent gets a task wrong, it tends to get it wrong the same way every time. A client that repeats the same illegal move on every retry is exactly the client a host exists for. Retrying without validation does not average the error away. It re-sends it.
What the host costs
People assume validation and redaction are expensive. In the card game it was four small payloads a move. Here it is one state lookup per mutating call and one projection per result, both cheaper than the model call they wrap. The cost that matters is design cost, and it is paid once: the tool server needs a "state as seen by agent X" view, the workflow needs explicit states with allowed transitions, and every mutating tool needs its three checks written down next to it.
There is a limit to the analogy, and it is worth naming. A card game has a finite rule book, so the host can be complete. A tool server's rules are only as complete as the states and transitions someone bothered to declare, and an action nobody modelled gets no gate. The discipline, then, is that a tool with no declared preconditions does not get wired to a mutating call at all. Read-only tools can be generous. Writes wait for their rules.
One more thing the game taught me, which I did not expect to carry over: the host's checks are also the best test suite. Every rejection the tool server produces is a labelled example of the model trying something the workflow forbids, and a week of those rejections tells you more about where the agent goes wrong than any benchmark. In Turup Chaal the rule violation log was how I found bugs in my own client. In an agent system it is how you find the prompts that need work, the tools that need clearer descriptions, and the states you forgot to declare.
What you get back is a security argument that fits in one sentence, and it is the same sentence as before. The host decides what happened, and each agent only ever learns its own part of it.
Get new posts by email
Occasional essays on engineering, AI, and building for the people technology leaves behind.
Subscribe with RSS