11 September 2026 · 5 min read

The part of your schema no grammar checks

A JSON Schema is compiled to a grammar for constrained decoding, and several keywords cannot survive the compilation. Know which, and validate that residue after decoding.

When a structured-output engine accepts your JSON Schema, it is not promising to enforce your JSON Schema. It is promising to enforce the part of it that can be turned into a grammar the decoder can check one token at a time. Several keywords cannot be, either because the constraint is not context-free at all or because the only grammar that expresses it is exponentially large, and every engine quietly drops or approximates those. If you do not know which keywords those are, you will ship outputs the schema forbids while believing they are guaranteed. I call the dropped part the schema residue, and the rule that follows is short: validate the residue after decoding, and only the residue.

How much gets dropped

The clearest measurement is the JSONSchemaBench study by Geng and colleagues, which ran ten thousand real-world schemas through six engines, Guidance, Outlines, llama.cpp, XGrammar, OpenAI and Gemini, and measured three things: whether an engine accepts a schema at all, whether its outputs then validate against the full schema, and how long compilation takes. Declared coverage, the share of schemas accepted without error, ranges from near-total on the easy collections to as low as 6 percent for the hosted engines on the hardest, and empirical coverage, the share of generated outputs that actually validate, is lower still, because an engine can accept a schema and then produce output that violates the part of it the grammar did not encode.

Declared coverage range across schema collections, by engine, from JSONSchemaBench Range bars from the lowest to the highest collection score for each engine, in share of schemas accepted: llama.cpp 0.54 to 0.98, Outlines 0.38 to 0.99, Guidance 0.35 to 0.98, XGrammar 0.12 to 1.00, OpenAI 0.06 to 0.89, Gemini 0.08 to 0.86. Every engine has at least one collection where a large share of schemas is rejected. No engine accepts everything, and the hard collections show it Share of schemas an engine accepts, lowest to highest collection, JSONSchemaBench 0 0.5 1.0 llama.cpp 0.540.98 Outlines 0.380.99 Guidance 0.350.98 XGrammar 0.121.00 OpenAI 0.060.89 Gemini 0.080.86
Source: coverage ranges as reported in JSONSchemaBench (Geng et al., 2025), Table 4, lowest to highest collection per engine; the hosted engines are highlighted because their supported subsets are documented rather than inferred.

The hosted engines are the honest case, because they publish the subset. OpenAI's structured outputs documentation lists what is supported, pattern and a fixed set of formats for strings, minimum, maximum and multipleOf for numbers, minItems and maxItems for arrays, and what is not: allOf, not, dependentRequired, dependentSchemas, if, then and else, with further keywords dropped for fine-tuned models, plus limits of 5,000 properties, ten levels of nesting, a thousand enum values and 120,000 characters of names. A schema outside the subset is rejected with an error, which is the right behaviour. The engines that accept a schema and silently ignore part of it are the ones the rule is for.

Why a keyword falls out

The keywords that drop are not arbitrary; each falls out for one of three reasons, and knowing the reason tells you what the post-decode check has to do.

A schema splitting into grammar-enforceable keywords and residue keywords, with the residue routed to a post-validator A schema box on the left feeds two boxes. The top box, grammar-enforceable: types, required, enum, const, additionalProperties, nesting, regular patterns, bounded lengths, and this box feeds the decoder. The bottom box, residue: uniqueItems, numeric bounds on decimals, patterns with backreferences, dependentSchemas and dependentRequired, format semantics, contains with counts, if-then-else across properties, allOf and not, multipleOf, and this box feeds a post-validator that runs on the decoded output. One schema, two destinations Schema as written Grammar-enforceable types, required, enum, const, nesting, regular patterns, bounded lengths Decoder enforced per token Residue uniqueItems, decimal bounds, backreferences, dependent keywords, if-then-else, format, contains counts, allOf, not, multipleOf Post-validator on decoded output the residue is validated with the full schema after decoding; the rest is guaranteed
Illustrative: the split as the checklist applies it; which keywords land in the residue depends on the engine, and the list shown is the union across common ones.

The first reason is that the constraint is not context-free. uniqueItems asks whether any two items in an array are equal, which is the copy language with extra steps, and no grammar recognises it. A pattern with a backreference is a regular expression in name only; backreferences make the language non-regular and in general non-context-free. allOf is an intersection, and the intersection of two context-free languages need not be context-free; not is a complement, with the same problem.

The second reason is that the constraint is context-free but only with an exponential grammar. A condition across properties, if this property has that value then that other property is required, is expressible in a context-free grammar only by enumerating the orderings in which the properties may appear, and JSON objects allow any order, so the grammar grows with the factorial of the property count. dependentSchemas and dependentRequired have the same shape. contains with minContains and maxContains is a counting constraint across an array; a grammar can count to a small fixed bound by unrolling, and the unrolling is the exponential.

The third reason is that the constraint is about meaning, not syntax. format says a string is an email address or a date, and a grammar can enforce the shape of a date but not that the thirtieth of February does not exist. A numeric minimum or maximum is a constraint on a number, and a number in JSON is a string of digits: a grammar can bound an integer by enumerating digit patterns, which is tedious but finite, and cannot bound a decimal with arbitrary precision without an automaton that grows with the precision. multipleOf is arithmetic and has no grammar at all.

How common the residue is

The residue would not matter if schemas rarely used those keywords. They use them constantly. I counted keyword usage across the JSON Schema Store catalogue, 984 public schemas for configuration files and manifests, and the residue keywords are in a large minority of them.

Share of Schema Store schemas using each keyword, with residue keywords highlighted Horizontal bars in percent of 984 schemas: pattern 40.1, format 32.4, minimum 31.3, allOf 27.8, maximum 22.1, uniqueItems 21.7, not 14.9, dependencies 12.0, if 11.8, propertyNames 4.7, multipleOf 1.6, contains 1.3. Residue keywords are gold; pattern is grey because most patterns are plain regular expressions and only those with backreferences drop, of which the catalogue has none. The residue is in a third of real schemas Percent of 984 public schemas using the keyword at least once, September 2026 pattern 40.1 format 32.4 minimum 31.3 allOf 27.8 maximum 22.1 uniqueItems 21.7 not 14.9 dependencies 12.0 if, then, else 11.8 propertyNames 4.7 multipleOf 1.6 contains 1.3 residue on at least one common engine regular in practice: no backreferences found
Source: computed by the author over the SchemaStore repository's schema catalogue, 984 parsed schemas, counting each keyword once per schema.

A third of these schemas use format, a third use a numeric minimum, over a fifth use uniqueItems, and one in eight uses a conditional across properties. None of those is exotic. A CRM-style schema, the kind our APIs emit and accept, has an email format on a contact, a minimum of zero on a deal amount, uniqueItems on a list of tags, and a conditional that requires a close date when the stage is won. Every one of those is residue on at least one common engine, and an output that violates any of them will pass the grammar. The pattern keyword is the exception that proves the shape of the rule: it is the most used keyword in the catalogue and it is enforceable, because a plain regular expression is a regular language and a grammar contains it, so it is grey in the chart. It would drop only for patterns with backreferences, and the catalogue has none, which says something about how rarely real schemas need them.

The residue checklist

The practice is a checklist run against the schema before it is handed to an engine. For each keyword in the residue list, uniqueItems, numeric bounds on non-integers, patterns with backreferences, dependentSchemas and dependentRequired, format, contains with counts, conditionals across properties, allOf, not, multipleOf, mark whether the schema uses it and whether the target engine documents support for it. What remains marked is the residue for that schema on that engine. The decoded output is then validated against the full schema with an ordinary validator implementing the validation vocabulary, and the validator's failures are, by construction, residue failures, because everything else was guaranteed by the grammar.

The "only the residue" half of the rule is about cost and about honesty. Validating the whole schema after decoding costs nothing extra, since the validator does it in one pass, so run the whole thing. What the checklist adds is that you know in advance which failures are possible, so that a failure on a keyword outside the residue is a bug in the engine rather than an expected rejection, and a failure inside it is the system working as designed. Reporting the residue rejection rate next to the schema validity rate is the same discipline as reporting the checker's failures for copy constraints: the grammar's guarantee is real, and the number that matters is what it does not cover.

There is a limit to the checklist that is worth stating. Engines change, and a keyword that is residue this quarter may be enforced next quarter, or the reverse, when an engine trades coverage for speed. The checklist is therefore per engine and per version, and the safest posture is the one the hosted engines take: reject what you cannot enforce, so that nobody is left believing a guarantee that was never made.

Structured OutputJSON SchemaValidation
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS