19 September 2026 · 6 min read

Which language your AI fails in

Building AI in India means shipping a product that works in English and fails quietly in the languages most users think in. Find and publish the failure languages before users do.

An AI product built in India is, by default, a product that works in English and fails quietly in the languages most of its users think in. Quietly is the important word. The product does not crash in Hindi; it answers a little worse, retrieves a little less, refuses a little more often, and costs several times as much per query, and none of that shows up in an evaluation run in English. The gap is structural, it comes from what the models were trained on, and the obligation that comes with building here is not to close it, which no single product can, but to find it and publish it before users find it themselves.

The gap is in the data

The web that language models learn from is overwhelmingly not in Indian languages. W3Techs' survey of website content languages, as of 17 September 2026, puts English at 49.5 percent of websites and lists Hindi, Bengali, Tamil, Urdu, Marathi and Telugu among the languages used by fewer than 0.1 percent each. Common Crawl, the corpus most open models are built from, publishes language statistics per crawl: in CC-MAIN-2026-34, the August 2026 crawl, English is 40.45 percent of pages, Hindi 0.21 percent, Bengali 0.11, Tamil 0.04, Urdu 0.03, Marathi 0.03 and Telugu 0.02. Add the six Indian languages together and they are under half a percent of the crawl for languages spoken by hundreds of millions of people.

Share of web pages by primary language in the August 2026 Common Crawl Horizontal bars on a logarithmic scale, percent of pages: English 40.45; Hindi 0.21; Bengali 0.11; Tamil 0.04; Urdu 0.03; Marathi 0.03; Telugu 0.02. The English bar is grey and the Indian-language bars are gold; the gap between them is more than two orders of magnitude. Two orders of magnitude between English and everything Indian Percent of pages by primary language, CC-MAIN-2026-34, log scale English 40.45 Hindi 0.21 Bengali 0.11 Tamil 0.04 Urdu 0.03 Marathi 0.03 Telugu 0.02 Log scale from 0.02 to 50 percent; W3Techs' website survey gives the same ordering.
Source: Common Crawl language statistics for CC-MAIN-2026-34; cross-checked against the W3Techs content-language survey of 17 September 2026, which reports English at 49.5 percent of websites and each Indian language below 0.1.

The models built in India are the response to that gap, and the India AI Impact Summit in New Delhi in February 2026 was where the response was shown: BharatGen's Param2, a 17-billion-parameter model built for 22 Indian languages, and Sarvam's 30 and 105-billion-parameter models. Those matter, and they do not remove the product obligation, because a product ships on whichever model it ships on, and the question for its users is not how good the best Indic model is but how the product they are holding behaves in their language.

Where the failure hides

The failure has four modes, and a product that measures only accuracy sees at most one of them. Refusal: the model declines in a language where it would have answered in English, because its safety training saw fewer examples and errs toward silence. Hallucination: the model answers fluently and wrongly, because fluency in the script was learned and the facts were not. Degraded retrieval: in a retrieval pipeline, the embedding model ranks passages worse in the query's language, so the generator is grounded on the wrong text and the citation is to the wrong place. And the token budget: the tokeniser, built on the frequencies of the web, spends several times as many tokens on the same sentence in an Indian script, which raises the cost per query and shortens the effective context by the same factor.

Tokens for the same sentence, the first article of the Universal Declaration of Human Rights, in four languages and two public tokenisers Grouped horizontal bars. cl100k tokeniser: English 33 tokens, Hindi 180, Bengali 213, Tamil 330. o200k tokeniser: English 33, Hindi 54, Bengali 56, Tamil 76. The newer tokeniser cuts the multiple from ten times down to about twice, and the gap is still there. The same sentence costs two to ten times as much Tokens for Article 1 of the Universal Declaration of Human Rights cl100k o200k English 33 33 Hindi 180 54 Bengali 213 56 Tamil 330 76 Same passage; the difference is what the tokeniser was trained to expect.
Source: computed by the author with tiktoken over the cl100k and o200k encodings, using the published translations of Article 1 of the Universal Declaration of Human Rights.

The token figure is the one I can compute exactly, and it is the one most teams never look at. With the older of the two public tokenisers, the Tamil sentence costs ten times the English one; with the newer, about two and a third times. A product priced per query in English is a product that loses money, or shortens its context, or both, for every Tamil user, and the team will discover this from the bill rather than from a benchmark.

The failure-language test

The obligation, then, is a test run before release, and it is not complicated. Take a fixed task set: the twenty or fifty things the product is actually for, with reference answers. Run it in every scheduled language a user could reasonably use with the product, which for most Indian products is a dozen rather than twenty-two. For each language, record the four modes: the refusal rate, an error rate judged by someone who reads the language, the retrieval quality if the product retrieves, and the tokens per task relative to English. Then publish the list of languages in which the product refuses, hallucinates, retrieves badly or blows the token budget, and say what "badly" meant.

The failure-language matrix: languages against failure modes, with the publish rule A grid with languages down the side, English, Hindi, Bengali, Tamil, Telugu, Marathi, and four failure modes across the top: refusal, hallucination, degraded retrieval, token budget. Cells are marked pass or fail with example results: English passes all; Hindi fails token budget; Tamil fails retrieval and token budget; Telugu fails refusal, retrieval and token budget. A rule beneath says any row with a fail is published as a failure language with the mode named. Run the same tasks in every language, publish every failing row REFUSAL HALLUCINATION RETRIEVAL TOKEN BUDGET English Hindi Tamil Telugu within the English tolerance fails: this row is published, with the mode named Cells are an example pattern, not a measurement; the shape is what the test produces.
Illustrative: the matrix the test fills in, with an example pattern of results; the languages shown are a subset of the scheduled languages a product would run.

A product that has not run the test cannot claim to serve India, and a product that has run it and published the failing rows can, honestly, because it has told its users where it works. That is the whole obligation. The test finds failure; it does not fix it, and pretending it does would be a worse dishonesty than the one it replaces. Fixing a failure language is a model choice, a retrieval choice, a tokeniser choice or a data choice, and each is a project. Publishing the list is an afternoon, once the task set exists, and it is the afternoon most products skip.

Running it without fooling yourself

Three details decide whether the test measures the product or measures something else. The first is where the task set comes from. It should be drawn from what users actually ask, from logs where they exist and from the product's own onboarding where they do not, so that the languages are tested on the product's job rather than on a translated benchmark that nobody uses it for. The second is translation. If the tasks are machine-translated into each language, the test measures the translator as much as the product, and a translation error becomes a product failure in the results; the tasks and the reference answers have to be written or checked by someone who speaks the language, which is a cost, and it is the cost that makes the test mean something. The third is the threshold: "fails" is defined relative to English on the same tasks, not against an absolute bar, because the obligation is to know where the product is worse than it claims, and the claim is calibrated in English.

Retrieval deserves a separate word, because it fails in a way that the other modes hide. A pipeline that indexes documents in English and receives a query in Tamil depends on a cross-lingual embedding to match them, and the match quality is a property of the embedding model in that language pair, not of the generator. A product can have a generator that handles Tamil adequately and still fail the retrieval column, because the passages it was handed were the wrong ones; the answer then looks like hallucination and is not. Measuring retrieval separately, by whether the right passage is in the top results for the translated query, is what tells the two apart, and it is what points the fix at the right component.

Why I hold to this

The pipeline I worked on during my internship was a retrieval system with grounded citations, served on fine-tuned open models against latency targets, and every stage of it was tuned in English because that is the language the documents, the evaluators and the latency budget were in. Nothing about it was wrong for English. What I could see, without measuring, was that retrieval quality and generation quality would move differently for a user who typed in Hindi, and that the latency target, which was set in tokens per second, would be met in English and missed for that user for reasons that had nothing to do with the hardware. I did not have the failure-language test then. I have it now, and it is the first thing I would run on that pipeline.

The reason it belongs to founders rather than to researchers is that researchers have already done their part: the gap is documented, the Indic models exist, the tokeniser cost is measurable with a public library in ten minutes. What remains is the decision, per product, to look, and to say what was found. Building AI in India comes with the languages of India attached. The test is how you find out which of them you have actually built for.

AI PolicyIndiaEvaluation
All writing

Written by Mohd Shayan

Get new posts by email

Occasional essays on engineering, AI, and building for the people technology leaves behind.

One email per new post. Unsubscribe any time.

Subscribe with RSS