Which language your AI fails in
Building AI in India means shipping a product that works in English and fails quietly in the languages most users think in. Find and publish the failure languages before users do.
An AI product built in India is, by default, a product that works in English and fails quietly in the languages most of its users think in. Quietly is the important word. The product does not crash in Hindi; it answers a little worse, retrieves a little less, refuses a little more often, and costs several times as much per query, and none of that shows up in an evaluation run in English. The gap is structural, it comes from what the models were trained on, and the obligation that comes with building here is not to close it, which no single product can, but to find it and publish it before users find it themselves.
The gap is in the data
The web that language models learn from is overwhelmingly not in Indian languages. W3Techs' survey of website content languages, as of 17 September 2026, puts English at 49.5 percent of websites and lists Hindi, Bengali, Tamil, Urdu, Marathi and Telugu among the languages used by fewer than 0.1 percent each. Common Crawl, the corpus most open models are built from, publishes language statistics per crawl: in CC-MAIN-2026-34, the August 2026 crawl, English is 40.45 percent of pages, Hindi 0.21 percent, Bengali 0.11, Tamil 0.04, Urdu 0.03, Marathi 0.03 and Telugu 0.02. Add the six Indian languages together and they are under half a percent of the crawl for languages spoken by hundreds of millions of people.
The models built in India are the response to that gap, and the India AI Impact Summit in New Delhi in February 2026 was where the response was shown: BharatGen's Param2, a 17-billion-parameter model built for 22 Indian languages, and Sarvam's 30 and 105-billion-parameter models. Those matter, and they do not remove the product obligation, because a product ships on whichever model it ships on, and the question for its users is not how good the best Indic model is but how the product they are holding behaves in their language.
Where the failure hides
The failure has four modes, and a product that measures only accuracy sees at most one of them. Refusal: the model declines in a language where it would have answered in English, because its safety training saw fewer examples and errs toward silence. Hallucination: the model answers fluently and wrongly, because fluency in the script was learned and the facts were not. Degraded retrieval: in a retrieval pipeline, the embedding model ranks passages worse in the query's language, so the generator is grounded on the wrong text and the citation is to the wrong place. And the token budget: the tokeniser, built on the frequencies of the web, spends several times as many tokens on the same sentence in an Indian script, which raises the cost per query and shortens the effective context by the same factor.
The token figure is the one I can compute exactly, and it is the one most teams never look at. With the older of the two public tokenisers, the Tamil sentence costs ten times the English one; with the newer, about two and a third times. A product priced per query in English is a product that loses money, or shortens its context, or both, for every Tamil user, and the team will discover this from the bill rather than from a benchmark.
The failure-language test
The obligation, then, is a test run before release, and it is not complicated. Take a fixed task set: the twenty or fifty things the product is actually for, with reference answers. Run it in every scheduled language a user could reasonably use with the product, which for most Indian products is a dozen rather than twenty-two. For each language, record the four modes: the refusal rate, an error rate judged by someone who reads the language, the retrieval quality if the product retrieves, and the tokens per task relative to English. Then publish the list of languages in which the product refuses, hallucinates, retrieves badly or blows the token budget, and say what "badly" meant.
A product that has not run the test cannot claim to serve India, and a product that has run it and published the failing rows can, honestly, because it has told its users where it works. That is the whole obligation. The test finds failure; it does not fix it, and pretending it does would be a worse dishonesty than the one it replaces. Fixing a failure language is a model choice, a retrieval choice, a tokeniser choice or a data choice, and each is a project. Publishing the list is an afternoon, once the task set exists, and it is the afternoon most products skip.
Running it without fooling yourself
Three details decide whether the test measures the product or measures something else. The first is where the task set comes from. It should be drawn from what users actually ask, from logs where they exist and from the product's own onboarding where they do not, so that the languages are tested on the product's job rather than on a translated benchmark that nobody uses it for. The second is translation. If the tasks are machine-translated into each language, the test measures the translator as much as the product, and a translation error becomes a product failure in the results; the tasks and the reference answers have to be written or checked by someone who speaks the language, which is a cost, and it is the cost that makes the test mean something. The third is the threshold: "fails" is defined relative to English on the same tasks, not against an absolute bar, because the obligation is to know where the product is worse than it claims, and the claim is calibrated in English.
Retrieval deserves a separate word, because it fails in a way that the other modes hide. A pipeline that indexes documents in English and receives a query in Tamil depends on a cross-lingual embedding to match them, and the match quality is a property of the embedding model in that language pair, not of the generator. A product can have a generator that handles Tamil adequately and still fail the retrieval column, because the passages it was handed were the wrong ones; the answer then looks like hallucination and is not. Measuring retrieval separately, by whether the right passage is in the top results for the translated query, is what tells the two apart, and it is what points the fix at the right component.
Why I hold to this
The pipeline I worked on during my internship was a retrieval system with grounded citations, served on fine-tuned open models against latency targets, and every stage of it was tuned in English because that is the language the documents, the evaluators and the latency budget were in. Nothing about it was wrong for English. What I could see, without measuring, was that retrieval quality and generation quality would move differently for a user who typed in Hindi, and that the latency target, which was set in tokens per second, would be met in English and missed for that user for reasons that had nothing to do with the hardware. I did not have the failure-language test then. I have it now, and it is the first thing I would run on that pipeline.
The reason it belongs to founders rather than to researchers is that researchers have already done their part: the gap is documented, the Indic models exist, the tokeniser cost is measurable with a public library in ten minutes. What remains is the decision, per product, to look, and to say what was found. Building AI in India comes with the languages of India attached. The test is how you find out which of them you have actually built for.
Get new posts by email
Occasional essays on engineering, AI, and building for the people technology leaves behind.
Subscribe with RSS