Blog · FEB 19, 2025 · 3 min read
Hindi Plus Eight: Choosing Our First Language Set
Nine languages isn't a marketing number, it's the smallest set that lets FieldVoice cover reps who don't think in English.
Deciding which languages FieldVoice should understand at launch sounds like a straightforward data problem, rank India's languages by speaker count, pick a cutoff, ship it. It is not quite that simple, because the relevant population is not "people who speak this language" but "medical representatives who think in this language while describing a clinical conversation," which is a narrower and more specific group than any census table captures directly.
How we actually chose
We started from field-force distribution data across the pharma companies we had spoken with, layered against language-comfort patterns from the same ride-alongs that shaped the rest of the product. Hindi was never in question, it covers the largest single share of field reps by a wide margin, across the belt from Punjab through Madhya Pradesh into large parts of the east. Beyond Hindi, the choice got harder, because speaker count alone would have pointed us toward languages already reasonably well served by English-medium business tools, and away from languages spoken by reps who had the least alternative.
- •Hindi, the largest single-language field force by far
- •Marathi, Gujarati, dense pharma territory concentration in Maharashtra and Gujarat
- •Tamil, Telugu, Kannada, Malayalam, the four major south Indian languages, each with distinct pharma hubs
- •Bengali, eastern coverage, another large single-state concentration
- •Punjabi, smaller in absolute numbers, but a territory where English and Hindi comfort both measurably drop
Nine languages, including English as the default, was the smallest set that let us say, honestly, that a rep in most of the country's major pharma territories could use FieldVoice in the language they actually think in, rather than the language their company's CRM happened to be built in.
We were deliberate about not treating this as a translation problem solved once and shipped. Regional pharma vocabulary, brand names, dosage terms, colloquial ways of describing an objection, varies enough across these languages that a model trained mainly on general conversational data misses a meaningful share of what a rep actually says in a visit. Getting that right per language took longer than adding the language itself, and it is the reason the list grew slowly rather than all at once.
Nine is not a ceiling. It is the honest floor for what we could ship well rather than ship broadly, and the next languages we add will earn their place the same way these did, by testing against real recordings from real territories, not by speaker count on a spreadsheet.