[Readmelater]

The Biases And Blinkers Of Big Tech’s Language Models

We spoke to researchers and practitioners across Bangladesh, India, and Bhutan to understand the promises and fallacies of linguistic diversity in AI, and the way ahead

Ask a chatbot a question in English and it will stun you with its eloquence. But ask the same question in Urdu or Bangla or any regional language, and it stumbles. Responses are garbled, reasoning is muddled, and in some cases, it produces “complete nonsense”, according to one investigation that studied ChatGPT’s responses.

ChatGPT is one of the several large language models (LLM) that mediates our relationships with the world, including ourselves, today. From welfare scheme portals and agri-advisory tools to mental health support and companionship, tech and AI infrastructure run on these language models. But a less prominent truth is language models are trained on vast swathes of data that is primarily in English.

Some languages are more present online than others. Bangla, for instance, is spoken by 242 million people globally, but there is not enough Bengali content available online, so there is no data to scrape from. It is thus considered a ‘low-resource’ language in tech circles. Across South Asia, there are more than 650 languages; some have limited computational resources dedicated to them, most have none at all.

Who decides which languages get digitised, whose knowledge becomes legible to a machine, and whose doesn’t is a technical question as much as it is a political one. Language models today mirror old imperial hierarchies, says Nova Ahmed, a natural language processing (NLP) researcher in Bangladesh.

The cost of linguistic invisibility is unevenly distributed among Dalit, Muslims, gender marginalised communities globally, says Aindriya Barua, who created ShhorAI, an AI-powered tool to combat hate speech in vernacular languages in India. “Any model will reproduce the same biases that exist in the real world…If one caste is overrepresented in the data, the model will learn that. It always comes down to where the data is collected, and from whom.”

In NLP, there’s a useful distinction between large language models — built for broad, general-purpose tasks, like Perplexity and Gemini — and task-specific models, trained on narrower, often community-specific data. The former is built for scale and efficiency, while the latter prizes cultural nuance and accuracy. There are efforts to build smaller, task-specific, local language models to address a community’s needs, but these come with their own risk. “I worry about this the most,” says Leki Choden, a tech practitioner from Bhutan. “A language that exists across many valleys, each with its own accent, vocabulary, and rhythm, gets flattened into one ‘official’ digital version.”

To understand what’s at stake in the politics of language, Saumya Kalia spoke with Nova, Aindriya, and Leki to understand the fallacies and promises of linguistic diversity in language models. Nova is cautiously hopeful about emerging initiatives, while Leki highlights the importance of centering communities. Aindriya argues that scale itself is the wrong goal, and the focus needs to shift toward community-led data.  Edited excerpts below.

How important is inclusive, good data to a language model?

Nova: The language model is learning from what it sees, just like a baby starts to learn from experiences. The more data it gets, the more it will learn, and anything missing will not have any impact. If some kind of data is missing – the model may make a wrong assumption or result about it. For example, in a recent study, it was found that women are not getting good credit scores, although they were rarely defaulters. It was because there was not enough data about women. So when it comes to decision making, the data, the data quality, how representative it is — everything matters.

How representative is the data today? What are LLMs like ChatGPT and Gemini predominantly trained on? 

Nova: The languages that dominate large language models today are, by and large, the same languages that dominated through colonisation. It mirrors old imperial hierarchies. In terms of representation, South Asian languages have a large number of speakers but the scale hasn’t translated into good-quality training datasets. What’s missing isn’t speakers but a coordinated effort — through organisations, funding, institutional will – to correct these imbalances.

Aindriya: When I was developing ShhorAI in 2018-19, I did not have resources to draw on because the existing datasets were really bad. All the NLP work that exists is in English or other ‘bigger’ languages spoken in the Global North. One example of what was missing is ‘code-mixed languages’ like Hinglish. I became interested in these in college when I was studying NLP, and at that time, Jio internet had entered the market and everyone had started coming online. There was an influx of people speaking various languages while using the English keyboard. I sensed that code-mixed languages will become an issue eventually.

Personally, I am skeptical of LLMs from an ethical perspective. They are often trained on enormous amounts of data collected without meaningful consent from the people who produced it, require significant computational and environmental resources, and concentrate power in the hands of a few large companies. They also tend to reproduce existing social biases and inequalities at scale. While some harms can be mitigated through better governance and transparency, I am not convinced that today’s dominant approach to building ever-larger models is ethical or necessary for many of the problems the industry claims to be solving.

Why are code-mixed languages [that blend vocabulary and expressions from different languages, such as Hinglish], or ‘low-resourced’ languages like Maithili or Newari, less represented than others? Is this a limitation of data, of model architecture, of missing funding and political will?

Aindriya: It’s all of those things. Initially, our language wasn’t digitised, we weren’t typing in our languages, so there wasn’t enough digital data. But over the years, more people have come online and started writing in regional languages — in messages, DMs, or making reels — so now a lot of our spoken language is digitally available. It’s not ‘low-resource’ in the way it used to be.

Now I would say it’s more a lack of political will. The major companies build these platforms and products and dump them into the global majority countries where there’s a high population, so they get a lot of users for their products. But they don’t really care about our specific context or problems. They get our data, but they don’t use it to solve our problems — only to sell our data to advertisers, and now also for AI training.

There’s enough data now for generative AI training, but not for specific, task-based problem-solving — like the hate speech work I do. Task-specific datasets in our languages still don’t exist, because the people who’d build these datasets for the benefit of their communities are people like me, who often don’t have the money or resources to do it. 

Leki: In my region, I find that Dzongkha, Bhutan’s national language, is the most underrepresented, followed by languages like Tshangla, Khengkha, Kurtopekha, and many other dialects, spoken by smaller communities. These are living, thriving languages — people wake up, argue, joke, grieve, and fall in love in them — but they are invisible in the digital space.

I think the gap exists for structural reasons. Language models can only learn from text that exists in digital form and is accessible to them. A language that has never been written down, or that exists mostly in oral tradition, or whose speakers have limited internet access, simply doesn’t make it into these systems. 

There is also an economic logic at work: companies building AI systems prioritise languages with the largest number of potential paying users. English, Mandarin, Spanish, and Hindi, which can be investments. Dzongkha, spoken by under a million people, is not viable, at least not from a commercial standpoint. So AI systems tend to ‘see’ these languages far less frequently than dominant global languages.

Why is it challenging to capture languages and dialects spoken by the global majority countries?

Aindriya: The main challenge, for instance, in code-mixed language and datasets, is the lack of fixed grammar and fixed spelling. Because people are phonetically thinking in their heads about how a word should look on the English keyboard, they end up typing it in different ways. “Main”,  [me in Hindi], can be written as “mai” (M-A-I); some people write “main” (M-A-I-N); some people just write “M.” In “udhar hai” — some write “hai” (H-A-I); some write “hain” (H-A-I-N); some just write “he” (H-E); some just “H”; some “ha” (H-A). It all depends on how people think of the word, and the reader then understands what it means from context. In English, you can use capitalisation to identify proper nouns, but here, there’s nothing fixed – there are no fixed spellings, and there’s no capitalisation because there’s no grammar. 

It leads to some issues. For instance, with content moderation on platforms owned by Meta. The only power these platforms give you is to let you change your settings to include words you don’t want to see in comments or DMs. That doesn’t work for code-mixed languages – there are no fixed spellings, people write things in countless ways, and they can also use stars, hashes, or emojis to obfuscate that they’re writing hate speech. 

What errors or biases emerge, or what can go wrong, when a language model is not trained on local languages? 

Leki: My company developed a chatbot for our e-commerce portal, but even this locally developed chatbot, built by our own team, has limited exposure to many local terms and expressions. The tool is often unable to fully meet users’ needs. For example, ‘bangchung’ is a commonly used Tshangla [a Sino-Tibetan language] term referring to a small woven bamboo basket, yet such culturally specific vocabulary is often absent from the data used to train AI systems. Another example: consider a flood early warning system, this is very critical for Bhutan, and these are being piloted in several South Asian countries. If the system’s instructions, alerts, or response prompts are only available in a national language that a remote community doesn’t fully read or speak fluently, the warning may as well not exist. And if an AI system is being used to translate or localise those alerts in real time, and it makes errors — saying “water level is stable” when it means “water level is rising”, the consequences can be fatal.

Nova: In my research, we interviewed young women in computing about their challenges and opportunities of building tech careers in Bangladesh. The interviews were translated, transcribed and annotated by humans, and at the same time, annotated using existing models. The NLP-based models missed the core emotion and sentiment in many cases. Sometimes when a sentence started with a positive emotion, the model interpreted the whole context as positive, even though the narratives were gritty and difficult. The subtle indications were easily missed. If not for human intervention, the machine would have misunderstood the entire tone of how one felt.

Aindriya: If a model isn’t trained on local languages, it definitely doesn’t understand local context, and marginalised people face the most harm from these systems. When we report hate speech or abusive content, most of the time it doesn’t get removed. I have never once gotten a response from Meta saying, “Yes, this is abusive, and we’re removing it.” It’s always that it doesn’t go against their community guidelines, because these systems simply aren’t built to understand this language or this kind of abuse in our context, [as documented here and here]. Take Pranshu’s case: there were thousands of hate speech comments, and that kid took his own life. People reported those comments again and again, and they were never removed — it kept saying the content didn’t violate community guidelines. 

When these systems are built in the Global North with a centralised understanding of what hate speech is, or what language is being used, they don’t account for the fact that we exist, that our languages are different, and that our contexts are different. The system completely fails our people, and it’s mostly marginalised people who pay the cost of those failures.

Aindriya, your project and other research showed AI systems do not recognise harm in low-resource languages, and this has led to real-world violence. What did that look like?

Aindriya: When we collected a random sample of hate speech for our training data, 40% of it was directed at gender and sexual minorities, and of that, 19% was specifically queerphobic. When we looked at data around election periods, most of the political hate speech was communal, and further analysis showed it was largely Islamophobic, so we could trace how certain groups were being targeted around elections. 

The reason I mention this is that all of this was a random sample of hate speech that existed on these platforms without being removed — it had been considered “safe” by the existing content moderation systems. Out of 50,000 comments, roughly half were hate speech, some of it really violent, targeting marginalised groups, and none of it was removed. Content moderation that wasn’t trained on our language and context failed to recognise hate speech against marginalised communities — and we have the data to show it. 

Hate speech isn’t something you can just delete and block your way out of. At a personal level, sure, you can delete a comment and block someone, but it’s part of a larger system of brigading against marginalised people. It normalises hate speech and the narratives behind it — you see so much of it that people get turned into jokes and trolling fodder, so that when real harm happens to them, people don’t care as much.

When the government brings in something like the Trans Act, and trans people are harmed, people don’t care enough because they’ve already been dehumanised by long-term exposure to hate speech and trolling.

During the recent West Bengal elections and SIR exercise, I was seeing a lot of reels around the beef ban. The ban wasn’t only spreading hate against Muslims — it was also hurting very poor Dalit people who depended on selling cattle. My first reaction was sadness, and I hoped that through the reel, people will understand the ground reality. But then I opened the comments, and people were laughing and trolling, saying things like, “Good, this should happen to them”. I was stunned. And I realised that’s what long-term exposure to hate speech and trolling does — it has dehumanised Dalit and Muslim people.

A concern with lacking diversity of language and knowledge systems, is that systems end up surfacing stereotypes. Where do AI models pick up biases — like stereotypes about women, or prejudice against Muslims or caste communities? Can those biases be fixed?

Nova: The language mode is learning from what it is being taught. So, if the majority of data considers a mother’s roles in certain ways or certain religions to be placed in a particular way, the model will learn in the same way. These biases are continuously being monitored and observed but the problem lies somewhere else. The various web-based platforms (e.g., social media platforms and communication platforms) do not monitor the abuse and harassment of minority communities; even worse, they monetise over the negative discussions. 

Aindriya: It really depends on whose data you’re training on and who tags the data. There are different kinds of hate speech — if an upper-caste celebrity is being trolled, that’s hate speech, but a marginalised person being trolled is also hate speech, and the language used is very different. It depends a lot on where I’m collecting data from — am I collecting it from an upper-caste actress being trolled, or from a Dalit artist or activist being trolled? Those are very different kinds of data, and it depends on who’s tagging it and deciding what counts as hate speech — that’s where bias comes from. Like Elon Musk saying that calling someone “cis” is hate speech — we’d have a completely different definition.

Any model will reproduce the same biases that exist in the real world. If the narrative it’s built on is patriarchal, the model will be patriarchal. If one caste is overrepresented in the data, the model will learn that. It always comes down to where and from whom the data is collected.

Under the flagship initiative BharatGen, India is investing close to Rs 1,200 crore to promote linguistic diversity and build indigenous LLMs in 22 official languages. The gold rush of ‘multilingual AI’ has raised concerns about accuracy and community consent among experts. When you digitise a language, what risks come with that? Can standardisation harm a living language? 

Nova: I have found some encouraging initiatives they allowed the collection of data around languages which were not previously documented. They could also, in the long run, serve the purpose of preserving these languages. But I am not sure about the efficacy of these models. Independent regulatory bodies need to ensure that the process does not create any power practice over the languages and how they are being trained and used.

Leki: I find myself thinking most about this question. Digitisation is not neutral. When you write down a spoken language, you immediately face choices: which dialect do you use as the standard? How do you represent sounds that don’t exist in any available script? Whose usage is considered “correct”? These choices, once encoded into a dataset and then into an AI system, can calcify. A language that exists across many valleys, each with its own accent, vocabulary, and rhythm, gets flattened into one “official” digital version. Over time, speakers, especially younger ones who interact with AI tools, may begin to adjust their usage toward what the machine understands, rather than what they grew up speaking.

Another challenge is that technology often prefers consistency. AI systems perform better when spelling, grammar, and vocabulary are standardised. But living languages are rarely uniform. They evolve across regions, generations, and communities. If a single dialect becomes the “official” digital version, other dialects may gradually become less visible. Over time, technology can unintentionally amplify one form of a language while marginalising others.

Another concern is cultural flattening. Languages carry local knowledge, humour, values, and ways of seeing the world. If digitisation focuses only on creating efficient machine-readable text, some of that richness may be lost. There is also a governance question: Who decides what counts as the correct version of a language? Linguists, governments, educators, elders, or speakers themselves may all have different perspectives.

How do you make sure a community’s consent and knowledge actually come first when you’re collecting their data? 

Aindriya: For my project, people joined and helped me voluntarily — it became a completely community-led, community-funded project, because I started without any institutional funding. I think a lot about ethical AI, and I don’t support generative AI at all; I don’t work with it. I work entirely with small-scale, task-specific data. I’m not collecting huge amounts of random data and feeding it all into a model — I’m specifically looking for hate speech and non-hate speech comments in a specific language, and taking only the minimum amount of data needed, partly because I don’t have infinite resources to work at a huge scale. All my work is done without a GPU [a graphic processing unit, the workhorses that power AI systems], without much money. So, it doesn’t harm the environment much, simply because I’m not using that many resources.

As for consent, I sent regular updates about how much data had been collected, how much was tagged, what the data looked like, what percentage represented different marginalised groups, so we’d have a good understanding of representation in the dataset itself. Everyone involved had lived experience and came from communities with multiple, intersecting marginalities. They had the power to decide what counts as hate speech, and the model understood hate speech as they defined it — which I think is really powerful – because in civil society work, we’re often the research subjects without having a say in the final decision or the final product. 

Leki: I think it is important to be honest about what we are asking of communities. We are asking people to give us a significant amount of data, their words – recorded speech, written text, stories, proverbs, conversational exchanges. We are asking elders to sit with recorders… we are asking communities to trust that their language will be treated with care by a technology that most of them have never used and may not fully understand.

In return, we try to offer something tangible, like digital literacy training, or a tool in their own language that solves a real problem they have identified, or commitments about how the data will be used, stored, and governed. Increasingly, communities, especially those with histories of having their cultural knowledge extracted without credit or benefit, are asking hard questions. Who owns this data? Can you use it to build a commercial product without our permission? What happens if we want it removed? These are entirely legitimate questions, and frankly, the AI industry has no good answers yet.

Lastly, what would it take to build a language model that’s ‘good’ for everyone, not just an average paying user? 

Aindriya: We’re currently seeing a trend where general-purpose GenAI is being integrated into almost every application, regardless of whether it’s actually needed. For example, many apps now have chatbots that can answer domain-specific questions, but also respond to completely unrelated prompts. I find that approach inefficient. Not every product needs a general-purpose chatbot, and I think we’re losing some of that discipline in the rush to adopt GenAI everywhere.

Most industry problems can be solved effectively with traditional machine learning approaches without requiring massive amounts of training data or large-scale GPU infrastructure. [Machine learning is a form of AI that analyses data and finds patterns, enabling technologies like GPS navigation systems and face recognition. A subset of this is Generative AI, which requires more data and computational power.]

For many applications, it is sufficient to train a model on a dataset tailored to a specific task and optimise it for that objective. This often results in systems that are more efficient, easier to maintain, and better aligned with the actual requirements of the product.

Nova: Technology was never neutral and, in many cases, not inclusive for all. When I was researching available tech around harassment reporting, I found that elderly women and women in rural areas needed it the most, but they were far beyond the tech-based reporting process. If we want to design for marginalised communities, we have to think about ensuring their data is protected and not misused or used for profit without consent, and the models are useful and bias-free. The models should be designed around feminist and humanitarian principles. I believe a language model will be really useful when a street vendor or a rural entrepreneur can use it seamlessly, and when the model responds, it is fair, just and validated.

  • Saumya Kalia is a Delhi-based journalist who writes about gender, labour, and social equity. She has won the Laadli and REACH Media Awards for her gender journalism, and reported on gender and healthcare as a Dr. Amit Sengupta Health Rights Fellow. At BehanBox she is working on developing editorial series, building quieter spaces, and redefining news engagement across different platforms. She is deeply interested in thinking about grief, care, community, and cities.

Malini Nair (Editor)

Malini Nair is a consulting editor with Behanbox. She is a culture writer with a keen interest in gender.

Support BehanBox

We believe everyone deserves equal access to accurate news. Support from our readers enables us to keep our journalism open and free for everyone, all over the world.

Donate Now