OpinionPREMIUM

AI needs lessons in African languages

When most people hear the term Artificial Intelligence their thoughts turn instinctively to a generative platform such as ChatGPT, the large language model now used by more than 200-million people each week.

We’ve even known for a decade that it can be sexist or racist.
We’ve even known for a decade that it can be sexist or racist. (123RF)

When most people hear the term Artificial Intelligence their thoughts turn instinctively to a generative platform such as ChatGPT, the large language model now used by more than 200-million people each week. One would expect a system with such reach to reflect the diversity of those it serves. But ask it, “How fluent are you in African languages compared to European ones?” and the answer is unambiguous: it isn’t. 

It acknowledges fluency in English, French, Spanish and German, attributing this to the volume and diversity of training data available in those languages. Its understanding of African languages, by contrast, is limited and uneven. While it can process basic content in widely spoken languages such as Swahili, Yoruba, Zulu and Afrikaans, it concedes that its grasp of grammar, idiom and regional variation remains rudimentary in comparison. 

The explanation is technical — training data, usage patterns, representation gaps — but the implication is structural. The language systems most deeply embedded in African life remain peripheral in the architectures now shaping the digital economy. 

And where language is peripheral, access soon follows. 

According to a recent Artificial Intelligence for Development (AI4D) study, the total of African languages accounts for 0.02% of all internet content — 2,650 times less than the 53% represented in English. Even the most digitally present African languages are statistically marginal: Afrikaans at 0.003% and Venda at 0.000115%. And these are the high points.

The implications for AI are profound. Large language models depend on the breadth, depth and quality of publicly available digital data to train their systems. When language representation is this unbalanced, so too is the intelligence it produces. As the global economy becomes mediated by AI, language exclusion translates directly into digital exclusion. Systems become less accessible and less representative for the populations that stand to benefit most from AI-enabled growth. 

Nowhere is this tension more operationally significant than in financial services. 

Financial inclusion is often mediated through language — whether in onboarding journeys, product terms and conditions, fraud alerts or customer support interactions. Across the continent, banks have accelerated their efforts to extend access to underserved populations. Mobile-first account origination, agent banking networks, simplified Know Your Customer (KYC) protocols and digital wallets have become core components of financial service design — engineered to reduce friction and expand participation across diverse operating environments.

Yet as banks deepen their reliance on AI to scale and personalise services — a matter of increasing focus in the sector — they do so on architectures that do not recognise how most Africans speak. And when the system cannot understand a customer’s language, it cannot fully serve their needs — no matter how smart the technology appears to be. 

That disconnect is beginning to draw the attention it warrants. 

In April, a landmark resolution at the Global AI Summit for Africa saw the commitment of $60bn to build a robust AI ecosystem across the continent — an inflection point that recognised both the urgency and the potential of African-led innovation. Banks, too, are beginning to shape a more linguistically inclusive digital future. Absa’s partnership with the University of Pretoria (UP), for instance, is exploring how machine learning and natural language processing can be harnessed to better extract meaning from African linguistic data — not only to safeguard heritage, but to embed it within the next generation of digital financial infrastructure.   

But the real test will lie in scale. 

Absa and UP have laid the groundwork. Now it’s time for the rest of South Africa to step up, support the initiative and help the nation unlock its full potential. Only through collective commitment can we fully realise the social and economic benefits of a truly inclusive AI landscape. 

The aim is not to localise systems after the fact, but to design them from the outset with linguistic relevance in mind. That means developing models attuned to the syntactic structures, tonal patterns and pragmatic cues that define African speech. It means training those models on data that reflects how people actually speak — not how systems expect them to. And it means embedding linguistic representation within the governance frameworks that now shape AI deployment in financial services, not leaving it to technical discretion. 

This will be central to expanding financial access in the age of AI.

As digital interfaces expand, the ability to engage users in their own languages will determine reach and relevance. Systems trained on African languages can onboard clients more effectively, resolve queries with greater accuracy and extend services to users previously excluded by language barriers — particularly in rural or linguistically diverse markets. 

The impact is both operational and relational. Clearer communication reduces transaction errors, improves fraud detection and strengthens compliance. But just as critically, it builds trust. When users can navigate financial systems in the languages they live by, adoption deepens and engagement becomes sustained. In this way, linguistic infrastructure becomes a foundation for enduring financial inclusion. 

• Christine Wu is CEO of everyday banking, Absa Group; Vukosi Marivate is professor of computer science, and holds the Absa UP chair of data science at the University of Pretoria. 


Would you like to comment on this article?
Sign up (it's quick and free) or sign in now.

Comment icon

Related Articles