Why AI Still Ignores Pashto and Dari — and What to Do About It

Published: August 9, 2026 — Pashto is spoken by roughly 60 million people, and Dari by tens of millions more. Yet in the AI era, they're afterthoughts — languages with a fraction of the training data, almost no dedicated models, and evaluations that barely measure them. This is the problem, why it persists, and what's finally changing in 2026.

🌍 Quick Takeaways

The Gap in Numbers

Pashto is one of the world's major languages by speakers — roughly 60 million people across Afghanistan, Pakistan, and diaspora communities — and Dari is the other official language of Afghanistan plus a major tongue in the region. But look at the NLP landscape and the numbers don't match the population:

In 2026 the picture is starting to shift. The PashtoCorp corpus — 1.25 billion words — is a major step, giving researchers a serious foundation where almost none existed. But one corpus doesn't close a decade of underinvestment.

Why It Happens

📉 Data scarcity

Modern LLMs are trained on web-scale text. Pashto and Dari have a fraction of the digitized content — and much of it is fragmented across dialects and scripts.

🔤 Script complexity

Arabic-script languages need careful tokenization and normalization. The Arabic script's context-dependent forms trip up generic pipelines.

🧮 The economics

Platforms prioritize markets where data is abundant and monetizable. Pashto and Dari fall below the commercial threshold — a market failure, not a technical one.

📊 Evaluation blind spots

"Multilingual" models are often measured on a handful of languages. What isn't measured isn't improved.

What's Finally Changing in 2026

None of this makes Pashto/Dari AI "solved." It makes the work possible — which is exactly where a capable builder with local-first tools can make a real difference.

What Actually Works, in Practice

  1. Glossary-locked translation. Curated domain glossaries (legal, medical, government) keep terminology consistent — the single highest-leverage technique for Pashto/Dari work.
  2. Fine-tuning on real data. A small, clean Pashto/Dari dataset adapted onto a multilingual base beats a generic model every time.
  3. Cross-lingual transfer. Persian resources are far richer — techniques that transfer improvements from Persian to Dari and Pashto pay off.
  4. Local, private deployment. Offline AI is often the only option for speakers in restricted environments — and it keeps the work out of cloud platforms that ignore these languages anyway.
  5. Evaluation beyond English. Build small, real evaluation sets in Pashto and Dari and measure what matters.

💡 Working examples on this site: Pashto and Dari in AI: What Works in 2026 covers the current model landscape, and Building a Multilingual Translation Pipeline with Local LLMs shows the glossary-locked pipeline in action.

Why It Matters

This isn't a niche concern. It's tens of millions of people whose access to AI-era tools — translation, education, legal help, government services — is being rationed by data economics. Global Voices put it bluntly in 2026: if the status quo holds, low-resource language communities keep losing ground in the race to unlock AI's potential. Building for Pashto and Dari is both a market opportunity and a correction.

Frequently Asked Questions (FAQ)

Why does AI perform poorly on Pashto and Dari?

Because these are low-resource languages: they have a fraction of the training data of English or even major regional languages, and most model families don't include them. Pashto — spoken by ~60 million people — has historically had very limited NLP resources, though new corpora like PashtoCorp (1.25 billion words) are changing that.

Is AI improving for Pashto and Dari?

Yes, slowly. 2026 brought the PashtoCorp 1.25-billion-word corpus, the SilkRoadNLP workshop for the Iranian language family (Persian, Dari, Tajiki), and better multilingual models. Progress is real but still far behind major languages.

Can I run AI in Pashto on my own device?

Yes. Multilingual open-weight models handle Pashto and Dari with varying quality, and local setups can be improved with glossaries and fine-tuning. Offline, private AI is often the only option for speakers in restricted environments.

What can be done to improve AI for Pashto and Dari?

Build and share more datasets, adapt multilingual models with glossaries and fine-tuning, evaluate beyond English benchmarks, and deploy local AI so speakers — not just platforms — benefit. Low-resource NLP techniques that work: data augmentation, cross-lingual transfer, and retrieval over curated glossaries.

Why should organizations care about Pashto/Dari AI?

Tens of millions of speakers, underserved communities, and real use cases: translation, education, legal and government services, journalism. Organizations that serve these speakers gain a genuine competitive and social edge.

🌍 Building AI for Pashto, Dari, or Persian?

I build multilingual AI for underrepresented languages — glossary-locked translation, fine-tuned local models, and private deployment. Contact me — this is exactly the kind of work I specialize in.