The Gap in Numbers
Pashto is one of the world's major languages by speakers — roughly 60 million people across Afghanistan, Pakistan, and diaspora communities — and Dari is the other official language of Afghanistan plus a major tongue in the region. But look at the NLP landscape and the numbers don't match the population:
- Datasets: a tiny fraction of English's web-scale corpora, and historically fragmented.
- Models: almost no dedicated Pashto LLMs; support inside multilingual models is partial and inconsistently evaluated.
- Evaluation: benchmarks are built for English (and a few major languages); Pashto and Dari rarely appear.
In 2026 the picture is starting to shift. The PashtoCorp corpus — 1.25 billion words — is a major step, giving researchers a serious foundation where almost none existed. But one corpus doesn't close a decade of underinvestment.
Why It Happens
📉 Data scarcity
Modern LLMs are trained on web-scale text. Pashto and Dari have a fraction of the digitized content — and much of it is fragmented across dialects and scripts.
🔤 Script complexity
Arabic-script languages need careful tokenization and normalization. The Arabic script's context-dependent forms trip up generic pipelines.
🧮 The economics
Platforms prioritize markets where data is abundant and monetizable. Pashto and Dari fall below the commercial threshold — a market failure, not a technical one.
📊 Evaluation blind spots
"Multilingual" models are often measured on a handful of languages. What isn't measured isn't improved.
What's Finally Changing in 2026
- PashtoCorp (1.25B words) — the first serious large-scale Pashto corpus, announced March 2026, with an evaluation suite attached.
- SilkRoadNLP — an ACL workshop dedicated to the Iranian language family: Persian, Dari, and Tajiki, with cross-lingual research across all three.
- AbjadNLP — a workshop focused on NLP for Arabic-script languages, covering low-resource and minority languages.
- Better multilingual bases — open-weight models with stronger low-resource coverage to fine-tune from.
None of this makes Pashto/Dari AI "solved." It makes the work possible — which is exactly where a capable builder with local-first tools can make a real difference.
What Actually Works, in Practice
- Glossary-locked translation. Curated domain glossaries (legal, medical, government) keep terminology consistent — the single highest-leverage technique for Pashto/Dari work.
- Fine-tuning on real data. A small, clean Pashto/Dari dataset adapted onto a multilingual base beats a generic model every time.
- Cross-lingual transfer. Persian resources are far richer — techniques that transfer improvements from Persian to Dari and Pashto pay off.
- Local, private deployment. Offline AI is often the only option for speakers in restricted environments — and it keeps the work out of cloud platforms that ignore these languages anyway.
- Evaluation beyond English. Build small, real evaluation sets in Pashto and Dari and measure what matters.
💡 Working examples on this site: Pashto and Dari in AI: What Works in 2026 covers the current model landscape, and Building a Multilingual Translation Pipeline with Local LLMs shows the glossary-locked pipeline in action.
Why It Matters
This isn't a niche concern. It's tens of millions of people whose access to AI-era tools — translation, education, legal help, government services — is being rationed by data economics. Global Voices put it bluntly in 2026: if the status quo holds, low-resource language communities keep losing ground in the race to unlock AI's potential. Building for Pashto and Dari is both a market opportunity and a correction.
Frequently Asked Questions (FAQ)
Why does AI perform poorly on Pashto and Dari?
Because these are low-resource languages: they have a fraction of the training data of English or even major regional languages, and most model families don't include them. Pashto — spoken by ~60 million people — has historically had very limited NLP resources, though new corpora like PashtoCorp (1.25 billion words) are changing that.
Is AI improving for Pashto and Dari?
Yes, slowly. 2026 brought the PashtoCorp 1.25-billion-word corpus, the SilkRoadNLP workshop for the Iranian language family (Persian, Dari, Tajiki), and better multilingual models. Progress is real but still far behind major languages.
Can I run AI in Pashto on my own device?
Yes. Multilingual open-weight models handle Pashto and Dari with varying quality, and local setups can be improved with glossaries and fine-tuning. Offline, private AI is often the only option for speakers in restricted environments.
What can be done to improve AI for Pashto and Dari?
Build and share more datasets, adapt multilingual models with glossaries and fine-tuning, evaluate beyond English benchmarks, and deploy local AI so speakers — not just platforms — benefit. Low-resource NLP techniques that work: data augmentation, cross-lingual transfer, and retrieval over curated glossaries.
Why should organizations care about Pashto/Dari AI?
Tens of millions of speakers, underserved communities, and real use cases: translation, education, legal and government services, journalism. Organizations that serve these speakers gain a genuine competitive and social edge.
🌍 Building AI for Pashto, Dari, or Persian?
I build multilingual AI for underrepresented languages — glossary-locked translation, fine-tuned local models, and private deployment. Contact me — this is exactly the kind of work I specialize in.