Speech & text data for the languages AI never heard.
We build open voice and text datasets for South Asia's underserved languages — from Urdu and Bengali to Torwali and Khowar — and publish them on Mozilla Data Collective for researchers, developers, and AI labs everywhere.
The gaps are the reason we exist.
Every filled cell is a published dataset on Mozilla Data Collective. Every empty cell is a language–capability pair that still has no open data at all. Solid cells are published; ringed cells are in development.
| Language | Corpus | Lexicon | MT | ASR | TTS | Vision | Multimodal |
|---|---|---|---|---|---|---|---|
| Indo-Aryan | |||||||
| Saraiki[srk] | |||||||
| Urdu[urd] | |||||||
| Bengali[ben] | |||||||
| Punjabi[pnb] | |||||||
| Sindhi[snd] | |||||||
| Hindi[hin] | |||||||
| Gujarati[guj] | |||||||
| Marathi[mar] | |||||||
| Chittagonian[ctg] | |||||||
| Noakhalian[oak] | |||||||
| Rangpuri[rkt] | |||||||
| Rohingya[rhg] | |||||||
| Sylheti[syl] | |||||||
| Gojri[gju] | |||||||
| Dardic | |||||||
| Khowar[khw] | |||||||
| Torwali[trw] | |||||||
| Kohistani Shina[plk] | |||||||
| Gawri[gwc] | |||||||
| Iranian | |||||||
| Hazargi[haz] | |||||||
| E. Balochi[bgp] | |||||||
| Dari[prs] | |||||||
| Persian[pes] | |||||||
| W. Balochi[bgn] | |||||||
| Dravidian | |||||||
| Brahui[brh] | |||||||
| Tamil[tam] | |||||||
| Kannada[kan] | |||||||
| Telugu[tel] | |||||||
| Malayalam[mal] | |||||||
| Multiple | |||||||
| Multi-language[multi] | |||||||
Cell colour = datasets published for that pair: 1234+ · none yet · ring = in development
Counts from the published catalogue below. 21% coverage after three years — the remaining 79% is the roadmap.
Miras means heritage.
We turn spoken heritage into open, machine-readable data — so the next generation of AI can speak with everyone, not just the internet's biggest languages.
Text corpora only; speech hours counted separately. 55M+ tokens published across the catalogue.
Peer-reviewed at LREC 2026. Of 532.6 hours recorded across 39 languages, 493.6 hours passed community validation.
63 datasets. Every row links to the source.
Text, speech, lexicon, and computer-vision resources for South Asian and low-resource languages. Every row links through to Mozilla Data Collective.
MirasAI is the listed steward for 19 of these datasets, shown with their verified licence below. The remainder were curated by MirasAI in partnership with publishers, community organisations, and language institutions, and are stewarded by those partners on Mozilla Data Collective.
| Dataset | ISO | Family | Task | Format | Volume | Licence |
|---|---|---|---|---|---|---|
| Hindi Literature & News Corpus | [hin] | Indo-Aryan | Corpus | TXT | 19.52 MB | CC-BY-NC-SA-4.0 |
| Tamil Literature Corpus | [tam] | Dravidian | Corpus | TXT | 41.85 MB | CC-BY-NC-SA-4.0 |
| Telugu Text Corpus | [tel] | Dravidian | Corpus | TXT | 19.87 MB | CC-BY-NC-SA-4.0 |
| Gujarati News & Blogs Corpus | [guj] | Indo-Aryan | Corpus | TXT | 18.93 MB | CC-BY-NC-4.0 |
| Marathi Blog & Literature Corpus | [mar] | Indo-Aryan | Corpus | TXT | 4.20 MB | CC-BY-NC-4.0 |
| Kannada Text Corpus | [kan] | Dravidian | Corpus | TXT | 2.61 MB | CC-BY-NC-SA-4.0 |
| Chittagonian Text Corpus | [ctg] | Indo-Aryan | Corpus | TXT | 6.59 MB | CC-BY-NC-SA-4.0 |
| Noakhalian Text Corpus | [oak] | Indo-Aryan | Corpus | TXT,DOCX | 3.61 MB | CC-BY-NC-4.0 |
| Rangpuri Text Corpus | [rkt] | Indo-Aryan | Corpus | TXT | 3.68 MB | CC-BY-NC-4.0 |
| Rohingya Literature Corpus | [rhg] | Indo-Aryan | Corpus | TXT,DOCX | 7.01 MB | CC-BY-NC-4.0 |
| Sylheti Text Corpus | [syl] | Indo-Aryan | Corpus | TXT | 3.53 MB | partner |
| Gojri Literature Corpus | [gju] | Indo-Aryan | Corpus | TXT | — | partner |
| Khowar Literature Corpus (FLI) | [khw] | Dardic | Corpus | TXT | — | partner |
| Khowar Word List | [khw] | Dardic | Lexicon | TXT | — | partner |
| Kohistani Shina Word List | [plk] | Dardic | Lexicon | TXT | — | partner |
| Western Balochi Literature Corpus | [bgn] | Iranian | Corpus | TXT | ~1.1M tok | partner |
| NAWA-E-WATAN Balochi Newspaper Corpus | [bgp] | Iranian | Corpus | TXT | ~1.02M tok | partner |
| Eastern Balochi Literature Corpus | [bgp] | Iranian | Corpus | TXT | — | partner |
| Gawri Magazine Corpus | [gwc] | Dardic | Corpus | TXT | — | partner |
| Saraiki–English Parallel Corpus | [srk] | Indo-Aryan | MT | TXT | 51,447 sent | partner |
| Jhoke Publisher Saraiki Newspaper Corpus | [srk] | Indo-Aryan | Corpus | TXT | 1.25M tok | partner |
| IBT Torwali Wordlist | [trw] | Dardic | Lexicon | TXT | 20,000 words | partner |
| Torwali Text Corpus | [trw] | Dardic | Corpus | TXT | — | partner |
| Elkhani Hazargi Literature Corpus | [haz] | Iranian | Corpus | TXT | 500K tok | partner |
| Dari 1 Million Text Corpus | [prs] | Iranian | Corpus | TXT | 1M tok | partner |
| Persian 1 Million Text Corpus | [pes] | Iranian | Corpus | TXT | 1M tok | partner |
| Jugantor Newspaper Corpus | [ben] | Indo-Aryan | Corpus | TXT | 10M words | partner |
| Kanto Newspaper Corpus | [ben] | Indo-Aryan | Corpus | TXT | 10M words | partner |
| Prothom Alo Newspaper Corpus | [ben] | Indo-Aryan | Corpus | TXT | 10M tok | partner |
| Hindi 10 Million Text Corpus | [hin] | Indo-Aryan | Corpus | TXT | 10M tok | partner |
| Punjabi Literature Corpus (Shahmukhi) | [pnb] | Indo-Aryan | Corpus | TXT | 1.04M tok | partner |
| Punjabi Dataset-02 | [pnb] | Indo-Aryan | Corpus | TXT | — | partner |
| Punjabi Dataset-03 | [pnb] | Indo-Aryan | Corpus | TXT | — | partner |
| Saraiki Literature Corpus (Kaleem Art Press) | [srk] | Indo-Aryan | Corpus | TXT | — | partner |
| Saraiki Corpus (Shabbir Baloch) | [srk] | Indo-Aryan | Corpus | TXT | — | partner |
| Saraiki Corpus (Sujaak Adbi Sangat) | [srk] | Indo-Aryan | Corpus | TXT | — | partner |
| Kaleem Art Press Urdu Literature Corpus | [urd] | Indo-Aryan | Corpus | TXT | 1.44M tok | partner |
| Weekly Kaleem Magazine Corpus | [urd] | Indo-Aryan | Corpus | TXT | ~1.4M tok | partner |
| Rana Printers Urdu Literature Corpus | [urd] | Indo-Aryan | Corpus | TXT | 1.68M tok | partner |
| Bismillah Graphics Publishers Urdu Corpus | [urd] | Indo-Aryan | Corpus | TXT | 1.62M tok | partner |
| Sindh Sujag Newspaper Corpus | [snd] | Indo-Aryan | Corpus | TXT | — | partner |
| Sindh Line Newspaper Corpus | [snd] | Indo-Aryan | Corpus | TXT | — | partner |
| Tamir Sindhi Newspaper Corpus | [snd] | Indo-Aryan | Corpus | TXT | 1.1M tok | partner |
| Jazab Sindhi Newspaper Corpus | [snd] | Indo-Aryan | Corpus | TXT | — | partner |
| Brahui Dataset 1 | [brh] | Dravidian | Corpus | TXT | — | partner |
| Brahui Dataset 2 | [brh] | Dravidian | Corpus | TXT | — | partner |
| Brahui Dataset 3 | [brh] | Dravidian | Corpus | TXT | — | partner |
| Hazaragi Dataset | [haz] | Iranian | Corpus | TXT | — | partner |
| Farsi Dataset | [pes] | Iranian | Corpus | TXT | — | partner |
| Dari Dataset | [prs] | Iranian | Corpus | TXT | — | partner |
| Kannada Time-Aligned Speech Corpus | [kan] | Dravidian | ASR | OGG,SRT | 355.77 MB | CC-BY-NC-SA-4.0 |
| Tamil Time-Aligned Speech Dataset | [tam] | Dravidian | ASR | OGG,SRT | 37.11 MB | CC-BY-NC-SA-4.0 |
| Multispeaker Hindi ASR Dataset | [hin] | Indo-Aryan | ASR | OGG,SRT | 63.03 MB | CC-BY-NC-SA-4.0 |
| Punjabi 10 Hours TTS | [pnb] | Indo-Aryan | TTS | WEBM,TSV | 481.96 MB | CC-BY-NC-SA-4.0 |
| Saraiki 10 Hours TTS Dataset | [srk] | Indo-Aryan | TTS | WEBM,TSV | 584.44 MB | CC-BY-NC-SA-4.0 |
| Malayalam Time-Aligned Speech Corpus | [mal] | Dravidian | ASR | SRT | 5 speakers | partner |
| Urdu 10 Hours TTS | [urd] | Indo-Aryan | TTS | WEBM,TSV | 10 hrs | partner |
| Pakistan Traffic Signs Dataset | [urd] | Indo-Aryan | Vision | JPEG,JSON | 2.64 GB | CC-BY-NC-4.0 |
| Bangladesh Traffic Signs Dataset | [ben] | Indo-Aryan | Vision | JSON,JPEG | 1.35 GB | CC-BY-NC-4.0 |
| India Traffic Signs Dataset | [multi] | Multiple | Vision | JPEG,JSON | 190.44 MB | CC-BY-ND-4.0 |
| Hazargi Corpus for Speech Recognition | [haz] | Iranian | ASR | MP3,TSV | 511.56 MB | partner |
| Hazargi–English Speech Translation Corpus | [haz] | Iranian | MT | MP3,TSV | 511.58 MB | partner |
| Hazargi Multimodal Dataset | [haz] | Iranian | Multimodal | MP4,EAF | 44.00 GB | partner |
published · in development · licences shown where MirasAI is the listed steward.
29 languages, 5 families.
Indo-Aryan, Iranian, Dravidian, Dardic, and Indo-European contact languages — many with tens of millions of speakers and almost no machine-readable data before this catalogue.
Building the language infrastructure South Asian AI needs next.
Two funded data initiatives are expanding what researchers, product teams, and communities can build for underrepresented languages — building on a completed, peer-reviewed corpus covering 39 languages of Pakistan.
Common Voice Pakistan: an open speech corpus for 39 languages
A one-year, Mozilla-funded effort spanning Pakistan's Indo-Aryan, Iranian, Dardic, Turkic, and isolate language families — from Balochistan and Sindh up through Gilgit-Baltistan and Kashmir. Locally authored texts, everyday speech, poetry, and folk song, collected with native speakers and community organizations rather than by top-down design.
Alam, M. & Tyers, F. M. (2026). Common Voice for Pakistan: Developing an Open Speech Corpus for Low-Resource Pakistani Languages. Proc. LREC 2026, pp. 3355–3359. ELRA. · Released CC0.
500-hour Urdu Text-to-Speech dataset
One of the largest planned Urdu TTS efforts of its kind, designed to support natural voice technology, accessibility tools, and speech applications. High-quality open Urdu speech data remains scarce; this closes that gap with structured, usable data.
60-hour multimodal dataset, six Pakistani languages
AI-ready video, audio, transcription, and English translation for languages severely underrepresented in today's models — supporting ASR, machine translation, and conversational AI.
Technology services from strategy through delivery.
IT consulting & delivery
Technology strategy, system planning, implementation support, and dependable delivery for organizations and growing teams.
AI & data systems
Applied AI, automation, data pipelines, custom models, and evaluation workflows designed around real operational needs.
Digital platforms
Web products, internal tools, data experiences, and technology infrastructure built for usability and long-term value.
Language technology
Text, speech, translation, ASR, TTS, NLP, and multimodal resources for South Asian and low-resource languages.
Campaign support in Pakistan
Technology, digital coordination, localized execution, and field-facing support for international campaigns and activities.
Research & data partnerships
Ethical collection, curation, annotation, quality assurance, and delivery for research institutions and technology teams.
Working with teams building useful, inclusive technology.
Licensing, access & commissioning.
What languages does MirasAI have datasets for?
Can I commission a custom speech or text dataset?
What licence are MirasAI datasets released under?
Where can I download the open datasets?
Can I use MirasAI datasets to train commercial AI models?
What formats are datasets delivered in?
Need a technology partner who understands the region?
Talk to us about commissioned datasets, AI and data projects, digital products, language technology, or campaign support in Pakistan.