MirasAI
MirasAI · Open language data for Pakistan & South Asia

Speech & text data for the languages AI never heard.

We build open voice and text datasets for South Asia's underserved languages — from Urdu and Bengali to Torwali and Khowar — and publish them on Mozilla Data Collective for researchers, developers, and AI labs everywhere.

0
open datasets
0
languages
0.6h
speech recorded
0M+
text tokens
$0K+
grant funding
Coverage — 29 languages × 7 capabilities

The gaps are the reason we exist.

Every filled cell is a published dataset on Mozilla Data Collective. Every empty cell is a language–capability pair that still has no open data at all. Solid cells are published; ringed cells are in development.

LanguageCorpusLexiconMTASRTTSVisionMultimodal
Indo-Aryan
Saraiki[srk]
Urdu[urd]
Bengali[ben]
Punjabi[pnb]
Sindhi[snd]
Hindi[hin]
Gujarati[guj]
Marathi[mar]
Chittagonian[ctg]
Noakhalian[oak]
Rangpuri[rkt]
Rohingya[rhg]
Sylheti[syl]
Gojri[gju]
Dardic
Khowar[khw]
Torwali[trw]
Kohistani Shina[plk]
Gawri[gwc]
Iranian
Hazargi[haz]
E. Balochi[bgp]
Dari[prs]
Persian[pes]
W. Balochi[bgn]
Dravidian
Brahui[brh]
Tamil[tam]
Kannada[kan]
Telugu[tel]
Malayalam[mal]
Multiple
Multi-language[multi]
Published MirasAI datasets by language and data type. Hatched = no dataset published yet.

Cell colour = datasets published for that pair: 1234+  ·  none yet  ·  ring = in development

Fill rate
43 / 203
language × capability cells filled
Thinnest capability
Multimodal1
MT2
Lexicon3
TTS3
Vision3
ASR5
Corpus26

Counts from the published catalogue below. 21% coverage after three years — the remaining 79% is the roadmap.

میراث

Miras means heritage.

We turn spoken heritage into open, machine-readable data — so the next generation of AI can speak with everyone, not just the internet's biggest languages.

Scale — published volume to date
Token volume by language · log scale
Bengali30M
Hindi10M
Urdu6.14M
Saraiki2.14M
Balochi2.12M
Sindhi1.1M
Punjabi1.04M
Dari1.0M
Persian1.0M
Hazargi0.5M

Text corpora only; speech hours counted separately. 55M+ tokens published across the catalogue.

Common Voice Pakistan · validation

Peer-reviewed at LREC 2026. Of 532.6 hours recorded across 39 languages, 493.6 hours passed community validation.

493.6 hrs validated (92.7%)39.0 hrs pending
139,000+sentences contributed
1,058speakers nationwide
39languages spanned
75language consultants
The open catalogue · published on Mozilla Data Collective

63 datasets. Every row links to the source.

Text, speech, lexicon, and computer-vision resources for South Asian and low-resource languages. Every row links through to Mozilla Data Collective.

MirasAI is the listed steward for 19 of these datasets, shown with their verified licence below. The remainder were curated by MirasAI in partnership with publishers, community organisations, and language institutions, and are stewarded by those partners on Mozilla Data Collective.

DatasetISOFamilyTaskFormatVolumeLicence
Hindi Literature & News Corpus[hin]Indo-AryanCorpusTXT19.52 MBCC-BY-NC-SA-4.0
Tamil Literature Corpus[tam]DravidianCorpusTXT41.85 MBCC-BY-NC-SA-4.0
Telugu Text Corpus[tel]DravidianCorpusTXT19.87 MBCC-BY-NC-SA-4.0
Gujarati News & Blogs Corpus[guj]Indo-AryanCorpusTXT18.93 MBCC-BY-NC-4.0
Marathi Blog & Literature Corpus[mar]Indo-AryanCorpusTXT4.20 MBCC-BY-NC-4.0
Kannada Text Corpus[kan]DravidianCorpusTXT2.61 MBCC-BY-NC-SA-4.0
Chittagonian Text Corpus[ctg]Indo-AryanCorpusTXT6.59 MBCC-BY-NC-SA-4.0
Noakhalian Text Corpus[oak]Indo-AryanCorpusTXT,DOCX3.61 MBCC-BY-NC-4.0
Rangpuri Text Corpus[rkt]Indo-AryanCorpusTXT3.68 MBCC-BY-NC-4.0
Rohingya Literature Corpus[rhg]Indo-AryanCorpusTXT,DOCX7.01 MBCC-BY-NC-4.0
Sylheti Text Corpus[syl]Indo-AryanCorpusTXT3.53 MBpartner
Gojri Literature Corpus[gju]Indo-AryanCorpusTXTpartner
Khowar Literature Corpus (FLI)[khw]DardicCorpusTXTpartner
Khowar Word List[khw]DardicLexiconTXTpartner
Kohistani Shina Word List[plk]DardicLexiconTXTpartner
Western Balochi Literature Corpus[bgn]IranianCorpusTXT~1.1M tokpartner
NAWA-E-WATAN Balochi Newspaper Corpus[bgp]IranianCorpusTXT~1.02M tokpartner
Eastern Balochi Literature Corpus[bgp]IranianCorpusTXTpartner
Gawri Magazine Corpus[gwc]DardicCorpusTXTpartner
Saraiki–English Parallel Corpus[srk]Indo-AryanMTTXT51,447 sentpartner
Jhoke Publisher Saraiki Newspaper Corpus[srk]Indo-AryanCorpusTXT1.25M tokpartner
IBT Torwali Wordlist[trw]DardicLexiconTXT20,000 wordspartner
Torwali Text Corpus[trw]DardicCorpusTXTpartner
Elkhani Hazargi Literature Corpus[haz]IranianCorpusTXT500K tokpartner
Dari 1 Million Text Corpus[prs]IranianCorpusTXT1M tokpartner
Persian 1 Million Text Corpus[pes]IranianCorpusTXT1M tokpartner
Jugantor Newspaper Corpus[ben]Indo-AryanCorpusTXT10M wordspartner
Kanto Newspaper Corpus[ben]Indo-AryanCorpusTXT10M wordspartner
Prothom Alo Newspaper Corpus[ben]Indo-AryanCorpusTXT10M tokpartner
Hindi 10 Million Text Corpus[hin]Indo-AryanCorpusTXT10M tokpartner
Punjabi Literature Corpus (Shahmukhi)[pnb]Indo-AryanCorpusTXT1.04M tokpartner
Punjabi Dataset-02[pnb]Indo-AryanCorpusTXTpartner
Punjabi Dataset-03[pnb]Indo-AryanCorpusTXTpartner
Saraiki Literature Corpus (Kaleem Art Press)[srk]Indo-AryanCorpusTXTpartner
Saraiki Corpus (Shabbir Baloch)[srk]Indo-AryanCorpusTXTpartner
Saraiki Corpus (Sujaak Adbi Sangat)[srk]Indo-AryanCorpusTXTpartner
Kaleem Art Press Urdu Literature Corpus[urd]Indo-AryanCorpusTXT1.44M tokpartner
Weekly Kaleem Magazine Corpus[urd]Indo-AryanCorpusTXT~1.4M tokpartner
Rana Printers Urdu Literature Corpus[urd]Indo-AryanCorpusTXT1.68M tokpartner
Bismillah Graphics Publishers Urdu Corpus[urd]Indo-AryanCorpusTXT1.62M tokpartner
Sindh Sujag Newspaper Corpus[snd]Indo-AryanCorpusTXTpartner
Sindh Line Newspaper Corpus[snd]Indo-AryanCorpusTXTpartner
Tamir Sindhi Newspaper Corpus[snd]Indo-AryanCorpusTXT1.1M tokpartner
Jazab Sindhi Newspaper Corpus[snd]Indo-AryanCorpusTXTpartner
Brahui Dataset 1[brh]DravidianCorpusTXTpartner
Brahui Dataset 2[brh]DravidianCorpusTXTpartner
Brahui Dataset 3[brh]DravidianCorpusTXTpartner
Hazaragi Dataset[haz]IranianCorpusTXTpartner
Farsi Dataset[pes]IranianCorpusTXTpartner
Dari Dataset[prs]IranianCorpusTXTpartner
Kannada Time-Aligned Speech Corpus[kan]DravidianASROGG,SRT355.77 MBCC-BY-NC-SA-4.0
Tamil Time-Aligned Speech Dataset[tam]DravidianASROGG,SRT37.11 MBCC-BY-NC-SA-4.0
Multispeaker Hindi ASR Dataset[hin]Indo-AryanASROGG,SRT63.03 MBCC-BY-NC-SA-4.0
Punjabi 10 Hours TTS[pnb]Indo-AryanTTSWEBM,TSV481.96 MBCC-BY-NC-SA-4.0
Saraiki 10 Hours TTS Dataset[srk]Indo-AryanTTSWEBM,TSV584.44 MBCC-BY-NC-SA-4.0
Malayalam Time-Aligned Speech Corpus[mal]DravidianASRSRT5 speakerspartner
Urdu 10 Hours TTS[urd]Indo-AryanTTSWEBM,TSV10 hrspartner
Pakistan Traffic Signs Dataset[urd]Indo-AryanVisionJPEG,JSON2.64 GBCC-BY-NC-4.0
Bangladesh Traffic Signs Dataset[ben]Indo-AryanVisionJSON,JPEG1.35 GBCC-BY-NC-4.0
India Traffic Signs Dataset[multi]MultipleVisionJPEG,JSON190.44 MBCC-BY-ND-4.0
Hazargi Corpus for Speech Recognition[haz]IranianASRMP3,TSV511.56 MBpartner
Hazargi–English Speech Translation Corpus[haz]IranianMTMP3,TSV511.58 MBpartner
Hazargi Multimodal Dataset[haz]IranianMultimodalMP4,EAF44.00 GBpartner

published  ·  in development  ·  licences shown where MirasAI is the listed steward.

The languages — each in its own script

29 languages, 5 families.

Indo-Aryan, Iranian, Dravidian, Dardic, and Indo-European contact languages — many with tens of millions of speakers and almost no machine-readable data before this catalogue.

سرائیکیSaraiki[srk] · 6 datasets · Indo-Aryan
اردوUrdu[urd] · 6 datasets · Indo-Aryan
هزارگیHazargi[haz] · 5 datasets · Iranian
বাংলাBengali[ben] · 4 datasets · Indo-Aryan
پنجابیPunjabi[pnb] · 4 datasets · Indo-Aryan
سنڌيSindhi[snd] · 4 datasets · Indo-Aryan
हिन्दीHindi[hin] · 3 datasets · Indo-Aryan
براہوئیBrahui[brh] · 3 datasets · Dravidian
தமிழ்Tamil[tam] · 2 datasets · Dravidian
ಕನ್ನಡKannada[kan] · 2 datasets · Dravidian
کھوارKhowar[khw] · 2 datasets · Dardic
بلوچیE. Balochi[bgp] · 2 datasets · Iranian
توروالیTorwali[trw] · 2 datasets · Dardic
دریDari[prs] · 2 datasets · Iranian
فارسیPersian[pes] · 2 datasets · Iranian
తెలుగుTelugu[tel] · 1 dataset · Dravidian
ગુજરાતીGujarati[guj] · 1 dataset · Indo-Aryan
मराठीMarathi[mar] · 1 dataset · Indo-Aryan
চাটগাঁইয়াChittagonian[ctg] · 1 dataset · Indo-Aryan
নোয়াখাইল্লাNoakhalian[oak] · 1 dataset · Indo-Aryan
রংপুরীRangpuri[rkt] · 1 dataset · Indo-Aryan
রোহিঙ্গাRohingya[rhg] · 1 dataset · Indo-Aryan
ছিলটীSylheti[syl] · 1 dataset · Indo-Aryan
گوجریGojri[gju] · 1 dataset · Indo-Aryan
شیناKohistani Shina[plk] · 1 dataset · Dardic
بلوچیW. Balochi[bgn] · 1 dataset · Iranian
گاوریGawri[gwc] · 1 dataset · Dardic
മലയാളംMalayalam[mal] · 1 dataset · Dravidian
Multi-language[multi] · 1 dataset · Multiple
Flagship programmes · 1 published, 2 funded & in development

Building the language infrastructure South Asian AI needs next.

Two funded data initiatives are expanding what researchers, product teams, and communities can build for underrepresented languages — building on a completed, peer-reviewed corpus covering 39 languages of Pakistan.

Published · LREC 2026

Common Voice Pakistan: an open speech corpus for 39 languages

A one-year, Mozilla-funded effort spanning Pakistan's Indo-Aryan, Iranian, Dardic, Turkic, and isolate language families — from Balochistan and Sindh up through Gilgit-Baltistan and Kashmir. Locally authored texts, everyday speech, poetry, and folk song, collected with native speakers and community organizations rather than by top-down design.

BaltiBurushaskiKhowarTorwali KalashaShinaWakhiOrmuri GojriPalulaHazargiBrahui +27 more

Alam, M. & Tyers, F. M. (2026). Common Voice for Pakistan: Developing an Open Speech Corpus for Low-Resource Pakistani Languages. Proc. LREC 2026, pp. 3355–3359. ELRA. · Released CC0.

532.6hours recorded
493.6hours validated
Access the published corpus on Mozilla Data Collective →
Funded · In development

500-hour Urdu Text-to-Speech dataset

One of the largest planned Urdu TTS efforts of its kind, designed to support natural voice technology, accessibility tools, and speech applications. High-quality open Urdu speech data remains scarce; this closes that gap with structured, usable data.

UrduTTSOpen data
500hours planned
230M+Urdu speakers
Collaborate on Urdu speech data →
Funded · In development

60-hour multimodal dataset, six Pakistani languages

AI-ready video, audio, transcription, and English translation for languages severely underrepresented in today's models — supporting ASR, machine translation, and conversational AI.

HazargiShinaSaraikiPunjabiPashtoKhowar
60hours planned
6languages
Discuss the multimodal project →
Services

Technology services from strategy through delivery.

Consulting

IT consulting & delivery

Technology strategy, system planning, implementation support, and dependable delivery for organizations and growing teams.

Applied AI

AI & data systems

Applied AI, automation, data pipelines, custom models, and evaluation workflows designed around real operational needs.

Product

Digital platforms

Web products, internal tools, data experiences, and technology infrastructure built for usability and long-term value.

Language tech

Language technology

Text, speech, translation, ASR, TTS, NLP, and multimodal resources for South Asian and low-resource languages.

Field ops

Campaign support in Pakistan

Technology, digital coordination, localized execution, and field-facing support for international campaigns and activities.

Research

Research & data partnerships

Ethical collection, curation, annotation, quality assurance, and delivery for research institutions and technology teams.

Common questions

Licensing, access & commissioning.

What languages does MirasAI have datasets for?
MirasAI has published data for 29 languages across five families — Indo-Aryan (Urdu, Punjabi, Saraiki, Sindhi, Bengali, Hindi, Rohingya, Sylheti and others), Iranian (Balochi, Dari, Persian, Hazargi), Dardic (Khowar, Torwali, Gawri, Shina), Dravidian (Tamil, Telugu, Kannada, Malayalam, Brahui) and multi-language sets. Across the wider body of work, including Common Voice Pakistan, the footprint covers more than 60 South Asian languages.
Can I commission a custom speech or text dataset?
Yes. MirasAI designs and delivers commissioned datasets — ASR corpora, text-to-speech data, parallel translation corpora, lexicons and multimodal video sets — for specific languages, dialects, domains and quality targets. Engagements typically cover collection, transcription, annotation, quality assurance and delivery in your required format.
What licence are MirasAI datasets released under?
Licensing varies by dataset, so check the licence shown on each dataset's page on Mozilla Data Collective before use. Openly published datasets are released for research and non-commercial use, most commonly under Creative Commons NonCommercial (NC) terms, with the Common Voice Pakistan corpus released under CC0. Compensated datasets, and any dataset licensed commercially, are governed by the MirasAI Data Licence Agreement.
Where can I download the open datasets?
The catalogue is hosted on Mozilla Data Collective, where the datasets are browsable and downloadable — some openly licensed, some compensated. Each row in the table on this page links straight through to Mozilla Data Collective. An MDC account is needed to download.
Can I use MirasAI datasets to train commercial AI models?
Not under the open terms. MirasAI's openly published datasets are for research and non-commercial use, and several carry Creative Commons NonCommercial (NC) licences that expressly exclude commercial use. Commercial use is available through the MirasAI Data Licence Agreement, which covers compensated datasets and datasets licensed directly from us. To train, fine-tune or deploy models commercially, contact meesum@mirasai.net and we will arrange access under those terms.
What formats are datasets delivered in?
Text corpora are delivered as TXT or DOCX. Speech datasets use OGG, MP3, WEBM or WAV with TSV or SRT time-aligned transcriptions. Vision datasets use JPEG with JSON annotations, and multimodal sets use MP4 with EAF annotation files. Custom formats can be specified for commissioned work.
Start a project

Need a technology partner who understands the region?

Talk to us about commissioned datasets, AI and data projects, digital products, language technology, or campaign support in Pakistan.