datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
english_dialects
Dataset Card for "english_dialects"
Dataset Summary
This dataset consists of 31 hours of transcribed high-quality audio of English sentences recorded by 120 volunteers speaking with different accents of the British Isles. The dataset is intended for linguistic analysis as well as use for speech technologies. The speakers self-identified as native speakers of Southern England, Midlands, Northern England, Welsh, Scottish and Irish varieties of English.
The recording scripts… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/english_dialects.malaysian-dialects-youtube
Malaysian dialects Youtube
Entire videos from https://www.youtube.com using 'malay dialects' keyword.
With total 398634 audio files, total 68607.6 hours.
how to download
huggingface-cli download --repo-type dataset \
--include '*.z*' \
--local-dir './' \
malaysia-ai/malaysian-dialects-youtube
https://gist.githubusercontent.com/huseinzol05/2e26de4f3b29d99e993b349864ab6c10/raw/9b2251f3ff958770215d70c8d82d311f82791b78/unzip.py
python3 unzip.py
Source code
Source… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/malaysian-dialects-youtube.arabic-dialects-gold20
arabic-dialects-gold20
660 sentences: 33 Arabic lects × 20 sentences, each with fully diacritized
dialectal orthography, the undiacritized surface form, gold IPA, an engine
draft, an English gloss, machine-verified phonetic feature tags, per-row
verification metadata, and notes citing the dialectological literature that
grounds the row.
Columns (TSV, UTF-8, one file per lect):
id, sentence, raw, ipa, ipa_o2i, gloss_en, features, notes, fable_corrections, verification… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-dialects-gold20.pseudolabel-dialects-youtube-whisper-large-v3
malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3
Pseudolabel malaysia-ai/malaysian-dialects-youtube using openai/whisper-large-v3
How to prepare the dataset
huggingface-cli download --repo-type dataset \
--include '*.zip' \
--local-dir './' \
--max-workers 20 \
malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3
wget https://gist.githubusercontent.com/huseinzol05/2e26de4f3b29d99e993b349864ab6c10/raw/9b2251f3ff958770215d70c8d82d311f82791b78/unzip.py
python3… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3.arab-dialects-20-countries-3m
Dataset evaluation: See EVALUATION.md for schema checks, indexing status, and quality limitations.
Viewer note: default is a lightweight preview; select full to load the complete corpus.
Current Hub Validation Status
Repository claim: 3,000,000 records
Dataset Server indexed rows: 1,183,361
Dataset Server estimate: 2,064,964
The 3M target figure is a raw-repository claim and is not yet fully verified by the Hub index. Validate the JSONL files before publishing a definitive… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arab-dialects-20-countries-3m.arabic-dialects-gold20-code-switch
gold20-code-switch
Code-switched Arabic sentences with IPA: 20 rows per lect across 33 Arabic
lects (the same roster as the sibling TigreGotico/arabic-dialects-gold20).
Each row embeds foreign material in a dialectal Arabic frame: inline
Latin-script English (and French, for the lects whose live contact language
is French), Arabic-script loanwords (سيرفس، كاش، موبايل-class), and Arabizi
(Latin-written Arabic with digit gutturals).
Columns (TSV, UTF-8, one file per lect):
id… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-dialects-gold20-code-switch.french-dialectsklang-dialects
Klang Dialects
Klang Dialects is an open benchmark for Swedish speech recognition, created by the Klang Research Team from recordings contributed through Knäck Klang.
The Swedish benchmark contains 1,804 recordings, 656 speaker IDs, and 5.15 hours of speech in its main configuration, sv-clean. The benchmark supports research on regional variation in Swedish speech recognition. We plan to expand the dataset with more sentences, speakers, and languages.
Configurations… See the full description on the dataset page: https://huggingface.co/datasets/KlangAI/klang-dialects.arabic_tweets_dialectsar-3-dialects-speechMulti-Arabic-dialectsMADIS5-spoken-arabic-dialects
Dataset Overview
MADIS-5 (Multi-domain Arabic Dialect Identification in Speech) is a manually curated dataset
designed to facilitate evaluation of cross-domain robustness of Arabic Dialect Identification (ADI) systems.
This dataset provides a comprehensive benchmark for testing out-of-domain generalization across different speech domains with diverse recording conditions and speaking styles.
Dataset Statistics
Total Duration: ~12 hours of speech
Total… See the full description on the dataset page: https://huggingface.co/datasets/badrex/MADIS5-spoken-arabic-dialects.Arabic-Dialects
Arabic Dialects Dataset (Bivalency & Code-Switching)
The Arabic Dialects Dataset is a specialised corpus designed for automatic dialect identification, with a focus on the linguistic phenomena of bivalency and written code-switching between major Arabic dialects and Modern Standard Arabic (MSA).It covers five varieties:
EGY – Egyptian Arabic
GLF – Gulf Arabic
LAV – Levantine Arabic
NOR – North African / Tunisian Arabic
MSA – Modern Standard Arabic
The dataset was created… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/Arabic-Dialects.portuguese-dialects-ipa-synthetic
portuguese-dialects-ipa-synthetic
920 dialect-register sentences — 20 per lect across 46 lects: 41 Portuguese varieties
(European regional, insular, Brazilian regional, African/Asian/border national norms,
medieval stages) plus the other languages of Portugal and their kin: Mirandese (3 lects),
Asturleonese of Portugal (Rionorese, Guadramilese) with Barranquenho,
and Galician-Portuguese. Each row carries two IPA columns with distinct provenance.
Schema
sentence —… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-dialects-ipa-synthetic.arabic-history-and-dialects
Dataset evaluation: See EVALUATION.md for schema checks, quality limits, and the fact-check plan.
Viewer note: default is a lightweight preview; select full to load the complete corpus.
مجموعة البيانات العربية الشاملة للذكاء الاصطناعي 🇸🇦🇪🇬🇱🇧🇲🇦
[!WARNING]
This is a small educational draft. The card flags three historical claims for fact-checking; verify them before using the dataset for factual QA or training.
Arabic Multi-Dialect & Civilization… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arabic-history-and-dialects.ukr-dialects-audio-dataset
Ukrainian Dialects Audio Dataset
Merged Ukrainian dialect speech dataset combining 5 speaker datasets, with train/validation/test splits.
Dataset Description
This dataset contains audio recordings of Ukrainian dialect speech, merged from the following source datasets:
NaUKMA-Audio-Dataset
Ivanna-Stefiuk-Audio-Dataset
Larysa-Irodenko-Audio-Dataset
Hutsulendia-Audio-Dataset
Dido-Yvanchyk-Audio-Dataset-v2
Dataset Structure
train: 27,675 samples… See the full description on the dataset page: https://huggingface.co/datasets/vladsfa/ukr-dialects-audio-dataset.ukr-dialects-audio-dataset
Ukrainian Dialects Audio Dataset
Merged Ukrainian dialect speech dataset combining 5 speaker datasets, with train/validation/test splits.
Dataset Description
This dataset contains audio recordings of Ukrainian dialect speech, merged from the following source datasets:
NaUKMA-Audio-Dataset
Ivanna-Stefiuk-Audio-Dataset
Larysa-Irodenko-Audio-Dataset
Hutsulendia-Audio-Dataset
Dido-Yvanchyk-Audio-Dataset-v2
Dataset Structure
train: 27,675 samples
validation: 3… See the full description on the dataset page: https://huggingface.co/datasets/KSE-RESEARCH-Group/ukr-dialects-audio-dataset.Arabic_Dialects_Datasetcasablanca-NADI-all-dialectsarabic-dialects
Lahga: a dictionary of Arabic dialects
لهجة: معجم اللهجات العربية. معجم تشاركي مفتوح، كل صفحة فيه تبدأ من معنى
بالعربية الفصحى، وتحته الأشكال التي يُقال بها في اللهجات، كل شكل منسوب إلى
لهجته. هذه نسخة من بيانات الموقع lahga.fyi بتاريخ 2026-10-02.
Lahga is an open, collaborative dictionary of spoken Arabic. Every
headword is a meaning stated in Modern Standard Arabic: a word, a short phrase
or a proverb. Under it are the forms the dialects use for that meaning, each… See the full description on the dataset page: https://huggingface.co/datasets/lahga/arabic-dialects.arabic_dialects_question_and_answerData Content
The file provided: Q/A Reasoning dataset
contains the following columns:
ID # : Denotes the reference ID for:
a. Question
b. Answer to the question
c. Hint
d. Reasoning
e. Word count for items a to d above
Dialects: Contains the following dialects in separate columns:
a. English
b. MSA
c. Emirati
d. Egyptian
e. Levantine Syria
f. Levantine Jordan
g. Levantine Palestine
h. Levantine Lebanon
Data Generation Process
The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.english_dialects_combinedThis dataset is the combined version of 🤗 ylacombe/english_dialects published at LREC 2020.
The goal is to prepare the dataset for audio models training on different tasks.
In addition to mergin 11 datasets together, gender and accent were added for each record as metadata.
@inproceedings{demirsahin-etal-2020-open,
title = "Open-source Multi-speaker Corpora of the {E}nglish Accents in the {B}ritish Isles",
author = "Demirsahin, Isin and
Kjartansson, Oddur and
Gutkin… See the full description on the dataset page: https://huggingface.co/datasets/AylinNaebzadeh/english_dialects_combined.Arabic_Dialects
Dataset Card for Arabic Dialects
Dataset Summary
The Arabic Dialects dataset is a collection of text samples representing multiple spoken Arabic dialects alongside Modern Standard Arabic (MSA). It is designed to help train and evaluate natural language processing (NLP) models on dialect identification, text classification, and understanding regional linguistic variations.
Languages and Dialects Included
Egyptian (EGY)
Gulf (GLF)
Levantine (LEV)… See the full description on the dataset page: https://huggingface.co/datasets/fatymahaly/Arabic_Dialects.swiss-dialects
Dataset Card for ArchiMod Corpus
Dataset Summary
The ArchiMob corpus represents German linguistic varieties spoken within the territory of Switzerland. This corpus is the first electronic resource containing long samples of transcribed text in Swiss German, intended for studying the spatial distribution of morphosyntactic features and for natural language processing.
Languages
Swiss-German
Dataset Structure
Data Instances
{ 'sentence':… See the full description on the dataset page: https://huggingface.co/datasets/statworx/swiss-dialects.habibi_dialects_dataBen_dialectsArabic_dialects_to_MSAarabic-speech-dialectsspoken-arabic-dialects-with-transcriptionsfiltered-malaysian-dialects-youtube
Filtered Malaysian Dialects Youtube
Originally from https://huggingface.co/datasets/malaysia-ai/malaysian-dialects-youtube, filtered using malay word dictionary.
how to download
huggingface-cli download --repo-type dataset \
--include '*.z*' \
--local-dir './' \
malaysia-ai/filtered-malaysian-dialects-youtube
wget https://gist.githubusercontent.com/huseinzol05/2e26de4f3b29d99e993b349864ab6c10/raw/9b2251f3ff958770215d70c8d82d311f82791b78/unzip.py
python3 unzip.py
