dialects
Datasets
All datasets matching “dialects”english_dialects
Dataset Card for "english_dialects"
Dataset Summary
This dataset consists of 31 hours of transcribed high-quality audio of English sentences recorded by 120 volunteers speaking with different accents of the British Isles. The dataset is intended for linguistic analysis as well as use for speech technologies. The speakers self-identified as native speakers of Southern England, Midlands, Northern England, Welsh, Scottish and Irish varieties of English.
The recording scripts… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/english_dialects.malaysian-dialects-youtube
Malaysian dialects Youtube
Entire videos from https://www.youtube.com using 'malay dialects' keyword.
With total 398634 audio files, total 68607.6 hours.
how to download
huggingface-cli download --repo-type dataset \
--include '*.z*' \
--local-dir './' \
malaysia-ai/malaysian-dialects-youtube
https://gist.githubusercontent.com/huseinzol05/2e26de4f3b29d99e993b349864ab6c10/raw/9b2251f3ff958770215d70c8d82d311f82791b78/unzip.py
python3 unzip.py
Source code
Source… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/malaysian-dialects-youtube.arabic-dialects-gold20
arabic-dialects-gold20
660 sentences: 33 Arabic lects × 20 sentences, each with fully diacritized
dialectal orthography, the undiacritized surface form, gold IPA, an engine
draft, an English gloss, machine-verified phonetic feature tags, per-row
verification metadata, and notes citing the dialectological literature that
grounds the row.
Columns (TSV, UTF-8, one file per lect):
id, sentence, raw, ipa, ipa_o2i, gloss_en, features, notes, fable_corrections, verification… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-dialects-gold20.pseudolabel-dialects-youtube-whisper-large-v3
malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3
Pseudolabel malaysia-ai/malaysian-dialects-youtube using openai/whisper-large-v3
How to prepare the dataset
huggingface-cli download --repo-type dataset \
--include '*.zip' \
--local-dir './' \
--max-workers 20 \
malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3
wget https://gist.githubusercontent.com/huseinzol05/2e26de4f3b29d99e993b349864ab6c10/raw/9b2251f3ff958770215d70c8d82d311f82791b78/unzip.py
python3… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3.arab-dialects-20-countries-3m
Dataset evaluation: See EVALUATION.md for schema checks, indexing status, and quality limitations.
Viewer note: default is a lightweight preview; select full to load the complete corpus.
Current Hub Validation Status
Repository claim: 3,000,000 records
Dataset Server indexed rows: 1,183,361
Dataset Server estimate: 2,064,964
The 3M target figure is a raw-repository claim and is not yet fully verified by the Hub index. Validate the JSONL files before publishing a definitive… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arab-dialects-20-countries-3m.arabic-dialects-gold20-code-switch
gold20-code-switch
Code-switched Arabic sentences with IPA: 20 rows per lect across 33 Arabic
lects (the same roster as the sibling TigreGotico/arabic-dialects-gold20).
Each row embeds foreign material in a dialectal Arabic frame: inline
Latin-script English (and French, for the lects whose live contact language
is French), Arabic-script loanwords (سيرفس، كاش، موبايل-class), and Arabizi
(Latin-written Arabic with digit gutturals).
Columns (TSV, UTF-8, one file per lect):
id… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-dialects-gold20-code-switch.
