Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ylacombe /english_dialects Dataset Card for "english_dialects" Dataset Summary This dataset consists of 31 hours of transcribed high-quality audio of English sentences recorded by 120 volunteers speaking with different accents of the British Isles. The dataset is intended for linguistic analysis as well as use for speech technologies. The speakers self-identified as native speakers of Southern England, Midlands, Northern England, Welsh, Scottish and Irish varieties of English. The recording scripts… See the full description on the dataset page: https://huggingface.co/datasets/ylacombe/english_dialects.audiotext-to-speech10K<n<100K38 likes1.6k downloads3y agoHugging Face02malaysia-ai /malaysian-dialects-youtube Malaysian dialects Youtube Entire videos from https://www.youtube.com using 'malay dialects' keyword. With total 398634 audio files, total 68607.6 hours. how to download huggingface-cli download --repo-type dataset \ --include '*.z*' \ --local-dir './' \ malaysia-ai/malaysian-dialects-youtube https://gist.githubusercontent.com/huseinzol05/2e26de4f3b29d99e993b349864ab6c10/raw/9b2251f3ff958770215d70c8d82d311f82791b78/unzip.py python3 unzip.py Source code Source… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/malaysian-dialects-youtube.0 likes1.5k downloads1y agoHugging Face03TigreGotico /arabic-dialects-gold20 arabic-dialects-gold20 660 sentences: 33 Arabic lects × 20 sentences, each with fully diacritized dialectal orthography, the undiacritized surface form, gold IPA, an engine draft, an English gloss, machine-verified phonetic feature tags, per-row verification metadata, and notes citing the dialectological literature that grounds the row. Columns (TSV, UTF-8, one file per lect): id, sentence, raw, ipa, ipa_o2i, gloss_en, features, notes, fable_corrections, verification… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-dialects-gold20.texttext-to-speechn<1K0 likes1.4k downloads3mo agoHugging Face04malaysia-ai /pseudolabel-dialects-youtube-whisper-large-v3 malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3 Pseudolabel malaysia-ai/malaysian-dialects-youtube using openai/whisper-large-v3 How to prepare the dataset huggingface-cli download --repo-type dataset \ --include '*.zip' \ --local-dir './' \ --max-workers 20 \ malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3 wget https://gist.githubusercontent.com/huseinzol05/2e26de4f3b29d99e993b349864ab6c10/raw/9b2251f3ff958770215d70c8d82d311f82791b78/unzip.py python3… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/pseudolabel-dialects-youtube-whisper-large-v3.audio1M<n<10M1 likes930 downloads1y agoHugging Face05ISLAM-PO /arab-dialects-20-countries-3m Dataset evaluation: See EVALUATION.md for schema checks, indexing status, and quality limitations. Viewer note: default is a lightweight preview; select full to load the complete corpus. Current Hub Validation Status Repository claim: 3,000,000 records Dataset Server indexed rows: 1,183,361 Dataset Server estimate: 2,064,964 The 3M target figure is a raw-repository claim and is not yet fully verified by the Hub index. Validate the JSONL files before publishing a definitive… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arab-dialects-20-countries-3m.texttext-generation1M<n<10M0 likes849 downloads16d agoHugging Face06TigreGotico /arabic-dialects-gold20-code-switch gold20-code-switch Code-switched Arabic sentences with IPA: 20 rows per lect across 33 Arabic lects (the same roster as the sibling TigreGotico/arabic-dialects-gold20). Each row embeds foreign material in a dialectal Arabic frame: inline Latin-script English (and French, for the lects whose live contact language is French), Arabic-script loanwords (سيرفس، كاش، موبايل-class), and Arabizi (Latin-written Arabic with digit gutturals). Columns (TSV, UTF-8, one file per lect): id… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-dialects-gold20-code-switch.texttext-to-speechn<1K0 likes693 downloads3mo agoHugging Face07rjnieto /french-dialectsaudio1K<n<10K0 likes314 downloads2y agoHugging Face08KlangAI /klang-dialects Klang Dialects Klang Dialects is an open benchmark for Swedish speech recognition, created by the Klang Research Team from recordings contributed through Knäck Klang. The Swedish benchmark contains 1,804 recordings, 656 speaker IDs, and 5.15 hours of speech in its main configuration, sv-clean. The benchmark supports research on regional variation in Swedish speech recognition. We plan to expand the dataset with more sentences, speakers, and languages. Configurations… See the full description on the dataset page: https://huggingface.co/datasets/KlangAI/klang-dialects.audioautomatic-speech-recognition1K<n<10K3 likes273 downloads25d agoHugging Face09amgadhasan /arabic_tweets_dialectstexttext-classification100K<n<1M0 likes252 downloads2y agoHugging Face10MahmoudIbrahim /ar-3-dialects-speechgatedaudio10K<n<100K0 likes134 downloads13h agoHugging Face11aminedjebbie /Multi-Arabic-dialectstext10K<n<100K1 likes128 downloads5y agoHugging Face12badrex /MADIS5-spoken-arabic-dialects Dataset Overview MADIS-5 (Multi-domain Arabic Dialect Identification in Speech) is a manually curated dataset designed to facilitate evaluation of cross-domain robustness of Arabic Dialect Identification (ADI) systems. This dataset provides a comprehensive benchmark for testing out-of-domain generalization across different speech domains with diverse recording conditions and speaking styles. Dataset Statistics Total Duration: ~12 hours of speech Total… See the full description on the dataset page: https://huggingface.co/datasets/badrex/MADIS5-spoken-arabic-dialects.audioaudio-classification1K<n<10K0 likes121 downloads1y agoHugging Face13drelhaj /Arabic-Dialects Arabic Dialects Dataset (Bivalency & Code-Switching) The Arabic Dialects Dataset is a specialised corpus designed for automatic dialect identification, with a focus on the linguistic phenomena of bivalency and written code-switching between major Arabic dialects and Modern Standard Arabic (MSA).It covers five varieties: EGY – Egyptian Arabic GLF – Gulf Arabic LAV – Levantine Arabic NOR – North African / Tunisian Arabic MSA – Modern Standard Arabic The dataset was created… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/Arabic-Dialects.texttext-classification10K<n<100K4 likes113 downloads11mo agoHugging Face14TigreGotico /portuguese-dialects-ipa-synthetic portuguese-dialects-ipa-synthetic 920 dialect-register sentences — 20 per lect across 46 lects: 41 Portuguese varieties (European regional, insular, Brazilian regional, African/Asian/border national norms, medieval stages) plus the other languages of Portugal and their kin: Mirandese (3 lects), Asturleonese of Portugal (Rionorese, Guadramilese) with Barranquenho, and Galician-Portuguese. Each row carries two IPA columns with distinct provenance. Schema sentence —… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-dialects-ipa-synthetic.texttext-to-speechn<1K0 likes106 downloads3mo agoHugging Face15ISLAM-PO /arabic-history-and-dialects Dataset evaluation: See EVALUATION.md for schema checks, quality limits, and the fact-check plan. Viewer note: default is a lightweight preview; select full to load the complete corpus. مجموعة البيانات العربية الشاملة للذكاء الاصطناعي 🇸🇦🇪🇬🇱🇧🇲🇦 [!WARNING] This is a small educational draft. The card flags three historical claims for fact-checking; verify them before using the dataset for factual QA or training. Arabic Multi-Dialect & Civilization… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arabic-history-and-dialects.textquestion-answeringn<1K0 likes89 downloads16d agoHugging Face16vladsfa /ukr-dialects-audio-dataset Ukrainian Dialects Audio Dataset Merged Ukrainian dialect speech dataset combining 5 speaker datasets, with train/validation/test splits. Dataset Description This dataset contains audio recordings of Ukrainian dialect speech, merged from the following source datasets: NaUKMA-Audio-Dataset Ivanna-Stefiuk-Audio-Dataset Larysa-Irodenko-Audio-Dataset Hutsulendia-Audio-Dataset Dido-Yvanchyk-Audio-Dataset-v2 Dataset Structure train: 27,675 samples… See the full description on the dataset page: https://huggingface.co/datasets/vladsfa/ukr-dialects-audio-dataset.audioautomatic-speech-recognition10K<n<100K0 likes80 downloads1mo agoHugging Face17KSE-RESEARCH-Group /ukr-dialects-audio-dataset Ukrainian Dialects Audio Dataset Merged Ukrainian dialect speech dataset combining 5 speaker datasets, with train/validation/test splits. Dataset Description This dataset contains audio recordings of Ukrainian dialect speech, merged from the following source datasets: NaUKMA-Audio-Dataset Ivanna-Stefiuk-Audio-Dataset Larysa-Irodenko-Audio-Dataset Hutsulendia-Audio-Dataset Dido-Yvanchyk-Audio-Dataset-v2 Dataset Structure train: 27,675 samples validation: 3… See the full description on the dataset page: https://huggingface.co/datasets/KSE-RESEARCH-Group/ukr-dialects-audio-dataset.audioautomatic-speech-recognition10K<n<100K1 likes77 downloads7mo agoHugging Face18arbml /Arabic_Dialects_Datasettext1K<n<10K3 likes76 downloads2y agoHugging Face19MohamedGomaa30 /casablanca-NADI-all-dialectsaudio10K<n<100K0 likes71 downloads4mo agoHugging Face20lahga /arabic-dialects Lahga: a dictionary of Arabic dialects لهجة: معجم اللهجات العربية. معجم تشاركي مفتوح، كل صفحة فيه تبدأ من معنى بالعربية الفصحى، وتحته الأشكال التي يُقال بها في اللهجات، كل شكل منسوب إلى لهجته. هذه نسخة من بيانات الموقع lahga.fyi بتاريخ 2026-10-02. Lahga is an open, collaborative dictionary of spoken Arabic. Every headword is a meaning stated in Modern Standard Arabic: a word, a short phrase or a proverb. Under it are the forms the dialects use for that meaning, each… See the full description on the dataset page: https://huggingface.co/datasets/lahga/arabic-dialects.tabulartranslation10K<n<100K0 likes60 downloads9d agoHugging Face21CNTXTAI0 /arabic_dialects_question_and_answerData Content The file provided: Q/A Reasoning dataset contains the following columns: ID # : Denotes the reference ID for: a. Question b. Answer to the question c. Hint d. Reasoning e. Word count for items a to d above Dialects: Contains the following dialects in separate columns: a. English b. MSA c. Emirati d. Egyptian e. Levantine Syria f. Levantine Jordan g. Levantine Palestine h. Levantine Lebanon Data Generation Process The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.tabularquestion-answeringn<1K6 likes57 downloads2y agoHugging Face22AylinNaebzadeh /english_dialects_combinedThis dataset is the combined version of 🤗 ylacombe/english_dialects published at LREC 2020. The goal is to prepare the dataset for audio models training on different tasks. In addition to mergin 11 datasets together, gender and accent were added for each record as metadata. @inproceedings{demirsahin-etal-2020-open, title = "Open-source Multi-speaker Corpora of the {E}nglish Accents in the {B}ritish Isles", author = "Demirsahin, Isin and Kjartansson, Oddur and Gutkin… See the full description on the dataset page: https://huggingface.co/datasets/AylinNaebzadeh/english_dialects_combined.audio10K<n<100K1 likes48 downloads5mo agoHugging Face23fatymahaly /Arabic_Dialects Dataset Card for Arabic Dialects Dataset Summary The Arabic Dialects dataset is a collection of text samples representing multiple spoken Arabic dialects alongside Modern Standard Arabic (MSA). It is designed to help train and evaluate natural language processing (NLP) models on dialect identification, text classification, and understanding regional linguistic variations. Languages and Dialects Included Egyptian (EGY) Gulf (GLF) Levantine (LEV)… See the full description on the dataset page: https://huggingface.co/datasets/fatymahaly/Arabic_Dialects.texttext-classificationn<1K1 likes48 downloads1mo agoHugging Face24statworx /swiss-dialects Dataset Card for ArchiMod Corpus Dataset Summary The ArchiMob corpus represents German linguistic varieties spoken within the territory of Switzerland. This corpus is the first electronic resource containing long samples of transcribed text in Swiss German, intended for studying the spatial distribution of morphosyntactic features and for natural language processing. Languages Swiss-German Dataset Structure Data Instances { 'sentence':… See the full description on the dataset page: https://huggingface.co/datasets/statworx/swiss-dialects.texttext-generation1K<n<10K1 likes46 downloads4y agoHugging Face25gagan3012 /habibi_dialects_datatext10K<n<100K0 likes42 downloads2y agoHugging Face26RafatK /Ben_dialectsaudio1K<n<10K0 likes41 downloads5mo agoHugging Face27PRAli22 /Arabic_dialects_to_MSAtext100K<n<1M10 likes40 downloads3y agoHugging Face28abdelfetteh /arabic-speech-dialectsaudio10K<n<100K0 likes35 downloads4mo agoHugging Face29Geethuzzz /spoken-arabic-dialects-with-transcriptionsaudio1K<n<10K0 likes32 downloads1y agoHugging Face30malaysia-ai /filtered-malaysian-dialects-youtube Filtered Malaysian Dialects Youtube Originally from https://huggingface.co/datasets/malaysia-ai/malaysian-dialects-youtube, filtered using malay word dictionary. how to download huggingface-cli download --repo-type dataset \ --include '*.z*' \ --local-dir './' \ malaysia-ai/filtered-malaysian-dialects-youtube wget https://gist.githubusercontent.com/huseinzol05/2e26de4f3b29d99e993b349864ab6c10/raw/9b2251f3ff958770215d70c8d82d311f82791b78/unzip.py python3 unzip.py 1 likes31 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.