datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikipedia_20240401_10-languages_bge-m3_qdrant_indexThis repository contains a Qdrant index created from preprocessed and chunked Wikipedia HTML dumps from 10 languages. The embedding model used is BAAI/bge-m3
This index is compatible with WikiChat v2.0.
Refer to the following for more information:
GitHub repository: https://github.com/stanford-oval/WikiChat
Papers:
WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on Wikipedia
SPAGHETTI: Open-Domain Question Answering from Heterogeneous Data Sources… See the full description on the dataset page: https://huggingface.co/datasets/stanford-oval/wikipedia_20240401_10-languages_bge-m3_qdrant_index.low_resource_languages_pretrain_datawikipedia_20240801_10-languages_bge-m3_qdrant_indexThis repository contains a Qdrant index created from preprocessed and chunked Wikipedia HTML dumps from 10 languages. The embedding model used is BAAI/bge-m3
This index is compatible with WikiChat v2.0.
Refer to the following for more information:
GitHub repository: https://github.com/stanford-oval/WikiChat
Papers:
WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on Wikipedia
SPAGHETTI: Open-Domain Question Answering from Heterogeneous Data Sources… See the full description on the dataset page: https://huggingface.co/datasets/stanford-oval/wikipedia_20240801_10-languages_bge-m3_qdrant_index.low_resource_languages_pretrain_data5mixed_multilingual_commonvoice_all_languages_100kBuild from mozilla commonvoice 13 using the script commited in this repo.
Used to teach a model to ignore languages that are not french
Marco-Bench-MIF-languagesafrican_languages_translationBiasShadesInterested in contributing? Speak a language not represented here? Disagree with an annotation? Please submit feedback in the Community tab!
Dataset Card for BiasShades
Note: This dataset may NOT be used as training data in any form (pre-training, fine-tuning, post-training, etc.) without express permission from creators.
Dataset Details
Version: 1.0
License: SHADES 1 Montreal Data License
Dataset Description
728 stereotypes and associated… See the full description on the dataset page: https://huggingface.co/datasets/LanguageShades/BiasShades.african-languages-hplt-filtered
VelkroLM African Languages Corpus
This repository contains a filtered, provenance-preserving text corpus derived from the language-specific HPLT v3.0 shards discovered at hplt-project.org/datasets/v3.0. It is organized by language and source shard so researchers can load only the languages they need. The upstream HPLT project describes its v3.0 data as multilingual web-corpus material; the upstream release, source metadata, and terms remain authoritative.
This publication is… See the full description on the dataset page: https://huggingface.co/datasets/rufatronics/african-languages-hplt-filtered.low_resource_languages_pretrainlow_resource_languages_pretrain_data8african-languages-corpus
VelkroLM African Languages Corpus
This repository contains a filtered, provenance-preserving text corpus derived from the language-specific HPLT v3.0 shards discovered at hplt-project.org/datasets/v3.0. It is organized by language and source shard so researchers can load only the languages they need. The upstream HPLT project describes its v3.0 data as multilingual web-corpus material; the upstream release, source metadata, and terms remain authoritative.
This publication is… See the full description on the dataset page: https://huggingface.co/datasets/VelkroLM/african-languages-corpus.asl_sign_languages_alphabets_v03low_resource_languages_pretrain_data2fleurs-r-neucodec-all-languages
FLEURS-R NeuCodec All Languages
FLEURS-R metadata, source audio and precomputed NeuCodec speech tokens for 102
locales, plus a speaker label FLEURS itself does not ship.
Layout
data/{locale}-{split}.parquet — metadata, one row per utterance (this is what the
viewer shows).
audio/{locale}-{split}.zip — source FLEURS-R audio, 24kHz mono PCM16 WAV, members
named audio/{locale}/{split}/{id}.wav (the path column).
neucodec/{locale}-{split}-rank{N}.zip — NeuCodec… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/fleurs-r-neucodec-all-languages.low_resource_languages_pretrain_data4mixed_multilingual_commonvoice_all_languagessa-languages-translation
South African Languages Translation Pairs
Parallel sentence pairs between all 10 non-English official South African languages and English, sourced from JW.org, the South African government magazine Vuk'uzenzele (GCIS/DSFSI), and the Autshumato Translation Project (North-West University).
Schema
Column
Description
lang
ISO code of the source language
text
Sentence in the source language
en_text
English translation
doc_id
Article ID (same across… See the full description on the dataset page: https://huggingface.co/datasets/Mehgoss/sa-languages-translation.african-languages-speech
WAXAL African-language speech — filtered wave 1
This repository contains bounded, filtered WAXAL audio/transcript pairs for Hausa (hau_tts), Yoruba (yor_tts), and Igbo (ibo_tts). Each configuration is split into tar archives containing an audio file plus a JSON record with transcript, language, speaker, source ID, and provenance.
The upstream source is google/WaxalNLP and the WAXAL paper is arXiv:2602.02734. Consult the upstream dataset card for the exact component license and… See the full description on the dataset page: https://huggingface.co/datasets/VelkroLM/african-languages-speech.ethiopian-languages-speech-dataset
Leyu Ethiopian Languages Speech Dataset
Audio recordings paired with corresponding text transcripts, collected on the Leyu Data Collection Platform — an open-source platform for crowdsourced speech data collection — for the Leyu Platform Competition, covering 4 languages: Amharic, Afaan Oromo, Sidama, Tigrinya.
Dataset Summary
Languages: Amharic (am), Afaan Oromo (om), Sidama (sid), Tigrinya (ti)
Total examples: 2750
License: CC-BY-4.0
Task categories: Automatic… See the full description on the dataset page: https://huggingface.co/datasets/Benji-fish/ethiopian-languages-speech-dataset.benji-ethiopian-languages-speech-dataset
Leyu Ethiopian Languages Speech Dataset
Audio recordings paired with corresponding text transcripts, collected on the Leyu Data Collection Platform — an open-source platform for crowdsourced speech data collection — for the Leyu Platform Competition, covering 4 languages: Amharic, Afaan Oromo, Sidama, Tigrinya.
Dataset Summary
Languages: Amharic (am), Afaan Oromo (om), Sidama (sid), Tigrinya (ti)
Total examples: 2750
License: CC-BY-4.0
Task categories: Automatic… See the full description on the dataset page: https://huggingface.co/datasets/LeyuCompetition/benji-ethiopian-languages-speech-dataset.gsm8k-translated
Multilingual GSM8K Translations
This dataset contains machine-translated versions of GSM8K in these languages:
French (fr)
German (de)
Hindi (hi)
Dataset Structure
For each language, we provide the original GSM8K train and test splits:
train: 7,473 samples
test: 1,319 samples
Each sample consists of a question and an answer.
The question describes a grade-school-level math word problem that requires multi-step mathematical reasoning. The answer contains a… See the full description on the dataset page: https://huggingface.co/datasets/math-across-languages/gsm8k-translated.indic-languages-speech
Indic Spontaneous Speech — Silencio
Spontaneous speech in 7 Indic languages from 161 speakers, one clip each. Hindi, Urdu, Bengali, Marathi, Nepali, Sindhi and Gujarati, recorded by speakers born in India, Pakistan, Bangladesh, Nepal and the diaspora, and labelled with 34 self-reported regional varieties. Audio and speaker metadata only. Human-validated transcription is available on request.
Hours
2.00
Clips
161
Speakers
161 (one clip each)
Languages
7… See the full description on the dataset page: https://huggingface.co/datasets/SilencioNetwork/indic-languages-speech.speech-recognition-congolese-languages
Speech Recognition Datasets for Congolese Languages
Dataset Details
Dataset Description
This dataset contains two new benchmark corpora designed for low-resource languages spoken in the Democratic Republic of the Congo: The Lingala Read Speech Corpus LRSC, with 4.3 hours of labelled audio, and the Congolese Speech Radio Corpus CSRC, which offers 741 hours of unlabeled audio spanning four significant low-resource languages of the region (Lingala, Tshiluba… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/speech-recognition-congolese-languages.TTS-For-Malian-Languages-Datasetworld-languages
World Languages: 7,130 Languages with Coordinates and Features
Where are the world's 7,000+ languages spoken, and what makes each one unique? This dataset provides geographic coordinates, linguistic features, and demographic data for every documented living language.
World Languages integrates three authoritative sources:
• Glottolog: The definitive catalog of the world's languages, dialects, and language families with ISO 639-3 codes and Glottocodes
• WALS (World Atlas of Language… See the full description on the dataset page: https://huggingface.co/datasets/lukeslp/world-languages.paloma_programming_languagesscam_ham_india_14_languages
Indian Multilingual Scam & Ham SMS Dataset (14 Languages)
A balanced benchmark dataset of 14,000 text samples curated for detecting fraud, phishing, and legitimate messages across 14 major Indian languages and scripts.
Dataset Summary
Total Records: 14,000 samples
Label Split: Exactly 7,000 Scam / 7,000 Ham (Balanced 50:50 distribution)
Languages Covered (1,000 samples each across 14 languages):
Assamese (as)
Bengali (bn)
English (en)
Gujarati (gu)
Hindi (hi)… See the full description on the dataset page: https://huggingface.co/datasets/anmolshrivastav/scam_ham_india_14_languages.ghanaian_languages_to_english_translation_and_transcription_datasetASR-For-Malian-Languages-Dataset
