Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01stanford-oval /wikipedia_20240401_10-languages_bge-m3_qdrant_indexThis repository contains a Qdrant index created from preprocessed and chunked Wikipedia HTML dumps from 10 languages. The embedding model used is BAAI/bge-m3 This index is compatible with WikiChat v2.0. Refer to the following for more information: GitHub repository: https://github.com/stanford-oval/WikiChat Papers: WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on Wikipedia SPAGHETTI: Open-Domain Question Answering from Heterogeneous Data Sources… See the full description on the dataset page: https://huggingface.co/datasets/stanford-oval/wikipedia_20240401_10-languages_bge-m3_qdrant_index.text-retrieval100M<n<1B0 likes5k downloads2y agoHugging Face02Aletheia-ng /low_resource_languages_pretrain_datatext100M<n<1B0 likes2.2k downloads1y agoHugging Face03stanford-oval /wikipedia_20240801_10-languages_bge-m3_qdrant_indexThis repository contains a Qdrant index created from preprocessed and chunked Wikipedia HTML dumps from 10 languages. The embedding model used is BAAI/bge-m3 This index is compatible with WikiChat v2.0. Refer to the following for more information: GitHub repository: https://github.com/stanford-oval/WikiChat Papers: WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on Wikipedia SPAGHETTI: Open-Domain Question Answering from Heterogeneous Data Sources… See the full description on the dataset page: https://huggingface.co/datasets/stanford-oval/wikipedia_20240801_10-languages_bge-m3_qdrant_index.text-retrieval100M<n<1B3 likes1.2k downloads2y agoHugging Face04Aletheia-ng /low_resource_languages_pretrain_data5text100M<n<1B0 likes994 downloads1y agoHugging Face05BrunoHays /mixed_multilingual_commonvoice_all_languages_100kBuild from mozilla commonvoice 13 using the script commited in this repo. Used to teach a model to ignore languages that are not french audio100K<n<1M0 likes796 downloads2y agoHugging Face06jbross-ibm-research /Marco-Bench-MIF-languagestext10K<n<100K0 likes677 downloads7mo agoHugging Face07Aletheia-ng /african_languages_translationtext1M<n<10M1 likes675 downloads1y agoHugging Face08LanguageShades /BiasShadesgatedInterested in contributing? Speak a language not represented here? Disagree with an annotation? Please submit feedback in the Community tab! Dataset Card for BiasShades Note: This dataset may NOT be used as training data in any form (pre-training, fine-tuning, post-training, etc.) without express permission from creators. Dataset Details Version: 1.0 License: SHADES 1 Montreal Data License Dataset Description 728 stereotypes and associated… See the full description on the dataset page: https://huggingface.co/datasets/LanguageShades/BiasShades.imagetext-classificationn<1K26 likes666 downloads3mo agoHugging Face09rufatronics /african-languages-hplt-filtered VelkroLM African Languages Corpus This repository contains a filtered, provenance-preserving text corpus derived from the language-specific HPLT v3.0 shards discovered at hplt-project.org/datasets/v3.0. It is organized by language and source shard so researchers can load only the languages they need. The upstream HPLT project describes its v3.0 data as multilingual web-corpus material; the upstream release, source metadata, and terms remain authoritative. This publication is… See the full description on the dataset page: https://huggingface.co/datasets/rufatronics/african-languages-hplt-filtered.text1M<n<10M0 likes484 downloads2mo agoHugging Face10Aletheia-ng /low_resource_languages_pretraintext100M<n<1B1 likes444 downloads1y agoHugging Face11BeardedMonster /low_resource_languages_pretrain_data8text100M<n<1B0 likes432 downloads10mo agoHugging Face12VelkroLM /african-languages-corpus VelkroLM African Languages Corpus This repository contains a filtered, provenance-preserving text corpus derived from the language-specific HPLT v3.0 shards discovered at hplt-project.org/datasets/v3.0. It is organized by language and source shard so researchers can load only the languages they need. The upstream HPLT project describes its v3.0 data as multilingual web-corpus material; the upstream release, source metadata, and terms remain authoritative. This publication is… See the full description on the dataset page: https://huggingface.co/datasets/VelkroLM/african-languages-corpus.text1M<n<10M0 likes351 downloads2mo agoHugging Face13Marxulia /asl_sign_languages_alphabets_v03imageimage-classification10K<n<100K9 likes320 downloads3y agoHugging Face14Aletheia-ng /low_resource_languages_pretrain_data2text100M<n<1B0 likes311 downloads1y agoHugging Face15malaysia-ai /fleurs-r-neucodec-all-languages FLEURS-R NeuCodec All Languages FLEURS-R metadata, source audio and precomputed NeuCodec speech tokens for 102 locales, plus a speaker label FLEURS itself does not ship. Layout data/{locale}-{split}.parquet — metadata, one row per utterance (this is what the viewer shows). audio/{locale}-{split}.zip — source FLEURS-R audio, 24kHz mono PCM16 WAV, members named audio/{locale}/{split}/{id}.wav (the path column). neucodec/{locale}-{split}-rank{N}.zip — NeuCodec… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/fleurs-r-neucodec-all-languages.audiotext-to-speech100K<n<1M4 likes296 downloads27d agoHugging Face16Aletheia-ng /low_resource_languages_pretrain_data4text100M<n<1B0 likes250 downloads1y agoHugging Face17BrunoHays /mixed_multilingual_commonvoice_all_languagesaudio10K<n<100K0 likes240 downloads2y agoHugging Face18Mehgoss /sa-languages-translation South African Languages Translation Pairs Parallel sentence pairs between all 10 non-English official South African languages and English, sourced from JW.org, the South African government magazine Vuk'uzenzele (GCIS/DSFSI), and the Autshumato Translation Project (North-West University). Schema Column Description lang ISO code of the source language text Sentence in the source language en_text English translation doc_id Article ID (same across… See the full description on the dataset page: https://huggingface.co/datasets/Mehgoss/sa-languages-translation.texttranslation1M<n<10M0 likes199 downloads4mo agoHugging Face19VelkroLM /african-languages-speech WAXAL African-language speech — filtered wave 1 This repository contains bounded, filtered WAXAL audio/transcript pairs for Hausa (hau_tts), Yoruba (yor_tts), and Igbo (ibo_tts). Each configuration is split into tar archives containing an audio file plus a JSON record with transcript, language, speaker, source ID, and provenance. The upstream source is google/WaxalNLP and the WAXAL paper is arXiv:2602.02734. Consult the upstream dataset card for the exact component license and… See the full description on the dataset page: https://huggingface.co/datasets/VelkroLM/african-languages-speech.0 likes178 downloads2mo agoHugging Face20Benji-fish /ethiopian-languages-speech-dataset Leyu Ethiopian Languages Speech Dataset Audio recordings paired with corresponding text transcripts, collected on the Leyu Data Collection Platform — an open-source platform for crowdsourced speech data collection — for the Leyu Platform Competition, covering 4 languages: Amharic, Afaan Oromo, Sidama, Tigrinya. Dataset Summary Languages: Amharic (am), Afaan Oromo (om), Sidama (sid), Tigrinya (ti) Total examples: 2750 License: CC-BY-4.0 Task categories: Automatic… See the full description on the dataset page: https://huggingface.co/datasets/Benji-fish/ethiopian-languages-speech-dataset.audioautomatic-speech-recognition1K<n<10K0 likes176 downloads2mo agoHugging Face21LeyuCompetition /benji-ethiopian-languages-speech-dataset Leyu Ethiopian Languages Speech Dataset Audio recordings paired with corresponding text transcripts, collected on the Leyu Data Collection Platform — an open-source platform for crowdsourced speech data collection — for the Leyu Platform Competition, covering 4 languages: Amharic, Afaan Oromo, Sidama, Tigrinya. Dataset Summary Languages: Amharic (am), Afaan Oromo (om), Sidama (sid), Tigrinya (ti) Total examples: 2750 License: CC-BY-4.0 Task categories: Automatic… See the full description on the dataset page: https://huggingface.co/datasets/LeyuCompetition/benji-ethiopian-languages-speech-dataset.audioautomatic-speech-recognition1K<n<10K0 likes160 downloads1mo agoHugging Face22math-across-languages /gsm8k-translated Multilingual GSM8K Translations This dataset contains machine-translated versions of GSM8K in these languages: French (fr) German (de) Hindi (hi) Dataset Structure For each language, we provide the original GSM8K train and test splits: train: 7,473 samples test: 1,319 samples Each sample consists of a question and an answer. The question describes a grade-school-level math word problem that requires multi-step mathematical reasoning. The answer contains a… See the full description on the dataset page: https://huggingface.co/datasets/math-across-languages/gsm8k-translated.textquestion-answering10K<n<100K0 likes123 downloads4mo agoHugging Face23SilencioNetwork /indic-languages-speech Indic Spontaneous Speech — Silencio Spontaneous speech in 7 Indic languages from 161 speakers, one clip each. Hindi, Urdu, Bengali, Marathi, Nepali, Sindhi and Gujarati, recorded by speakers born in India, Pakistan, Bangladesh, Nepal and the diaspora, and labelled with 34 self-reported regional varieties. Audio and speaker metadata only. Human-validated transcription is available on request. Hours 2.00 Clips 161 Speakers 161 (one clip each) Languages 7… See the full description on the dataset page: https://huggingface.co/datasets/SilencioNetwork/indic-languages-speech.audioaudio-classificationn<1K0 likes116 downloads18d agoHugging Face24Svngoku /speech-recognition-congolese-languages Speech Recognition Datasets for Congolese Languages Dataset Details Dataset Description This dataset contains two new benchmark corpora designed for low-resource languages spoken in the Democratic Republic of the Congo: The Lingala Read Speech Corpus LRSC, with 4.3 hours of labelled audio, and the Congolese Speech Radio Corpus CSRC, which offers 741 hours of unlabeled audio spanning four significant low-resource languages of the region (Lingala, Tshiluba… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/speech-recognition-congolese-languages.audioautomatic-speech-recognition1K<n<10K5 likes110 downloads2y agoHugging Face25Panga-Azazia /TTS-For-Malian-Languages-Datasetgatedaudio100K<n<1M0 likes109 downloads3d agoHugging Face26lukeslp /world-languages World Languages: 7,130 Languages with Coordinates and Features Where are the world's 7,000+ languages spoken, and what makes each one unique? This dataset provides geographic coordinates, linguistic features, and demographic data for every documented living language. World Languages integrates three authoritative sources: • Glottolog: The definitive catalog of the world's languages, dialects, and language families with ISO 639-3 codes and Glottocodes • WALS (World Atlas of Language… See the full description on the dataset page: https://huggingface.co/datasets/lukeslp/world-languages.geospatialfeature-extraction1K<n<10K3 likes106 downloads6mo agoHugging Face27WillHeld /paloma_programming_languagestext10K<n<100K0 likes105 downloads1y agoHugging Face28anmolshrivastav /scam_ham_india_14_languages Indian Multilingual Scam & Ham SMS Dataset (14 Languages) A balanced benchmark dataset of 14,000 text samples curated for detecting fraud, phishing, and legitimate messages across 14 major Indian languages and scripts. Dataset Summary Total Records: 14,000 samples Label Split: Exactly 7,000 Scam / 7,000 Ham (Balanced 50:50 distribution) Languages Covered (1,000 samples each across 14 languages): Assamese (as) Bengali (bn) English (en) Gujarati (gu) Hindi (hi)… See the full description on the dataset page: https://huggingface.co/datasets/anmolshrivastav/scam_ham_india_14_languages.texttext-classification10K<n<100K0 likes97 downloads12d agoHugging Face29theonlyamos /ghanaian_languages_to_english_translation_and_transcription_datasetaudion<1K0 likes91 downloads2y agoHugging Face30Panga-Azazia /ASR-For-Malian-Languages-Datasetgatedaudio100K<n<1M0 likes90 downloads3d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.