Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceH4 /Multilingual-Thinking Dataset summary Multilingual-Thinking is a reasoning dataset where the chain-of-thought has been translated from English into one of 4 languages: Spanish, French, Italian, and German. The dataset was created by sampling 1k training samples from the SystemChat subset of SmolTalk2 and translating the reasoning traces with another language model. This dataset was used in the OpenAI Cookbook to fine-tune the OpenAI gpt-oss models. You can load the dataset using: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/Multilingual-Thinking.texttext-generation1K<n<10K119 likes9k downloads1y agoHugging Face02baber /multilingual_mmluMMLU professionally translated into 14 languages using professional human translators, sourced from OpenAI's simple-eval. Original files: english: https://openaipublic.blob.core.windows.net/simple-evals/mmlu.csv multilingual: https://openaipublic.blob.core.windows.net/simple-evals/mmlu_{language}.csv where language one of "AR-XY", "BN-BD", "DE-DE", "ES-LA", "FR-FR", "HI-IN", "ID-ID", "IT-IT", "JA-JP", "KO-KR", "PT-BR", "ZH-CN", "SW-KE", "YO-NG", "EN-US" texttext-generation100K<n<1M1 likes5k downloads2y agoHugging Face03Multilingual-Multimodal-NLP /McEvalMcEval benchmark data as described in the McEval Paper. Code for the evaluation can be found on Github as McEval. texttext-generation10K<n<100K21 likes4.2k downloads2y agoHugging Face04ellamind /hle-multilingual HLE Multilingual Multilingual translations of HLE (Humanity's Last Exam), an expert-level QA benchmark with questions across math, science, humanities, and engineering designed to challenge even domain experts. Source: cais/hle (test split, 2,158 text-only questions out of 2,500 total) Languages Config Language Examples ces Czech 50 dan Danish 50 deu German 800 fin Finnish 50 fra French 50 ita Italian 50 nld Dutch 50 pol Polish 50 spa… See the full description on the dataset page: https://huggingface.co/datasets/ellamind/hle-multilingual.textquestion-answering1K<n<10K0 likes2.6k downloads7mo agoHugging Face05nvidia /Nemotron-SFT-Multilingual-v2 Dataset Description: Nemotron-SFT-Multilingual-v2 is a multilingual supervised fine-tuning (SFT) dataset for post-training text-generation models. It is generated by translating seed data from Nemotron-Math-v2, Nemotron-Competitive-Programming-v1, and Nemotron-Science-v1, adding multilingual coverage for Hindi (hi), Korean (ko), Brazilian Portuguese (pt-br), and refreshed Japanese (ja) data. The dataset is generated with a new data processing pipeline that avoids line-breaking… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Multilingual-v2.texttext-generation100K<n<1M14 likes2.5k downloads4mo agoHugging Face06CohereLabs /tiny-aya-l2-thinker-multilingual-reasoning Tiny Aya L2 Multilingual Reasoning (44 languages) Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker. Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English. Data source Prompts from AM-DeepSeek-R1-0528-Distilled Thinking traces and outputs distilled from gpt-oss-120b Translated with command-a-translate and DeepSeek-V3 Languages (44) Language Train… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/tiny-aya-l2-thinker-multilingual-reasoning.texttext-generation100K<n<1M8 likes1.6k downloads25d agoHugging Face07ellamind /gsm8k-platinum-multilingual GSM8K Platinum Multilingual Multilingual translations of GSM8K Platinum, a rigorously cleaned and verified version of GSM8K containing 1,209 elementary math word problems requiring multi-step arithmetic reasoning. Source: madrylab/gsm8k-platinum (test split, 1,209 questions) Languages Config Language Examples ces Czech 100 dan Danish 100 deu German 1,209 fin Finnish 100 fra French 100 ita Italian 100 nld Dutch 100 pol Polish 100 spa Spanish… See the full description on the dataset page: https://huggingface.co/datasets/ellamind/gsm8k-platinum-multilingual.textquestion-answering1K<n<10K1 likes1.3k downloads6mo agoHugging Face08lightonai /Dolci-Think-SFT-32B-Multilingual Dolci-Think-SFT-32B-Multilingual Dolci-Think-SFT-32B-Multilingual is a large-scale multilingual long chain-of-thought (CoT) reasoning corpus spanning six languages: English, French, German, Spanish, Chinese, and Swahili. Each sample includes a question, a long-form reasoning trace, and a final answer, all translated into the target language, with sequences up to 32,768 tokens. It is released alongside the paper Rethinking the Multilingual Reasoning Gap with Layer Swap.… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/Dolci-Think-SFT-32B-Multilingual.texttext-generation1M<n<10M2 likes1.3k downloads4mo agoHugging Face09eddie-OB /gsm8k-multilingual-reasoning gsm8k-multilingual-reasoning GSM8K with reasoning translated to multiple languages Schema {"prompt": "...", "answer": "...", "reasoning": "...", "metadata": {...}} Usage from datasets importload_dataset ds = load_dataset("eddie-OB/gsm8k-multilingual-reasoning") print(ds["train"][0]) Source Derived from OpenAI GSM8K. texttext-generationn<1K1 likes826 downloads9mo agoHugging Face10ellamind /gpqa-multilingualgated GPQA Multilingual Multilingual translations of GPQA (Graduate-Level Google-Proof Q&A), a challenging multiple-choice benchmark requiring graduate-level expertise in biology, physics, and chemistry. Source: Idavidrein/gpqa (gpqa_main, 448 questions) Languages Config Language Examples ces Czech 448 dan Danish 448 deu German 448 fin Finnish 50 fra French 448 ita Italian 448 nld Dutch 448 pol Polish 448 spa Spanish 448 More to be added later.… See the full description on the dataset page: https://huggingface.co/datasets/ellamind/gpqa-multilingual.textquestion-answering1K<n<10K0 likes704 downloads7mo agoHugging Face11drooryck /multilingual-macaroni-corpus multilingual-macaroni training corpora The 8 training corpora behind our BabyLM 2026 Multilingual-track study of code-switched pretraining curricula (English / Dutch / Chinese) — see the paper Feeding BabyLMs Macaroni: Code-Switching Curricula Cause Cross-Lingual Convergence. All are derived from the official BabyLM 2026 multilingual corpora (babylm-{eng,nld,zho}) — no external text — and held to the same 100M byte-premium-adjusted-word budget, split across the three languages.… See the full description on the dataset page: https://huggingface.co/datasets/drooryck/multilingual-macaroni-corpus.texttext-generation1M<n<10M0 likes563 downloads7d agoHugging Face12Gabrui /multilingual_TinyStories Dataset Card for Multilingual TinyStories Dataset Details Dataset Description The Multilingual TinyStories dataset contains translations of the original TinyStories dataset, which consists of synthetically generated short stories using a small vocabulary suitable for 3 to 4-year-olds. These stories were originally generated by GPT-3.5 and GPT-4. The multilingual versions have been translated into various languages, including Spanish, Chinese, German, Turkish… See the full description on the dataset page: https://huggingface.co/datasets/Gabrui/multilingual_TinyStories.texttext-generation10M<n<100M1 likes536 downloads2y agoHugging Face13eddie-OB /gsm8k-multilingual gsm8k-multilingual GSM8K translated to multiple languages (no reasoning) Schema {"prompt": "...", "answer": "...", "metadata": {...}} Usage from datasets import load_dataset ds = load_dataset("eddie-OB/gsm8k-multilingual") print(ds["train"][0]) Source Derived from OpenAI GSM8K. texttext-generationn<1K0 likes530 downloads9mo agoHugging Face14hafsteinn /bfcl-v3-multilingual BFCL v3 in eight European languages Not manually verified yet. Every request is a machine translation. No native speaker has reviewed the set. Expect some unidiomatic phrasing and possibly a few meaning errors. Report problems in the Community tab. This is a translation of the five non-live, AST-checked categories of the Berkeley Function Calling Leaderboard v3 (simple 400, multiple 200, parallel 200, parallel_multiple 200, irrelevance 240; 1,240 items). The languages are… See the full description on the dataset page: https://huggingface.co/datasets/hafsteinn/bfcl-v3-multilingual.texttext-generation10K<n<100K0 likes526 downloads7d agoHugging Face15Similoluwa /african-multilingual-tokenizer-challenge African Multilingual Tokenizer Challenge dataset The frozen public corpus for the African Multilingual Tokenizer Challenge. It contains one balanced multilingual train split and one balanced validation split. Split Per language Total Train 40,000 240,000 Validation 4,000 24,000 Languages are English (en), French (fr), Hausa (ha), Swahili (sw), Yoruba (yo) and Amharic (am). Official public-test and private-test text are deliberately absent from this repository.… See the full description on the dataset page: https://huggingface.co/datasets/Similoluwa/african-multilingual-tokenizer-challenge.texttext-generation100K<n<1M0 likes470 downloads1mo agoHugging Face16Multilingual-Multimodal-NLP /McEval-InstructMcEval-Instruct data as described in the McEval Paper. Code for the evaluation and sft can be found on Github as McEval. texttext-generation10K<n<100K39 likes446 downloads2y agoHugging Face17giuliolovisotto /openai_multilingual_mmluMMLU professionally translated into 14 languages using professional human translators, sourced from OpenAI's simple-eval. Original files: english: https://openaipublic.blob.core.windows.net/simple-evals/mmlu.csv multilingual: https://openaipublic.blob.core.windows.net/simple-evals/mmlu_{language}.csv where language one of "AR-XY", "BN-BD", "DE-DE", "ES-LA", "FR-FR", "HI-IN", "ID-ID", "IT-IT", "JA-JP", "KO-KR", "PT-BR", "ZH-CN", "SW-KE", "YO-NG", "EN-US" texttext-generation100K<n<1M1 likes440 downloads2y agoHugging Face18ellamind /simpleqa-verified-multilingual SimpleQA Verified Multilingual Multilingual translations of SimpleQA Verified, a 1,000-prompt factuality benchmark from Google DeepMind that evaluates short-form parametric knowledge (facts stored in model weights). Source: google/simpleqa-verified (eval split, 1,000 examples) Languages Config Language Examples ces Czech 100 dan Danish 100 deu German 1,000 fra French 100 ita Italian 100 nld Dutch 100 pol Polish 100 spa Spanish 100 More to… See the full description on the dataset page: https://huggingface.co/datasets/ellamind/simpleqa-verified-multilingual.textquestion-answering1K<n<10K1 likes437 downloads7mo agoHugging Face19erenyeager-1 /tiny-aya-l2-thinker-multilingual-reasoning Tiny Aya L2 Multilingual Reasoning (44 languages) Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker. Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English. Languages (44) Language Train Test Total Amharic (am) 3,807 448 4,255 Arabic (ar) 22,968 2,538 25,506 Bulgarian (bg) 4,177 452 4,629 Bengali (bn) 3,803 422 4,225 Catalan (ca) 4,251 512 4,763 Czech… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/tiny-aya-l2-thinker-multilingual-reasoning.texttext-generation100K<n<1M0 likes425 downloads26d agoHugging Face20Dxniz /TinyStories-Multilingual Novelist: TinyStories Multilingual Edition Dataset Summary The TinyStories Multilingual Edition is a high-fidelity synthetic dataset of short, child-safe fiction designed to stress-test literary consistency, emotional warmth, and multilingual fluency in small models. Derived from the broader Novelist ecosystem, this subset focuses on narrative simplicity paired with complex moral and social themes. The dataset contains 15,688 high-quality stories across 28 languages. Each… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/TinyStories-Multilingual.texttext-generation10K<n<100K1 likes372 downloads7mo agoHugging Face21agentlans /multilingual-text Multilingual Text Dataset This dataset contains a curated selection of rows from multiple input datasets, where each row includes a text chunk of approximately 2000 tokens (as measured by Llama 3.1 tokenizer) verified to be written in the correct language. Only rows with properly classified language chunks are retained, ensuring high-quality multilingual data for analysis or model training. Preprocessing Steps Normalized whitespace, punctuation, Unicode characters, and… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/multilingual-text.texttext-generation1M<n<10M5 likes364 downloads1y agoHugging Face22aryashah00 /multilingual-sycophancy Multilingual Sycophancy A Parallel Benchmark for Cross-Lingual Alignment Failure across 38 Languages, 33 Opinion Categories, and 3 Resource Tiers. This dataset accompanies the research paper Sycophancy as a Multilingual Alignment Failure: How Safety Degrades Across Languages, Topics, and Models. It contains 188,100 parallel records (4,950 per language × 38 languages) — each a triple of (prompt, sycophantic response, non-sycophantic response) — designed for forced-choice… See the full description on the dataset page: https://huggingface.co/datasets/aryashah00/multilingual-sycophancy.texttext-classification100K<n<1M0 likes329 downloads3mo agoHugging Face23argilla /databricks-dolly-15k-curated-multilingual Dataset Card for "databricks-dolly-15k-curated-multilingual" A curated and multilingual version of the Databricks Dolly instructions dataset. It includes a programmatically and manually corrected version of the original en dataset. See below. STATUS: Currently, the original Dolly v2 English version has been curated combining automatic processing and collaborative human curation using Argilla (~400 records have been manually edited and fixed). The following graph shows a summary… See the full description on the dataset page: https://huggingface.co/datasets/argilla/databricks-dolly-15k-curated-multilingual.texttext-generation10K<n<100K54 likes324 downloads3y agoHugging Face24JRQi /DeepResearch-Bench-Multilingual DeepResearch Bench Multilingual Prompts This dataset provides prompt-level multilingual translations for the 100 research tasks used in muset-ai/DeepResearch-Bench-Dataset. The translations cover eight languages: en zh es it ar bn ja el What is included This repository focuses on the benchmark prompts only. On the Hugging Face Hub, the Dataset Viewer is configured with one default subset named all plus nine explicit subset configurations: source_prompt, en, zh, es… See the full description on the dataset page: https://huggingface.co/datasets/JRQi/DeepResearch-Bench-Multilingual.texttext-generation1K<n<10K1 likes304 downloads6mo agoHugging Face25agentlans /high-quality-multilingual-sentences High Quality Multilingual Sentences This dataset contains multilingual sentences derived from the agentlans/LinguaNova dataset. It includes 1.58 million rows across 51 different languages, each in its own configuration. Example row (from the all config): { "text": "امام جمعه اصفهان گفت: میزان نیاز آب شرب اصفهان ۱۱.۵ متر مکعب است که تمام استان اصفهان را پوشش میدهد و نسبت به قبل از انقلاب یکی از پیشرفتها در حوزه آب بوده است.", "fasttext": "fa", "gcld3": "fa" } Fields:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/high-quality-multilingual-sentences.texttext-generation1M<n<10M9 likes293 downloads2y agoHugging Face26agentlans /multilingual-sentences Multilingual Sentences Dataset contains sentences from 50 languages, grouped by their two-letter ISO 639-1 codes. The "all" configuration includes sentences from all languages. Dataset Overview Multilingual Sentence Dataset is a comprehensive collection of high-quality, linguistically diverse sentences. Dataset is designed to support a wide range of natural language processing tasks, including but not limited to language modeling, machine translation, and cross-lingual… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/multilingual-sentences.texttext-generation10M<n<100M6 likes277 downloads2y agoHugging Face27nhagar /c4_urls_multilingual Dataset Card for c4_urls_multilingual This dataset provides the URLs and top-level domains associated with training records in allenai/c4 (multilingual variant). It is part of a collection of datasets curated to make exploring LLM training datasets more straightforward and accessible. Dataset Details Dataset Description This dataset was created by downloading the source data, extracting URLs and top-level domains, and retaining only those record… See the full description on the dataset page: https://huggingface.co/datasets/nhagar/c4_urls_multilingual.texttext-generation1B<n<10B1 likes244 downloads1y agoHugging Face28karmx /TinyQuery-Tools-Multilingual TinyQuery Tools Multilingual GitHub: source code, setup guide, streaming examples and tests A synthetic, fictitious dataset for learning schema-conditioned SQL and tool actions from English, imperfect English, Hindi and Hinglish. Created for a four-hour, from-scratch small-model experiment. It contains no real user databases. Split Examples Purpose Train 347,376 Semantic scenarios, teacher language and randomized context variants Validation 1,200 Held-out domain… See the full description on the dataset page: https://huggingface.co/datasets/karmx/TinyQuery-Tools-Multilingual.texttext-generation100K<n<1M1 likes243 downloads23d agoHugging Face29Scicom-intl /Multilingual-Normalizer Multilingual TTS text normalizer (written → spoken) Training pairs for fine-tuning a small LLM as a text-to-speech normalizer: text is a sentence the way people type it (digits, currency symbols, dates, phone numbers, …) and normalized is the exact spoken form, in the same language, with nothing left that a TTS model cannot say. 52,698 rows — 16 monolingual locales and 6 Malaysian code-switched pairs. Every row is digit-free on the spoken side. from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Multilingual-Normalizer.texttext-generation100K<n<1M0 likes228 downloads1mo agoHugging Face30textdetox /multilingual_paradetoxMultilingual Text Detoxification with Parallel Data This is the multilingual parallel dataset for the text detoxification task. Prepared for TextDetox Shared Task. 📰 Updates [2025] The second edition of TextDetox shared task! webpage [2025] We extend our data to new languages! Now also included: Italian, French, Hebrew, Hinglish, Japanese, Tatar. Check our test part. [2025]We dived into the explainability of our data in our new COLING paper! [2024] You can check additional releases for… See the full description on the dataset page: https://huggingface.co/datasets/textdetox/multilingual_paradetox.texttext-generation1K<n<10K12 likes209 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.