Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01KodCode /KodCode-V1-SFT-R1 🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks. It contains 12 distinct subsets spanning various domains (from algorithmic to package-specific knowledge) and difficulty levels (from basic coding exercises to interview and competitive programming challenges). KodCode is designed for both supervised fine-tuning (SFT) and RL tuning. 🕸️… See the full description on the dataset page: https://huggingface.co/datasets/KodCode/KodCode-V1-SFT-R1.tabularquestion-answering100K<n<1M40 likes21k downloads2y agoHugging Face02KodCode /KodCode-V1-SFT-4o 🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding KodCode is the largest fully-synthetic open-source dataset providing verifiable solutions and tests for coding tasks. It contains 12 distinct subsets spanning various domains (from algorithmic to package-specific knowledge) and difficulty levels (from basic coding exercises to interview and competitive programming challenges). KodCode is designed for both supervised fine-tuning (SFT) and RL tuning. 🕸️… See the full description on the dataset page: https://huggingface.co/datasets/KodCode/KodCode-V1-SFT-4o.tabularquestion-answering100K<n<1M10 likes9.7k downloads2y agoHugging Face03geodesic-research /pa-warm-start-sft-xl-50b-mix geodesic-research/pa-warm-start-sft-xl-50b-mix Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/pa-warm-start-sft-xl-50b-mix", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-warm-start-sft-xl-50b-mix.tabular10M<n<100M0 likes9.2k downloads27d agoHugging Face04geodesic-research /pa-warm-start-sft-heavy-25b-mix geodesic-research/pa-warm-start-sft-heavy-25b-mix Auto-generated by dataset-builder. Each config below is a separate dataset produced from a versioned YAML build config. Load with: from datasets import load_dataset ds = load_dataset("geodesic-research/pa-warm-start-sft-heavy-25b-mix", "<config_name>", revision="<commit-sha>") Pin revision= to the specific commit SHA you want; without it, you get the current HEAD of the dataset repo, which may change when the builder re-pushes.… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/pa-warm-start-sft-heavy-25b-mix.tabular10M<n<100M0 likes6.2k downloads1mo agoHugging Face05OpenFormosa /barbet-long-context-sft Barbet long-context SFT Release a9fe3ba5b4c2869855e4e75118262797572d695c146eaa8fed7a180ec3544381 preserves 6633 active records. This is one joint assistant-only SFT dataset; no Barbet model training has been run. The skill-prefill migration has revised 1245 of 1254 records from its fixed base snapshot. Revisions replace their original records in the explicit shard lists above. Old bundles and releases remain available at their pinned commits. Additional records from other… See the full description on the dataset page: https://huggingface.co/datasets/OpenFormosa/barbet-long-context-sft.tabular1K<n<10K1 likes6k downloads2d agoHugging Face06mlfoundations-dev /Eurus-2-7B-SFT_eval_2e29 mlfoundations-dev/Eurus-2-7B-SFT_eval_2e29 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 MMLUPro JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 Accuracy 2.3 21.0 30.6 11.0 11.4 10.4 6.8 1.5 2.1 1.3 4.1 4.4 AIME24 Average Accuracy: 2.33% ± 0.67% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 0.00% 0 30 2 3.33% 1… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/Eurus-2-7B-SFT_eval_2e29.tabular1K<n<10K0 likes5.6k downloads1y agoHugging Face07SHSLab /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLab/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.tabulartext-generation10M<n<100M5 likes5.3k downloads1mo agoHugging Face08joshycodes /sorrel-sft-voicetabular100K<n<1M0 likes2.9k downloads23d agoHugging Face09geodesic-research /pa-warm-start-sft-xl-smoketabular10K<n<100K0 likes2.6k downloads28d agoHugging Face10openeurollm /Dolci-Think-SFT-translated Dolci-Think-SFT-translated Machine translations of the Dolci-Think-SFT-32B dataset, produced with gemma-4-31B-it. The samples selected for translation are those where content_quality == "excellent" according to the propella annotations. Columns Each row is a translated conversation plus the result of a post-translation quality filter: id — source record id. messages — the translated conversation (list of {content, role}). filter_pass — true if the row passed… See the full description on the dataset page: https://huggingface.co/datasets/openeurollm/Dolci-Think-SFT-translated.tabulartext-generation1M<n<10M0 likes1.7k downloads23m agoHugging Face11nyu-dice-lab /lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private Dataset Card for Evaluation run of princeton-nlp/Llama-3-Base-8B-SFT-RDPO Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-Base-8B-SFT-RDPO The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private.tabular100K<n<1M0 likes919 downloads2y agoHugging Face12laion /tts-realspeech-sft-en-de LAION TTS Real-Speech SFT — English + German, emotion-balanced 1,947,272 real recorded utterances — no synthetic voices — selected from freely-licensed corpora and balanced across 40 emotions x 2 languages. 6,996 hours, 313,844,544 MOSS frames (3,766,134,528 audio tokens), 79,337,527 aligned words. Each row is a self-contained TTS example: a corrected procedural caption, the transcript with word-level timestamps, the original audio, and the target MOSS-Audio-Tokenizer-v2 codes.… See the full description on the dataset page: https://huggingface.co/datasets/laion/tts-realspeech-sft-en-de.tabulartext-to-speech1M<n<10M0 likes879 downloads1mo agoHugging Face13geodesic-research /pa-warm-start-sft-xl-calibrationtabular100K<n<1M0 likes870 downloads28d agoHugging Face14geodesic-research /pa-warm-start-sft-heavy-25b-mix-longtabular1M<n<10M0 likes844 downloads1mo agoHugging Face15geodesic-research /pa-warm-start-sft-xl-50b-mix-metagaming-filteredtabular10M<n<100M0 likes839 downloads18d agoHugging Face16FineEnvs /SmolDataEnvs-sft 🛠️ SmolDataEnvs: SFT 5.5K+ RL tasks for hill-climbing small models in code and data science. A 2B model on these tasks. Left: what it optimises. Right: 144 held-out tasks it never trains on. Two runs over the same 5,000 tasks: shuffled against a curriculum ordered easiest to hardest. 4,677 worked examples of an agent doing data science the right way. Each row is a complete, verified-correct trajectory: read the question, poke at the data with a shell tool… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/SmolDataEnvs-sft.tabulartext-generation1K<n<10K2 likes821 downloads16d agoHugging Face17MercanAI /turkce-sft-qa-3.7m 🇹🇷 Turkish SFT/QA — Birleştirilmiş ve Tekrarsız Veri Seti 3,723,264 örnek. 24 açık Türkçe SFT/QA veri setinin, satır düzeyinde tekrar temizliği ve kalite kontrolünden geçirilmiş birleşimi. Her satır hangi veri setinden geldiğini taşır. English: A merged, row-level deduplicated and quality-filtered collection of 24 open Turkish SFT/QA datasets (3,723,264 examples). Every row carries its source dataset, source URL and original license. 🙏 Teşekkür /… See the full description on the dataset page: https://huggingface.co/datasets/MercanAI/turkce-sft-qa-3.7m.tabulartext-generation1M<n<10M0 likes753 downloads2mo agoHugging Face18argo11 /0399-tv-valid-clean-sft-tokenized-llmjp4-8btabular1M<n<10M0 likes749 downloads3mo agoHugging Face19Congliu /Chinese-DeepSeek-R1-Distill-data-110k-SFT 中文基于满血DeepSeek-R1蒸馏数据集(Chinese-Data-Distill-From-R1) 🤗 Hugging Face   |   🤖 ModelScope    |   🚀 Github    |   📑 Blog 注意:该版本为,可以直接SFT使用的版本,将原始数据中的思考和答案整合成output字段,大部分SFT代码框架均可直接直接加载训练。 本数据集为中文开源蒸馏满血R1的数据集,数据集中不仅包含math数据,还包括大量的通用类型数据,总数量为110K。 为什么开源这个数据? R1的效果十分强大,并且基于R1蒸馏数据SFT的小模型也展现出了强大的效果,但检索发现,大部分开源的R1蒸馏数据集均为英文数据集。 同时,R1的报告中展示,蒸馏模型中同时也使用了部分通用场景数据集。 为了帮助大家更好地复现R1蒸馏模型的效果,特此开源中文数据集。该中文数据集中的数据分布如下: Math:共计36568个样本, Exam:共计2432个样本, STEM:共计12648个样本,… See the full description on the dataset page: https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k-SFT.tabulartext-generation100K<n<1M225 likes716 downloads2y agoHugging Face20toroe /ReasonXL-SFT ReasonXL: A Multilingual Cross-Domain Reasoning Corpus ReasonXL is a large-scale multilingual reasoning corpus spanning five languages, with 2,538,450 positionally aligned examples per language (12,692,250 rows total). It is designed to support supervised fine-tuning of reasoning models with in-language chain-of-thought traces across diverse technical domains. Data Generation English source samples were drawn from 10 existing reasoning datasets, filtered and… See the full description on the dataset page: https://huggingface.co/datasets/toroe/ReasonXL-SFT.tabular10M<n<100M0 likes704 downloads2mo agoHugging Face21AdaMLLab /smolkalam-arabic-conversational-sft SmolKalam SmolKalam is a quality-filtered Arabic SFT dataset of 1,790,478 examples (~2.45B tokens), built as an ensemble translation of SmolTalk2. It covers multi-turn dialogue (23% of rows), reasoning traces (19% carry <think>), tool and function calling (4.4%), and long context, categories that are underrepresented in existing Arabic post-training data. The SmolTalk2 source mixtures are kept as subsets. Released with the paper SmolKalam: Ensemble Quality-Filtered Translation… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/smolkalam-arabic-conversational-sft.tabulartext-generation1M<n<10M3 likes677 downloads2mo agoHugging Face22Polygl0t /gigaverbo-v2-sft GigaVerbo-v2 SFT: A Large-Scale Portuguese Instruction-Tuning Dataset Dataset Summary GigaVerbo-v2 SFT is a large-scale instruction-tuning dataset designed for supervised fine-tuning of language models in Portuguese. The dataset comprises approximately 2.1 billion tokens (~4.4 GB) across 4 million instruction-following examples, organized into 12 distinct task categories. It is entirely composed of high-quality, LLM-generated data that has been carefully curated and… See the full description on the dataset page: https://huggingface.co/datasets/Polygl0t/gigaverbo-v2-sft.imagetext-generation1M<n<10M3 likes577 downloads7mo agoHugging Face23LingweiGu /openswe-success-sft-v1 Open-SWE-Traces successes — cleaned and tokenized for MiniCPM5 SFT (v1) Successful software-engineering agent trajectories from nvidia/Open-SWE-Traces (revision f8fb5b3d2c787f85f8a00f5fe04fe3f1a11088ef), filtered, validated and pre-tokenized with the native MiniCPM5-2B-Midtrain tokenizer and chat template (revision 0a45344e) for supervised fine-tuning with a 131,072-token context. Split Trajectories Unique tasks Repositories Input tokens Supervised tokens Longest… See the full description on the dataset page: https://huggingface.co/datasets/LingweiGu/openswe-success-sft-v1.tabulartext-generation10K<n<100K0 likes573 downloads12d agoHugging Face24ByteDance /VR-X-SFT-RL VR-X: Visual Reasoning Benchmark for UniVR VR-X contains three independent data blocks: SFT data organized by capability. VR-X-RL data for visual-reasoning reinforcement learning. VR-X-Eval held-out evaluation data. VR-X-RL and VR-X-Eval are independent from SFT and must be loaded separately. Public repository paths use anonymous source codes; no source-to-code mapping is published. Repository layout . ├── Robot Manipulation/ # SFT only │ └── RM-###/ │… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance/VR-X-SFT-RL.tabularvisual-question-answering100K<n<1M0 likes517 downloads3mo agoHugging Face25kelexine /fable-5-sft-traces Fable-5 SFT Traces Author / maintainer: kelexine (github.com/kelexine) A cleaned, anonymised, schema-normalised derivative of Kelexine/Fable-5-traces — agentic traces from Fable-5 (claude-fable-5), the model now publicly known as Claude Mythos — Anthropic's top-of-family frontier model at time of collection. The dataset supports three fine-tuning shapes off a single JSONL with no preprocessing required: Mode Fields used Full SFT (thinking + response) messages or… See the full description on the dataset page: https://huggingface.co/datasets/kelexine/fable-5-sft-traces.tabulartext-generation1K<n<10K14 likes510 downloads4mo agoHugging Face26geodesic-research /pa-warm-start-sft-xl-1b-smoketabular1K<n<10K0 likes497 downloads29d agoHugging Face27Xirui1208 /readall-sft-stage-a-b ReadAll / ReadTwice SFT:Stage A + Stage B 当前 ReadAll 模型的 SFT 数据和可移植训练包。训练链为: Qwen/Qwen2.5-7B-Instruct → Stage A (ReadTwice step286) → Stage B (ReadAll union step448) 阶段 训练行数 验证行数 Parquet 分片 全局 batch 学习率 1 epoch 更新数 A 73,416 1,676 15 + 1 256 1e-5 286 B 57,287 无独立验证集 29 128 5e-7 448 这些计数是 SFT 消息样本行数,包含 SKIM / UPDATE / FINAL,并非独立问题数。45 个 Parquet 共 1,030,251,959 字节。原始数据分片和 manifest 原样保留,逐一核对原始 SHA-256;没有删列、重新筛选或重新生成。 Stage A 使用普通多轮 assistant-token SFT、12,288 token… See the full description on the dataset page: https://huggingface.co/datasets/Xirui1208/readall-sft-stage-a-b.tabulartext-generation100K<n<1M0 likes495 downloads19d agoHugging Face28divelab /combined_gsm8k_math_dataset_dapo_math_17k_Qwen3-4B_ntokens2048_sfttabular100K<n<1M0 likes490 downloads7mo agoHugging Face29soketlabs /bhasha-sft Bhasha SFT Bhasha SFT is a massive collection of multiple open sourced Supervised Fine-Tuning datasets for training Multilingual Large Language Models. The dataset contains collation of over 13 million instances of instruction-response data for 3 Indian languages (Hindi, Gujarati, Bengali) and English having both human annotated and synthetic data. Curated by: Soket AI Labs Language(s) (NLP): [English, Hindi, Bengali, Gujarati] License: [cc-by-4.0, apache-2.0, mit]… See the full description on the dataset page: https://huggingface.co/datasets/soketlabs/bhasha-sft.tabularquestion-answering10M<n<100M4 likes473 downloads2y agoHugging Face30YijiaFan /UMM-Reflection-SFT-Data UMM-Reflection SFT Data The reflection-SFT data of UMM-Reflection (Learning Native Reflection in Unified Models). It trains UMM-Reflection-BAGEL-SFT. Research use only, non-commercial. The rows are derived from datasets with different licenses, some of them non-commercial. Each row records its source dataset and license in source_dataset and source_license, and each row follows the terms of its source. See LICENSE.md. Contents Part Rows Shards Size… See the full description on the dataset page: https://huggingface.co/datasets/YijiaFan/UMM-Reflection-SFT-Data.tabulartext-to-image10K<n<100K3 likes454 downloads12d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.