Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01IFM /Code-Reasoning Code-Reasoning Dataset Description Code problem-solving data with reasoning, direct-answer, and task-synthesis subsets. This repository is part of the K2 Horizon collection. The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards, which supports Dataset Viewer inspection and streaming access. K2 Horizon Dataset Series Dataset repository Focus Subsets IFM/TxT360-v2 Web and… See the full description on the dataset page: https://huggingface.co/datasets/IFM/Code-Reasoning.texttext-generation100M<n<1B99 likes51k downloads1mo agoHugging Face02smshahbaj /verifiable-code-reasoning Verifiable Code Reasoning Execution-verified Python problems with chain-of-thought Sandbox-checked solutions · Multi-test unit checks · Deduplicated instances · Training-ready sft_text Overview Verifiable Code Reasoning is a large-scale dataset of Python coding problems where every kept solution has passed sandboxed unit tests. Unlike scraped contest dumps or unverified LLM traces, an example enters this release only if: a reference… See the full description on the dataset page: https://huggingface.co/datasets/smshahbaj/verifiable-code-reasoning.texttext-generation1M<n<10M2 likes2.9k downloads24d agoHugging Face03LexyJawa /Code-Reasoning Code-Reasoning Dataset Description Code problem-solving data with reasoning, direct-answer, and task-synthesis subsets. This repository is part of the K2 Horizon collection. The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards, which supports Dataset Viewer inspection and streaming access. K2 Horizon Dataset Series Dataset repository Focus Subsets IFM/TxT360-v2 Web and… See the full description on the dataset page: https://huggingface.co/datasets/LexyJawa/Code-Reasoning.texttext-generation100M<n<1B0 likes1.5k downloads22d agoHugging Face04AMAImedia /NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54 ⚡ Each donation funds the next large quant. I host free GGUF or MoE quants as independent research. Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro. Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant. 🎉 Boosty🦄 &nbsp;|&nbsp; ☕ Buy Me a Coffee🦄 &nbsp;|&nbsp; ⭐ DonationAlerts🦄 💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.text1M<n<10M2 likes300 downloads23d agoHugging Face05efficientscaling /Z1-Code-Reasoning-107K Z1: Efficient Test-time Scaling with Code Train Large Language Model to Reason with Shifted Thinking [📜 Paper] • [🤗 HF Models] • [🐱 GitHub] Details Please refer to https://github.com/efficientscaling/Z1. Usage from datasets import load_dataset ds = load_dataset("efficientscaling/Z1-Code-Reasoning-107K")["train"] ds[0] Citation @misc{yu2025efficientscaling, title={Z1: Efficient Test-time Scaling with Code}… See the full description on the dataset page: https://huggingface.co/datasets/efficientscaling/Z1-Code-Reasoning-107K.text100K<n<1M21 likes271 downloads2y agoHugging Face06AMAImedia /NOESIS-50K-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54 ⚡ Each donation funds the next large quant. I host free GGUF or MoE quants as independent research. Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro. Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant. 🎉 Boosty🦄 &nbsp;|&nbsp; ☕ Buy Me a Coffee🦄 &nbsp;|&nbsp; ⭐ DonationAlerts🦄 💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/NOESIS-50K-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.text10K<n<100K7 likes249 downloads23d agoHugging Face07SMH-DEV-AI /NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54 NOESIS DORA SFT Dataset Multilingual supervised fine-tuning dataset built for the NOESIS QwQ+DeepSeek-R1 MoE pipeline. Released as part of the NOESIS Professional Multilingual Dubbing Automation Platform(framework: DHCF-FNO — Deterministic Hybrid Control Framework for Frozen Neural Operators). Founder: Ilia Bolotnikov Organization: AMAImedia.com X (Twitter): @AMAImediacom LinkedIn: Ilia Bolotnikov Telegram: @djbionicl NOESIS version: v14.8-NT89 Build date: 2026-04… See the full description on the dataset page: https://huggingface.co/datasets/SMH-DEV-AI/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.texttext-generation1M<n<10M1 likes236 downloads5mo agoHugging Face08ykarout /code-reasoning-phi4-templatetext1M<n<10M1 likes203 downloads1y agoHugging Face09sagea-ai /SAGE-Code-Reasoningtext100K<n<1M3 likes184 downloads9mo agoHugging Face10allenai /code-meta-reasoning-cleaned-final-string-idtext100K<n<1M5 likes165 downloads1y agoHugging Face11Banaxi-Tech /Deepseek-V4-Reasoning-Code-2500 DeepSeek Reasoning and Code Distillation Dataset This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research. The dataset file is: train.csv It contains 2… See the full description on the dataset page: https://huggingface.co/datasets/Banaxi-Tech/Deepseek-V4-Reasoning-Code-2500.tabulartext-generation1K<n<10K13 likes158 downloads5mo agoHugging Face12saurabh5 /open-code-reasoning-sfttext100K<n<1M0 likes140 downloads1y agoHugging Face13GetSoloTech /Code-Reasoning Code-Reasoning: Quality Filtered Dataset A high-quality, curated subset of the OpenCodeReasoning-2 dataset for training competitive programming models. This dataset has been processed with rigorous quality filtering and question reconstruction to ensure optimal training data. 📊 Dataset Overview This dataset contains competitive programming problems with high-quality solutions, specifically filtered and processed from the original nvidia/OpenCodeReasoning-2 dataset.… See the full description on the dataset page: https://huggingface.co/datasets/GetSoloTech/Code-Reasoning.textquestion-answering10K<n<100K8 likes135 downloads1y agoHugging Face14allenai /code-meta-reasoning-filteredtabular100K<n<1M1 likes112 downloads1y agoHugging Face15cublya /Code-Reasoning Code-Reasoning: Quality Filtered Dataset A high-quality, curated subset of the OpenCodeReasoning-2 dataset for training competitive programming models. This dataset has been processed with rigorous quality filtering and question reconstruction to ensure optimal training data. 📊 Dataset Overview This dataset contains competitive programming problems with high-quality solutions, specifically filtered and processed from the original nvidia/OpenCodeReasoning-2 dataset.… See the full description on the dataset page: https://huggingface.co/datasets/cublya/Code-Reasoning.textquestion-answering10K<n<100K0 likes92 downloads8mo agoHugging Face16ykarout /code-reasoning-thinking-3k-5ktext10K<n<100K2 likes82 downloads1y agoHugging Face17saurabh5 /open-code-reasoning-sft-n-32text100K<n<1M1 likes69 downloads1y agoHugging Face18MaLA-LM /mala-code-reasoning-v3text100M<n<1B5 likes68 downloads1y agoHugging Face19Lucien-shark /Linny-Code-Reasoning-Shrunkentext100K<n<1M1 likes54 downloads7mo agoHugging Face20saurabh5 /open-code-reasoning-rlvrtext1K<n<10K2 likes48 downloads1y agoHugging Face21naimulislam /reasoning_code_advanced_1m 💻 Reasoning Code Advanced 1M 📖 Dataset Summary Reasoning Code Advanced 1M is a massive-scale, synthetic dataset specifically engineered to improve the algorithmic reasoning and problem-solving capabilities of Large Language Models (LLMs). Featuring 1,000,000 unique coding samples, this dataset spans multiple programming languages (Python, JS, C++, etc.) and focuses on logic-heavy development tasks. A key feature of this dataset is its Adaptive Reasoning Architecture.… See the full description on the dataset page: https://huggingface.co/datasets/naimulislam/reasoning_code_advanced_1m.texttext-generation1M<n<10M1 likes48 downloads10mo agoHugging Face22efficientscaling /Z1-Code-Reasoning-Shortest-90Ktext10K<n<100K2 likes42 downloads1y agoHugging Face23efficientscaling /Z1-Code-Reasoning-Longest-33Ktext10K<n<100K1 likes40 downloads1y agoHugging Face24MaLA-LM /mala-code-reasoning-v2 MaLA Corpus: Massive Language Adaptation Corpus This MaLA code and reasoning dataset (V2) is used for training EMMA-500 Llama 3(.1) Mono/Bi model series. 🤗MaLA-LM/emma-500-llama3-8b-mono: CPT model trained on monolingual data mix in 500+ languages 🤗MaLA-LM/emma-500-llama3-8b-bi: CPT model trained on monolingual data mix in 500+ languages + bilingual translation data in 2,500+ language pairs 🤗MaLA-LM/emma-500-llama3.1-8b-mono: CPT model trained on monolingual data mix in… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-code-reasoning-v2.texttext-generation10M<n<100M9 likes40 downloads1y agoHugging Face25ykarout /nvidia-code-reasoning-cleantext10K<n<100K0 likes35 downloads1y agoHugging Face26lucsaint /Deepseek-V4-Reasoning-Code-2500 DeepSeek Reasoning and Code Distillation Dataset This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research. The dataset file is: train.csv It contains… See the full description on the dataset page: https://huggingface.co/datasets/lucsaint/Deepseek-V4-Reasoning-Code-2500.tabulartext-generation1K<n<10K0 likes35 downloads3mo agoHugging Face27XythicK /CodeReasoningPro CodeReasoningPro Dataset Summary CodeReasoningPro is a large-scale synthetic dataset comprising 1,785,725 competitive programming problems in Python, created by XythicK, an MLOps Engineer. Designed for supervised fine-tuning (SFT) of machine learning models for coding tasks, it draws inspiration from datasets like OpenCodeReasoning. The dataset includes problem statements, Python solutions, and reasoning explanations, covering algorithmic topics such as arrays, subarrays… See the full description on the dataset page: https://huggingface.co/datasets/XythicK/CodeReasoningPro.texttext-generation1M<n<10M4 likes33 downloads1y agoHugging Face28MaLA-LM /mala-code-reasoning MaLA Corpus: Massive Language Adaptation Corpus This MaLA code and reasoning dataset is used for training 🤗MaLA-LM/emma-500-llama2-7b. Dataset Summary The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. This subset contains code, reasoning data, and scientific papers. Project page: https://mala-lm.github.io Paper: https://arxiv.org/abs/2409.17892… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-code-reasoning.texttext-generation10M<n<100M5 likes33 downloads1y agoHugging Face29AlgorithmicResearchGroup /math_reasoning_automated_problem_solving_with_code_track_3text1K<n<10K1 likes31 downloads2y agoHugging Face30provie17 /react_native_code_review-reasoning-SFTtext1K<n<10K3 likes31 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.