Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01IFM /Code-Reasoning Code-Reasoning Dataset Description Code problem-solving data with reasoning, direct-answer, and task-synthesis subsets. This repository is part of the K2 Horizon collection. The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards, which supports Dataset Viewer inspection and streaming access. K2 Horizon Dataset Series Dataset repository Focus Subsets IFM/TxT360-v2 Web and… See the full description on the dataset page: https://huggingface.co/datasets/IFM/Code-Reasoning.texttext-generation100M<n<1B97 likes54k downloads1mo agoHugging Face02smshahbaj /verifiable-code-reasoning Verifiable Code Reasoning Execution-verified Python problems with chain-of-thought Sandbox-checked solutions · Multi-test unit checks · Deduplicated instances · Training-ready sft_text Overview Verifiable Code Reasoning is a large-scale dataset of Python coding problems where every kept solution has passed sandboxed unit tests. Unlike scraped contest dumps or unverified LLM traces, an example enters this release only if: a reference… See the full description on the dataset page: https://huggingface.co/datasets/smshahbaj/verifiable-code-reasoning.texttext-generation1M<n<10M2 likes1.8k downloads20d agoHugging Face03LexyJawa /Code-Reasoning Code-Reasoning Dataset Description Code problem-solving data with reasoning, direct-answer, and task-synthesis subsets. This repository is part of the K2 Horizon collection. The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards, which supports Dataset Viewer inspection and streaming access. K2 Horizon Dataset Series Dataset repository Focus Subsets IFM/TxT360-v2 Web and… See the full description on the dataset page: https://huggingface.co/datasets/LexyJawa/Code-Reasoning.texttext-generation100M<n<1B0 likes1.5k downloads18d agoHugging Face04efficientscaling /Z1-Code-Reasoning-107K Z1: Efficient Test-time Scaling with Code Train Large Language Model to Reason with Shifted Thinking [📜 Paper] • [🤗 HF Models] • [🐱 GitHub] Details Please refer to https://github.com/efficientscaling/Z1. Usage from datasets import load_dataset ds = load_dataset("efficientscaling/Z1-Code-Reasoning-107K")["train"] ds[0] Citation @misc{yu2025efficientscaling, title={Z1: Efficient Test-time Scaling with Code}… See the full description on the dataset page: https://huggingface.co/datasets/efficientscaling/Z1-Code-Reasoning-107K.text100K<n<1M21 likes266 downloads2y agoHugging Face05ykarout /code-reasoning-phi4-templatetext1M<n<10M1 likes249 downloads1y agoHugging Face06sagea-ai /SAGE-Code-Reasoningtext100K<n<1M3 likes243 downloads9mo agoHugging Face07SMH-DEV-AI /NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54 NOESIS DORA SFT Dataset Multilingual supervised fine-tuning dataset built for the NOESIS QwQ+DeepSeek-R1 MoE pipeline. Released as part of the NOESIS Professional Multilingual Dubbing Automation Platform(framework: DHCF-FNO — Deterministic Hybrid Control Framework for Frozen Neural Operators). Founder: Ilia Bolotnikov Organization: AMAImedia.com X (Twitter): @AMAImediacom LinkedIn: Ilia Bolotnikov Telegram: @djbionicl NOESIS version: v14.8-NT89 Build date: 2026-04… See the full description on the dataset page: https://huggingface.co/datasets/SMH-DEV-AI/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.texttext-generation1M<n<10M1 likes243 downloads5mo agoHugging Face08AMAImedia /NOESIS-50K-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54 ⚡ Each donation funds the next large quant. I host free GGUF or MoE quants as independent research. Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro. Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant. 🎉 Boosty🦄 &nbsp;|&nbsp; ☕ Buy Me a Coffee🦄 &nbsp;|&nbsp; ⭐ DonationAlerts🦄 💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/NOESIS-50K-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.text10K<n<100K7 likes227 downloads20d agoHugging Face09AMAImedia /NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54 ⚡ Each donation funds the next large quant. I host free GGUF or MoE quants as independent research. Local hardware: Mechrevo Kuangshi GM7AG0M — RTX 3060 Laptop 6GB GDDR6, 64GB DDR5, i7-12700H (14C/20T, 4.7GHz), Windows 11, Samsung 990 Pro. Good for imatrix and 0.6–35B-class work in RAM. 9B+ and searches need rented H200/Blackwell, typically $100 per quant. 🎉 Boosty🦄 &nbsp;|&nbsp; ☕ Buy Me a Coffee🦄 &nbsp;|&nbsp; ⭐ DonationAlerts🦄 💚 Thanks to Hugging Face for extra storage.🦄… See the full description on the dataset page: https://huggingface.co/datasets/AMAImedia/NOESIS-1M-reasoning-router-code-math-psych-opus47-deepseek4-qwen36-gemini31-r1-gpt54.text1M<n<10M2 likes214 downloads20d agoHugging Face10allenai /code-meta-reasoning-cleaned-final-string-idtext100K<n<1M5 likes181 downloads1y agoHugging Face11GetSoloTech /Code-Reasoning Code-Reasoning: Quality Filtered Dataset A high-quality, curated subset of the OpenCodeReasoning-2 dataset for training competitive programming models. This dataset has been processed with rigorous quality filtering and question reconstruction to ensure optimal training data. 📊 Dataset Overview This dataset contains competitive programming problems with high-quality solutions, specifically filtered and processed from the original nvidia/OpenCodeReasoning-2 dataset.… See the full description on the dataset page: https://huggingface.co/datasets/GetSoloTech/Code-Reasoning.textquestion-answering10K<n<100K8 likes162 downloads1y agoHugging Face12Banaxi-Tech /Deepseek-V4-Reasoning-Code-2500 DeepSeek Reasoning and Code Distillation Dataset This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research. The dataset file is: train.csv It contains 2… See the full description on the dataset page: https://huggingface.co/datasets/Banaxi-Tech/Deepseek-V4-Reasoning-Code-2500.tabulartext-generation1K<n<10K13 likes160 downloads5mo agoHugging Face13saurabh5 /open-code-reasoning-sfttext100K<n<1M0 likes150 downloads1y agoHugging Face14allenai /code-meta-reasoning-filteredtabular100K<n<1M1 likes129 downloads1y agoHugging Face15cublya /Code-Reasoning Code-Reasoning: Quality Filtered Dataset A high-quality, curated subset of the OpenCodeReasoning-2 dataset for training competitive programming models. This dataset has been processed with rigorous quality filtering and question reconstruction to ensure optimal training data. 📊 Dataset Overview This dataset contains competitive programming problems with high-quality solutions, specifically filtered and processed from the original nvidia/OpenCodeReasoning-2 dataset.… See the full description on the dataset page: https://huggingface.co/datasets/cublya/Code-Reasoning.textquestion-answering10K<n<100K0 likes96 downloads8mo agoHugging Face16ykarout /code-reasoning-thinking-3k-5ktext10K<n<100K2 likes87 downloads1y agoHugging Face17Lucien-shark /Linny-Code-Reasoning-Shrunkentext100K<n<1M1 likes82 downloads6mo agoHugging Face18MaLA-LM /mala-code-reasoning-v3text100M<n<1B5 likes71 downloads1y agoHugging Face19saurabh5 /open-code-reasoning-sft-n-32text100K<n<1M1 likes67 downloads1y agoHugging Face20naimulislam /reasoning_code_advanced_1m 💻 Reasoning Code Advanced 1M 📖 Dataset Summary Reasoning Code Advanced 1M is a massive-scale, synthetic dataset specifically engineered to improve the algorithmic reasoning and problem-solving capabilities of Large Language Models (LLMs). Featuring 1,000,000 unique coding samples, this dataset spans multiple programming languages (Python, JS, C++, etc.) and focuses on logic-heavy development tasks. A key feature of this dataset is its Adaptive Reasoning Architecture.… See the full description on the dataset page: https://huggingface.co/datasets/naimulislam/reasoning_code_advanced_1m.texttext-generation1M<n<10M1 likes60 downloads10mo agoHugging Face21saurabh5 /open-code-reasoning-rlvrtext1K<n<10K2 likes49 downloads1y agoHugging Face22efficientscaling /Z1-Code-Reasoning-Shortest-90Ktext10K<n<100K2 likes45 downloads1y agoHugging Face23MaLA-LM /mala-code-reasoning-v2 MaLA Corpus: Massive Language Adaptation Corpus This MaLA code and reasoning dataset (V2) is used for training EMMA-500 Llama 3(.1) Mono/Bi model series. 🤗MaLA-LM/emma-500-llama3-8b-mono: CPT model trained on monolingual data mix in 500+ languages 🤗MaLA-LM/emma-500-llama3-8b-bi: CPT model trained on monolingual data mix in 500+ languages + bilingual translation data in 2,500+ language pairs 🤗MaLA-LM/emma-500-llama3.1-8b-mono: CPT model trained on monolingual data mix in… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-code-reasoning-v2.texttext-generation10M<n<100M9 likes44 downloads1y agoHugging Face24efficientscaling /Z1-Code-Reasoning-Longest-33Ktext10K<n<100K1 likes42 downloads1y agoHugging Face25ykarout /nvidia-code-reasoning-cleantext10K<n<100K0 likes39 downloads1y agoHugging Face26XythicK /CodeReasoningPro CodeReasoningPro Dataset Summary CodeReasoningPro is a large-scale synthetic dataset comprising 1,785,725 competitive programming problems in Python, created by XythicK, an MLOps Engineer. Designed for supervised fine-tuning (SFT) of machine learning models for coding tasks, it draws inspiration from datasets like OpenCodeReasoning. The dataset includes problem statements, Python solutions, and reasoning explanations, covering algorithmic topics such as arrays, subarrays… See the full description on the dataset page: https://huggingface.co/datasets/XythicK/CodeReasoningPro.texttext-generation1M<n<10M4 likes37 downloads1y agoHugging Face27lucsaint /Deepseek-V4-Reasoning-Code-2500 DeepSeek Reasoning and Code Distillation Dataset This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research. The dataset file is: train.csv It contains… See the full description on the dataset page: https://huggingface.co/datasets/lucsaint/Deepseek-V4-Reasoning-Code-2500.tabulartext-generation1K<n<10K0 likes34 downloads3mo agoHugging Face28ansulev /deepseek-v4-reasoning-code-2500 Mirror: lucsaint/Deepseek-V4-Reasoning-Code-2500 Pinned snapshot / mirror of lucsaint/Deepseek-V4-Reasoning-Code-2500, re-hosted for PROTISEC research reproducibility. Redistributed under the upstream license (apache-2.0) with attribution — all credit to the original author. Original author: lucsaint Source dataset: lucsaint/Deepseek-V4-Reasoning-Code-2500 License: apache-2.0 Family: coding_traces Mode: full Rows cached: 2556 Changes vs upstream: cached snapshot, possibly… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/deepseek-v4-reasoning-code-2500.tabular1K<n<10K0 likes34 downloads2mo agoHugging Face29AlgorithmicResearchGroup /math_reasoning_automated_problem_solving_with_code_track_3text1K<n<10K1 likes32 downloads2y agoHugging Face30MaLA-LM /mala-code-reasoning MaLA Corpus: Massive Language Adaptation Corpus This MaLA code and reasoning dataset is used for training 🤗MaLA-LM/emma-500-llama2-7b. Dataset Summary The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. This subset contains code, reasoning data, and scientific papers. Project page: https://mala-lm.github.io Paper: https://arxiv.org/abs/2409.17892… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-code-reasoning.texttext-generation10M<n<100M5 likes32 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.