Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CoderOfCode /ship-tracking-data2 likes95k downloads3mo agoHugging Face02IFM /Code-Reasoning Code-Reasoning Dataset Description Code problem-solving data with reasoning, direct-answer, and task-synthesis subsets. This repository is part of the K2 Horizon collection. The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards, which supports Dataset Viewer inspection and streaming access. K2 Horizon Dataset Series Dataset repository Focus Subsets IFM/TxT360-v2 Web and… See the full description on the dataset page: https://huggingface.co/datasets/IFM/Code-Reasoning.texttext-generation100M<n<1B97 likes54k downloads1mo agoHugging Face03microsoft /rStar-Coder rStar-Coder Dataset Project GitHub | Paper Dataset Description rStar-Coder is a large-scale competitive code problem dataset containing 418K programming problems, 580K long-reasoning solutions, and rich test cases of varying difficulty levels. This dataset aims to enhance code reasoning capabilities in large language models, particularly in handling competitive code problems. Experiments on Qwen models (1.5B-14B) across various code reasoning benchmarks demonstrate… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/rStar-Coder.text1M<n<10M249 likes9k downloads1y agoHugging Face04togethercomputer /CoderForge-Preview CoderForge-Preview: SOTA Open Dataset for Training Efficient Agents CoderForge-Preview is the largest open test-verified coding agent dataset. Fine-tuning Qwen-3 32B on it, we boost SWE-Bench Verified performance 23.0% → 59.4% pass@1 and rank #1 among open-data and #2 among open-weight models ≤32B parameters. Limitations Adaptability to different scaffolds: We generated all trajectories using a single scaffold and fixed tool set (no permutations). Models trained via… See the full description on the dataset page: https://huggingface.co/datasets/togethercomputer/CoderForge-Preview.text100K<n<1M176 likes4.1k downloads7mo agoHugging Face05tsinghua-sigs-robot-lab /VeriLoop-Coder-E1-Evaluation-Evidence VeriLoop Coder-E1 Evaluation Evidence This repository contains the public evaluation-evidence packages referenced by the official VeriLoop Coder-E1 benchmark result files. Model repository: tsinghua-sigs-robot-lab/veriloop-coder-e1 Evidence packages Benchmark Evidence directory DeepSWE veriloop-coder-e1-deepswe-evaluation-evidence-v1.0.0 SWE-bench Pro veriloop-coder-e1-swe-bench-pro-evaluation-evidence-v1.0.0 SWE-bench Verified… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-Coder-E1-Evaluation-Evidence.0 likes1.8k downloads2mo agoHugging Face06code-rag-bench /github-repos-pythonThe Github repository retrieval source for [code-rag-bench], containing all Python files from the entire GitHub dump (in github-repos) text100K<n<1M2 likes1.6k downloads2y agoHugging Face07Lite-Coder /LiteCoder-Terminal-RL-preview LiteCoder-Terminal-RL-preview Paper | Code | Blog Post This dataset contains 602 standardized Harbor terminal environments and was released as part of the paper LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents. Unlike static text-only instructions, these environments are fully executable and are designed to support the training of terminal-based agents. Environment Generation Pipeline The lack of high-quality, executable… See the full description on the dataset page: https://huggingface.co/datasets/Lite-Coder/LiteCoder-Terminal-RL-preview.tabulartext-generationn<1K6 likes1.6k downloads3mo agoHugging Face08LexyJawa /Code-Reasoning Code-Reasoning Dataset Description Code problem-solving data with reasoning, direct-answer, and task-synthesis subsets. This repository is part of the K2 Horizon collection. The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards, which supports Dataset Viewer inspection and streaming access. K2 Horizon Dataset Series Dataset repository Focus Subsets IFM/TxT360-v2 Web and… See the full description on the dataset page: https://huggingface.co/datasets/LexyJawa/Code-Reasoning.texttext-generation100M<n<1B0 likes1.5k downloads18d agoHugging Face09inclusionAI /Ling-Coder-SFT 🤗 Hugging Face 🤖 ModelScope 🖥️ GitHub Ling-Coder Dataset The Ling-Coder Dataset comprises the following components: Ling-Coder-SFT: A subset of SFT data used for training Ling-Coder Lite, containing more than 5 million samples. Ling-Coder-DPO: A subset of DPO data used for training Ling-Coder Lite, containing 250k samples. Ling-Coder-SyntheticQA: A subset of synthetic data used for annealing training of Ling-Coder Lite, containing more… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/Ling-Coder-SFT.texttext-generation1M<n<10M45 likes1.4k downloads2y agoHugging Face10ronantakizawa /github-codereview Code Review Dataset A large-scale dataset of the best human-written code reviews from top GitHub repositories. Each row captures a moment where a human code reviewer left an inline comment on a pull request, and the author subsequently modified the code in response. The dataset also includes negative examples — code from the same PRs that passed review without comments — to help models learn when code is acceptable. This provides a natural signal for training models to: Generate… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-codereview.tabulartext-generation100K<n<1M64 likes1.2k downloads7mo agoHugging Face11QuixiAI /dolphin-coder dolphin-coder This dataset is transformed from https://www.kaggle.com/datasets/erichartford/leetcode-rosetta it is used to train dolphin-coder model text100K<n<1M62 likes857 downloads3y agoHugging Face12Fraser /dream-coder Program Synthesis Data Generated program synthesis datasets used to train dreamcoder. Currently just supports text & list data. text1K<n<10K6 likes809 downloads4y agoHugging Face13Crownelius /High-Coder-Reasoning-Multi-Turn High-Coder-Reasoning-Multi-Turn Dataset Description This dataset contains high-quality, multi-turn coding conversations focused on code critique, transformation (fixing, translating, and repurposing), and architectural analysis. It was generated using a proprietary pipeline targeting the openrouter/hunter-alpha model to simulate expert-level software engineering workflows. Pipeline Details: Each sample consists of three turns: Critique: A detailed… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/High-Coder-Reasoning-Multi-Turn.text10K<n<100K16 likes753 downloads3mo agoHugging Face14ricdomolm /mini-coder-trajs-400kGenerated using Qwen 3 Coder 30B A3B, mini-swe-agent, and SWE-smith. Used to train the mini-coder models Citation @article{olmedo2026computational, title={Computational Arbitrage in AI Model Markets}, author={Olmedo, Ricardo and Sch{\"o}lkopf, Bernhard and Hardt, Moritz}, journal={The International Conference on Machine Learning}, year={2026} } text100K<n<1M17 likes751 downloads2mo agoHugging Face15IIGroup /X-Coder-SFT-376k X-Coder: Advancing Competitive Programming with Fully Synthetic Tasks, Solutions, and Tests Dataset Overview X-Coder-SFT-376k is a large-scale, fully synthetic dataset for advancing competitive programming. The dataset comprises 4 subsets with a total of 887,321 synthetic records across 423,883 unique queries. It is designed for supervised fine-tuning and suitbale for cold start to train code reasoning foundations. X-Coder-SFT-376k is curated by sota reasoning models.… See the full description on the dataset page: https://huggingface.co/datasets/IIGroup/X-Coder-SFT-376k.texttext-generation100K<n<1M21 likes742 downloads8mo agoHugging Face16coderchen01 /MMSD2.0 MMSD2.0: Towards a Reliable Multi-modal Sarcasm Detection System This is a copy of the dataset uploaded on Hugging Face for easy access. The original data comes from this work, which is an improvement upon a previous study. Usage from typing import TypedDict, cast import pytorch_lightning as pl from datasets import Dataset, load_dataset from torch import Tensor from torch.utils.data import DataLoader from transformers import CLIPProcessor class… See the full description on the dataset page: https://huggingface.co/datasets/coderchen01/MMSD2.0.imagefeature-extraction10K<n<100K7 likes731 downloads2y agoHugging Face17nebula2025 /CodeR-Pile Towards A Generalist Code Embedding Model Based On Massive Data Synthesis Introduction This repository contains the synthetic training data introduced in the paper Towards A Generalist Code Embedding Model Based On Massive Data Synthesis. The dataset is designed to enhance text embeddings for code retrieval tasks. For more details, please refer to our Github repo: CodeR. Load Dataset Simple Example An example to load the dataset:… See the full description on the dataset page: https://huggingface.co/datasets/nebula2025/CodeR-Pile.text1M<n<10M4 likes708 downloads1y agoHugging Face18coderofpears /ClankerDatasettext1M<n<10M0 likes648 downloads6mo agoHugging Face19code-rag-bench /github-reposThe entire dump of GitHub repositories. text100K<n<1M3 likes632 downloads2y agoHugging Face20code-review-bench /code-review-bench Code Review Bench A paired online-offline benchmark for AI code review. Splits online — Stratified sample of 1,135 bot-reviewed PRs, scraped from open-source Github repositories and scored by the online benchmark (15 tools, Feb–Apr 2026). offline — 136 expert-curated golden issues across 50 PRs (5 repositories). Provenance The offline golden issues extend the 50-PR benchmark originally created by Greptile (2025) and refined by Augment (2025). Our… See the full description on the dataset page: https://huggingface.co/datasets/code-review-bench/code-review-bench.tabulartext-generation1K<n<10K1 likes555 downloads2mo agoHugging Face21coderofpears /clanker-data0 likes542 downloads3mo agoHugging Face22Coder-AN /StreakNet-Dataset StreakNet-Dataset StreakNet-Dataset is an underwater laser imaging dataset for UCLR systems, introduced in the paper StreakNet-Arch: An Anti-scattering Network-based Architecture for Underwater Carrier LiDAR-Radar Imaging. It comprises a collection of streak-tube images captured by a UCLR system at distances of 10m, 13m, 15m, and 20m, contributing 2,695,168 real-world underwater 3D point cloud data. For the associated source code, models, and comprehensive usage instructions… See the full description on the dataset page: https://huggingface.co/datasets/Coder-AN/StreakNet-Dataset.image-to-3d0 likes528 downloads1y agoHugging Face23gudo7208 /CAD-Coder CAD-Coder Dataset This is the official dataset for the paper "CAD-Coder: Text-to-CAD Generation with Chain-of-Thought and Geometric Reward". Accepted at NeurIPS 2025 (Poster) Dataset Description CAD-Coder Dataset is a large-scale Text-to-CadQuery dataset containing natural language descriptions of 3D CAD models paired with executable CadQuery Python code. The dataset enables training and evaluating language models to generate parametric CAD code from textual descriptions.… See the full description on the dataset page: https://huggingface.co/datasets/gudo7208/CAD-Coder.texttext-generation100K<n<1M5 likes528 downloads9mo agoHugging Face24MnemicAI /Ling-Coder-SFT-English-Clean Ling-Coder-SFT-English-Clean A cleaned, English-only version of inclusionAI/Ling-Coder-SFT — one of the largest open-source coding instruction datasets (~5.1M samples). Split by programming language for easy access. Curated by MnemicAI Origin Story While building our Mnemic COCM-COT training pipeline — a multi-language coding instruction dataset with stratified topic sampling — we discovered that 11.44% of Ling-Coder-SFT contains Chinese/CJK characters mixed into what… See the full description on the dataset page: https://huggingface.co/datasets/MnemicAI/Ling-Coder-SFT-English-Clean.texttext-generation1M<n<10M1 likes454 downloads6mo agoHugging Face25fan-shu /swe-mt-combined-coderforge-hero-lego-nex-swezero fan-shu/swe-mt-combined-coderforge-hero-lego-nex-swezero Concatenated mid-train dataset for Qwen3 Thinking SFT. Each source subset is loaded in order and concatenated into a single config so one training epoch visits every trajectory exactly once (no interleave / no oversampling). Built from fan-shu/swe-instruct-trajectories-empty-think-inserted. Source subsets (7) togethercomputer__CoderForge-Preview nvidia__SWE-Zero-openhands-trajectories nex-agi__agent-sft… See the full description on the dataset page: https://huggingface.co/datasets/fan-shu/swe-mt-combined-coderforge-hero-lego-nex-swezero.text100K<n<1M0 likes445 downloads3mo agoHugging Face26open-llm-leaderboard-old /details_deepseek-ai__deepseek-coder-1.3b-instruct Dataset Card for Evaluation run of deepseek-ai/deepseek-coder-1.3b-instruct Dataset Summary Dataset automatically created during the evaluation run of model deepseek-ai/deepseek-coder-1.3b-instruct on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_deepseek-ai__deepseek-coder-1.3b-instruct.0 likes432 downloads3y agoHugging Face27Lite-Coder /LiteCoder-Terminal-SFT LiteCoder-SFT-Terminal Paper | Code | Blog Post LiteCoder-SFT-Terminal is a dataset of 11,255 agent trajectories in terminal environments, introduced in the paper LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agents. Fine-tuned on this data, the LiteCoder-Terminal-30b-a3b-sft model achieves 31.5% Pass@1 on Terminal Bench Pro, while the LiteCoder-Terminal-4b-sft model shows distinct gains over its baseline. Released Artifacts… See the full description on the dataset page: https://huggingface.co/datasets/Lite-Coder/LiteCoder-Terminal-SFT.texttext-generation10K<n<100K9 likes427 downloads4mo agoHugging Face28inclusionAI /Ling-Coder-SyntheticQA 🤗 Hugging Face 🤖 ModelScope 🖥️ GitHub Ling-Coder Dataset The Ling-Coder Dataset comprises the following components: Ling-Coder-SFT: A subset of SFT data used for training Ling-Coder Lite, containing more than 5 million samples. Ling-Coder-DPO: A subset of DPO data used for training Ling-Coder Lite, containing 250k samples. Ling-Coder-SyntheticQA: A subset of synthetic data used for annealing training of Ling-Coder Lite, containing more… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/Ling-Coder-SyntheticQA.texttext-generation10M<n<100M17 likes409 downloads2y agoHugging Face29smcleod /golang-coderQ&A style combined, deduplicated dataset including portions of: Golang best practices and coding guides (general Q&A) https://huggingface.co/datasets/smcleod/golang-programming-style-best-practices (MIT) Golang questions (general Q&A) https://huggingface.co/datasets/ExAi/Code-Golang-QA-2k (Apache2) Golang functions (code & description) https://huggingface.co/datasets/google/code_x_glue_ct_code_to_text (c-uda) Golang snippets (code & description)… See the full description on the dataset page: https://huggingface.co/datasets/smcleod/golang-coder.texttext-generation100K<n<1M19 likes401 downloads2y agoHugging Face30IIGroup /X-Coder-RL-40k X-Coder-RL-40k X-Coder-RL-40k is a fully synthetic reinforcement learning dataset for competitive programming, containing 40k high-quality tasks with verified test cases. Dataset Structure The dataset is organized by difficulty level: File Difficulty part_0000.parquet Easiest part_0001.parquet Easy part_0002.parquet Medium part_0003.parquet Hard part_0004.parquet Hardest Task Difficulty Distribution Table: Distribution of Proprietary… See the full description on the dataset page: https://huggingface.co/datasets/IIGroup/X-Coder-RL-40k.textn<1K2 likes378 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.