Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01livebench /coding Dataset Card for "livebench/coding" LiveBench is a benchmark for LLMs designed with test set contamination and objective evaluation in mind. It has the following properties: LiveBench is designed to limit potential contamination by releasing new questions monthly, as well as having questions based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses. Each question has verifiable, objective ground-truth answers, allowing hard questions to be scored… See the full description on the dataset page: https://huggingface.co/datasets/livebench/coding.textn<1K12 likes9.7k downloads2y agoHugging Face02PrimeIntellect /verifiable-coding-problems SYNTHETIC-1 This is a subset of the task data used to construct SYNTHETIC-1. You can find the full collection here text100K<n<1M45 likes5.4k downloads2y agoHugging Face03SHSLab /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLab/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.tabulartext-generation10M<n<100M5 likes5.3k downloads1mo agoHugging Face04open-r1 /verifiable-coding-problems-python Dataset Card for Verifiable Coding Problems Python 10k This dataset contains all Python problems from PrimeIntellect's verifiable-coding-problems dataset. We have formatted the verification_info and metadata columns to be proper dictionaries, but otherwise the data is the same. Please see their dataset for more details. text10K<n<100K12 likes2.2k downloads2y agoHugging Face05Manusagents /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🌌 Omni-Frontier Distillation SFT The Definitive Evolution of Open-Source Distillation & Human-Crafted Expertise Repository: Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection "The most comprehensive multi‑domain SFT corpus ever assembled — fusing 6.86 million cleaned distillation samples with 9.14 million human‑crafted expert examples across medical, cybersecurity, chemical, robotics, humanities, and more. 16 million… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.texttext-generation10M<n<100M7 likes1.8k downloads2mo agoHugging Face06nvidia /Nemotron-RL-coding-competitive_coding Dataset Description: The Nemotron-RL-coding-competitive_coding dataset is a python-only, reasoning-based, synthetic dataset. It contains competitive coding style problems and their unit test cases. These questions and test cases are collected from CodeContests (deepmind/code_contests), and Open-R1 (open-r1/codeforces) . This dataset is released as part of NVIDIA NeMo Gym, a framework for building reinforcement learning environments to train large language models. NeMo Gym… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-coding-competitive_coding.text10K<n<100K27 likes1.5k downloads12d agoHugging Face07DSFFGFG456 /fable-5-coding-and-debugging-traces Claude Fable 5 Agent Traces 2,380 TRAJECTORIES · 12,490 TRAINING ROWS · 14 MB PARQUET · 663 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task… See the full description on the dataset page: https://huggingface.co/datasets/DSFFGFG456/fable-5-coding-and-debugging-traces.tabulartext-generation10K<n<100K4 likes1.5k downloads2mo agoHugging Face08greghavens /kimi-k3-coding-and-debugging-traces Kimi K3 Coding, Tool Use & Instruction Following Traces 582 TRAJECTORIES · 3,956 TRAINING ROWS · 3 MB PARQUET · 72 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/kimi-k3-coding-and-debugging-traces.tabulartext-generation1K<n<10K65 likes1.4k downloads2mo agoHugging Face09Lelonthecodeur /web-research-coding-5m Web Research + GitHub + Website Coding Dataset Version: 1.1.0 Total examples: 5,000,000 Splits train: 4,750,000 validation: 125,000 test: 125,000 Core capabilities Web search Web research Evidence extraction Fact verification Multi-hop research Multi-layer technical analysis Architecture analysis Root-cause analysis Security analysis Performance analysis UX analysis Design analysis Refactoring Code review Debugging Website coding Design systems… See the full description on the dataset page: https://huggingface.co/datasets/Lelonthecodeur/web-research-coding-5m.text1M<n<10M2 likes1.1k downloads22d agoHugging Face10hi-todayis-jh /Nemotron-RL-coding-quality-filtered Nemotron coding — quality pool, revision 2 8,201 retained tasks from 16,083 upstream train rows (51.0%). This is a static quality screen for Python standard-input/standard-output coding RL. It is not a reference-verified gold dataset. All difficulty levels are eligible. There is no model-accuracy filter, rollout generation, rating cutoff, random subsampling, or 3,200-row cap. Source: NVIDIA Nemotron-RL-coding-competitive_coding, revision 755d5910fc8646b385e3926eec08c152051cdc07.… See the full description on the dataset page: https://huggingface.co/datasets/hi-todayis-jh/Nemotron-RL-coding-quality-filtered.texttext-generation1K<n<10K0 likes1k downloads11d agoHugging Face11netpreme /coding_agent_tracesWe release coding agent traces using Claude Code for Opus ISL, OSL, ISL_new counts GPT-oss-120B ISL, OSL, ISL_new counts and their raw texts For Opus, only the locally saved files from the harness were used for analysis. Coding agents take multiple turns to carry out a task from the input prompt. To analyze the token distribution, two models were selected: Anthropic's Opus and OpenAI's gpt-oss-120B. The input sequence length (ISL), output sequence length (OSL) and the uncached, new input… See the full description on the dataset page: https://huggingface.co/datasets/netpreme/coding_agent_traces.text1K<n<10K3 likes1k downloads4mo agoHugging Face12open-r1 /verifiable-coding-problems-python_decontaminated-testedtext10K<n<100K0 likes793 downloads2y agoHugging Face13open-r1 /verifiable-coding-problems-python_decontaminated-tested-shuffledtext10K<n<100K2 likes588 downloads2y agoHugging Face14greghavens /glm-5.2-coding-and-debugging-traces GLM 5.2 Agent Traces 207 TRAJECTORIES · 1,821 TRAINING ROWS · 1 MB PARQUET · 35 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from GLM 5.2 (glm-5.2). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task domain. This is an actively growing… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/glm-5.2-coding-and-debugging-traces.tabulartext-generation1K<n<10K21 likes587 downloads3mo agoHugging Face15open-r1 /verifiable-coding-problems-python_decontaminatedtext10K<n<100K5 likes542 downloads2y agoHugging Face16hi-todayis-jh /Nemotron-LeetCode-coding-clean-3.2k Nemotron + LeetCode Coding Clean 3.2k 3,200 distinct training problems, seed 42, intended for Python coding reinforcement learning. This is a training mix, not a held-out benchmark. It combines the pinned default train Parquet split of Nemotron-RL-coding-competitive_coding with the train JSONL of LeetCodeDataset. Composition Source Questions Selection Nemotron / Codeforces 1,856 1000–1600, inclusive Nemotron / AtCoder 366 300–2000 display difficulty… See the full description on the dataset page: https://huggingface.co/datasets/hi-todayis-jh/Nemotron-LeetCode-coding-clean-3.2k.texttext-generation1K<n<10K0 likes507 downloads11d agoHugging Face17wanglab /variant_effect_coding 🧬 BioReasonIncentivizing Multimodal Biological Reasoning within a DNA-LLM Model Variant Effect Coding Dataset 50,083 core variant entries from GPN-MSA study using ClinVar pathogenic variants and gnomAD benign variants (MAF>5%), split by chromosome (Chr 1-7,9-22,X,Y for train, Chr 8 for test) for pathogenic/benign classification. Usage from datasets import load_dataset dataset = load_dataset("wanglab/variant_effect_coding") example = dataset["train"][0]… See the full description on the dataset page: https://huggingface.co/datasets/wanglab/variant_effect_coding.text10K<n<100K14 likes488 downloads1y agoHugging Face18katarinagresova /Genomic_Benchmarks_demo_coding_vs_intergenomic_seqs Dataset Card for "Genomic_Benchmarks_demo_coding_vs_intergenomic_seqs" More Information needed text100K<n<1M4 likes447 downloads3y agoHugging Face19Manusagents /Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.tabulartext-generation10M<n<100M1 likes440 downloads1mo agoHugging Face2011-47 /Organized_PreTrain_Coding_239ktext0 likes439 downloads2mo agoHugging Face21Self-Improving-Coding-Agents /SI2CA-Training-TrajectoriesDataset Card for SI2CA-Training-Trajectories [🌐 Website] • [🤗 Dataset] • [📜 Paper] • [🐱 GitHub] 💡 Introduction This dataset consists of 32,340 coding-agent trajectories generated by Qwen3.5-122B-A10B on the same 10,780 executable Python SWE tasks under the three trajectory-curation settings of Section 4.4 of the paper: standard sampling, full self-judgement, and an efficient discovered strategy found by the recursive self-improvement framework. Each task is… See the full description on the dataset page: https://huggingface.co/datasets/Self-Improving-Coding-Agents/SI2CA-Training-Trajectories.tabulartext-generation10K<n<100K0 likes405 downloads19d agoHugging Face22jonas-is-coding /german-wikipedia-articlestext1M<n<10M2 likes385 downloads2y agoHugging Face2311-47 /kimi-k3-coding-and-debugging-traces Kimi K3 Coding, Tool Use & Instruction Following Traces 582 TRAJECTORIES · 3,956 TRAINING ROWS · 3 MB PARQUET · 72 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/11-47/kimi-k3-coding-and-debugging-traces.tabulartext-generation1K<n<10K0 likes382 downloads24d agoHugging Face24zetomatoz /guidellm-agentic-coding-trajectories GuideLLM agentic coding trajectories A sampled serving-load benchmark derived from Thoughtworks agentic-coding-trajectories, for GuideLLM and an OpenAI-compatible /v1/chat/completions endpoint. There are 630 rows representing 481 unique source sessions, across the same 8turn, 24turn, and 48turn configurations as the earlier version. The configuration names now refer to original logical steps, not always HTTP request counts. Native tool steps expand into a tool-call request and a… See the full description on the dataset page: https://huggingface.co/datasets/zetomatoz/guidellm-agentic-coding-trajectories.tabulartext-generationn<1K5 likes362 downloads11d agoHugging Face25thoughtworks /agentic-coding-trajectories agentic-coding-trajectories A unified, tokenized corpus of 15,000 multi-turn agentic-coding sessions (618K turns, 41 turns/session avg) drawn from three publicly-released upstream datasets. Built for benchmarking LLM serving systems on realistic multi-turn coding-agent workloads. Why this exists Most LLM serving benchmarks use single-shot prompts. Real coding agents work in long multi-turn loops where each turn appends to a growing prompt. This corpus captures that shape… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/agentic-coding-trajectories.tabulartext-generation10K<n<100K1 likes309 downloads5mo agoHugging Face26rmems /agentic-coding-trajectories-grok46 Agentic Coding Trajectories (Grok 4.6) Rights & intended use: public research corpus, not training data. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. License: Synthetic Factory Research-Only License v1.0 (license: other, see LICENSE) (non-commercial). Release status: the raw… See the full description on the dataset page: https://huggingface.co/datasets/rmems/agentic-coding-trajectories-grok46.textn<1K1 likes306 downloads1mo agoHugging Face27MergeBench /coding_valtext1K<n<10K0 likes284 downloads1y agoHugging Face28rmems /agentic-coding-trajectories Agentic Coding Trajectories Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated coding-episode payload is published under data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/agentic-coding-trajectories.textn<1K1 likes279 downloads22d agoHugging Face29rogue-security /coding-agent-security-benchmark Coding Agent Security Benchmark A benchmark for evaluating whether an LLM can correctly identify security violations in the behavior of an autonomous coding agent - spanning dangerous shell commands, credential leakage, prompt injection, supply-chain risk, privacy leaks, and more. Each row is a single message sampled from a coding-agent session (a user instruction, a tool call the agent issued, a tool's response, or the agent's own output) paired with a ground-truth security… See the full description on the dataset page: https://huggingface.co/datasets/rogue-security/coding-agent-security-benchmark.textn<1K2 likes276 downloads2mo agoHugging Face3011-47 /glm-5.2-coding-and-debugging-traces GLM 5.2 Agent Traces 207 TRAJECTORIES · 1,821 TRAINING ROWS · 1 MB PARQUET · 35 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from GLM 5.2 (glm-5.2). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task domain. This is an actively growing… See the full description on the dataset page: https://huggingface.co/datasets/11-47/glm-5.2-coding-and-debugging-traces.tabulartext-generation1K<n<10K0 likes265 downloads24d agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.