Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SHSLab /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLab/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.tabulartext-generation10M<n<100M5 likes5.3k downloads1mo agoHugging Face02DSFFGFG456 /fable-5-coding-and-debugging-traces Claude Fable 5 Agent Traces 2,380 TRAJECTORIES · 12,490 TRAINING ROWS · 14 MB PARQUET · 663 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task… See the full description on the dataset page: https://huggingface.co/datasets/DSFFGFG456/fable-5-coding-and-debugging-traces.tabulartext-generation10K<n<100K4 likes1.5k downloads2mo agoHugging Face03greghavens /kimi-k3-coding-and-debugging-traces Kimi K3 Coding, Tool Use & Instruction Following Traces 582 TRAJECTORIES · 3,956 TRAINING ROWS · 3 MB PARQUET · 72 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/kimi-k3-coding-and-debugging-traces.tabulartext-generation1K<n<10K65 likes1.4k downloads2mo agoHugging Face04greghavens /glm-5.2-coding-and-debugging-traces GLM 5.2 Agent Traces 207 TRAJECTORIES · 1,821 TRAINING ROWS · 1 MB PARQUET · 35 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from GLM 5.2 (glm-5.2). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task domain. This is an actively growing… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/glm-5.2-coding-and-debugging-traces.tabulartext-generation1K<n<10K21 likes587 downloads3mo agoHugging Face05Manusagents /Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.tabulartext-generation10M<n<100M1 likes440 downloads1mo agoHugging Face06Self-Improving-Coding-Agents /SI2CA-Training-TrajectoriesDataset Card for SI2CA-Training-Trajectories [🌐 Website] • [🤗 Dataset] • [📜 Paper] • [🐱 GitHub] 💡 Introduction This dataset consists of 32,340 coding-agent trajectories generated by Qwen3.5-122B-A10B on the same 10,780 executable Python SWE tasks under the three trajectory-curation settings of Section 4.4 of the paper: standard sampling, full self-judgement, and an efficient discovered strategy found by the recursive self-improvement framework. Each task is… See the full description on the dataset page: https://huggingface.co/datasets/Self-Improving-Coding-Agents/SI2CA-Training-Trajectories.tabulartext-generation10K<n<100K0 likes405 downloads19d agoHugging Face0711-47 /kimi-k3-coding-and-debugging-traces Kimi K3 Coding, Tool Use & Instruction Following Traces 582 TRAJECTORIES · 3,956 TRAINING ROWS · 3 MB PARQUET · 72 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/11-47/kimi-k3-coding-and-debugging-traces.tabulartext-generation1K<n<10K0 likes382 downloads24d agoHugging Face08zetomatoz /guidellm-agentic-coding-trajectories GuideLLM agentic coding trajectories A sampled serving-load benchmark derived from Thoughtworks agentic-coding-trajectories, for GuideLLM and an OpenAI-compatible /v1/chat/completions endpoint. There are 630 rows representing 481 unique source sessions, across the same 8turn, 24turn, and 48turn configurations as the earlier version. The configuration names now refer to original logical steps, not always HTTP request counts. Native tool steps expand into a tool-call request and a… See the full description on the dataset page: https://huggingface.co/datasets/zetomatoz/guidellm-agentic-coding-trajectories.tabulartext-generationn<1K5 likes362 downloads11d agoHugging Face09thoughtworks /agentic-coding-trajectories agentic-coding-trajectories A unified, tokenized corpus of 15,000 multi-turn agentic-coding sessions (618K turns, 41 turns/session avg) drawn from three publicly-released upstream datasets. Built for benchmarking LLM serving systems on realistic multi-turn coding-agent workloads. Why this exists Most LLM serving benchmarks use single-shot prompts. Real coding agents work in long multi-turn loops where each turn appends to a growing prompt. This corpus captures that shape… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/agentic-coding-trajectories.tabulartext-generation10K<n<100K1 likes309 downloads5mo agoHugging Face1011-47 /glm-5.2-coding-and-debugging-traces GLM 5.2 Agent Traces 207 TRAJECTORIES · 1,821 TRAINING ROWS · 1 MB PARQUET · 35 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from GLM 5.2 (glm-5.2). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task domain. This is an actively growing… See the full description on the dataset page: https://huggingface.co/datasets/11-47/glm-5.2-coding-and-debugging-traces.tabulartext-generation1K<n<10K0 likes265 downloads24d agoHugging Face11dacorvo /transformers-coding-session-captures dacorvo/transformers-coding-session-captures HTTP captures of agent ↔ model interactions — one parquet row per /v1/chat/completions call. Produced by agentcap. Native session traces for the same runs live in companion datasets named transformers-coding-session-<agent>-traces. They're all grouped under the transformers-coding-session Collection alongside this dataset. Join on run_id. Loading from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/dacorvo/transformers-coding-session-captures.tabular10K<n<100K0 likes220 downloads4mo agoHugging Face12Aratako /Synthetic-JP-EN-Coding-Dataset-801k Synthetic-JP-EN-Coding-Dataset-801k Magpieによって作成したコードSFTデータセットであるAratako/Synthetic-JP-EN-Coding-Dataset-Magpie-69kを元に、Evol-Instructのような手法を用いて複数のinstructionとresonseを生成し拡張して作成した、日英混合801262件のコードSFT用合成データセットです。 日本語: 173849件 英語: 627413件 元のinstructionの作成に利用したモデルは以下の通りです。modelキーに該当レコードの作成に利用したモデル情報があります。 nvidia/Nemotron-4-340B-Instruct microsoft/Phi-3-medium-4k-instruct mistralai/Mixtral-8x22B-Instruct-v0.1… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Synthetic-JP-EN-Coding-Dataset-801k.tabulartext-generation100K<n<1M17 likes211 downloads2y agoHugging Face13eitans-pu /appellate-coding-inputtabular100K<n<1M0 likes199 downloads11mo agoHugging Face14songlab /ukb_finemapped_coding UKBB finemapped coding variants Predictions from all models tabular1K<n<10K1 likes195 downloads10mo agoHugging Face15jiajiale9 /kimi-k3-coding-and-debugging-traces Kimi K3 Coding, Tool Use & Instruction Following Traces 697 TRAJECTORIES · 4,890 TRAINING ROWS · 3 MB PARQUET · 89 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/jiajiale9/kimi-k3-coding-and-debugging-traces.tabulartext-generation1K<n<10K0 likes172 downloads3mo agoHugging Face16rashdan1 /glm-5.2-coding-and-debugging-traces GLM 5.2 Agent Traces 207 TRAJECTORIES · 1,821 TRAINING ROWS · 1 MB PARQUET · 35 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from GLM 5.2 (glm-5.2). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task domain. This is an actively growing… See the full description on the dataset page: https://huggingface.co/datasets/rashdan1/glm-5.2-coding-and-debugging-traces.tabulartext-generation1K<n<10K0 likes159 downloads3mo agoHugging Face17codingninja /fineweb-edutabular10M<n<100M0 likes140 downloads2y agoHugging Face18TozAI /TR-coding-data TR Coding Data (v0.2) English | Türkçe Documents / Doküman: 28 Tokens / Token: 317,427 (Qwen/Qwen3-0.6B tokenizer) from datasets import load_dataset ds = load_dataset("TozAI/TR-coding-data", split="train") English Raw-text documents that explain programming in Turkish with embedded code, intended for pretraining / mid-training of language models. Each row is one self-contained document (tutorial, project walkthrough, debugging session, ...). How it… See the full description on the dataset page: https://huggingface.co/datasets/TozAI/TR-coding-data.tabulartext-generationn<1K1 likes120 downloads15d agoHugging Face19ArkhAngelLifeJiggy /glm-5.2-coding-and-debugging-traces GLM 5.2 Agent Traces 207 TRAJECTORIES · 1,821 TRAINING ROWS · 1 MB PARQUET · 35 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from GLM 5.2 (glm-5.2). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task domain. This is an actively growing… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/glm-5.2-coding-and-debugging-traces.tabulartext-generation1K<n<10K0 likes109 downloads3mo agoHugging Face20ansulev /gpt-5-6-sol-coding-and-debugging-traces Mirror: greghavens/gpt-5.6-sol-coding-and-debugging-traces Pinned snapshot / mirror of greghavens/gpt-5.6-sol-coding-and-debugging-traces, re-hosted for PROTISEC research reproducibility. Redistributed under the upstream license (cc-by-4.0) with attribution — all credit to the original author. Original author: greghavens Source dataset: greghavens/gpt-5.6-sol-coding-and-debugging-traces License: cc-by-4.0 Family: coding_traces Mode: stream Rows cached: 17939 Changes vs… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/gpt-5-6-sol-coding-and-debugging-traces.tabular10K<n<100K0 likes99 downloads2mo agoHugging Face21squeezebits /synthesized-coding-assistant-dataset Synthesized Coding Assistant Dataset Overview Coding assistants are increasingly used for real-world software engineering workflows. However, there are relatively few datasets that closely resemble how such assistants operate in practice. Many existing coding datasets are based on single-turn or single-iteration tasks, where a model receives one coding request and directly produces an answer or patch. In contrast, practical coding assistants often work through… See the full description on the dataset page: https://huggingface.co/datasets/squeezebits/synthesized-coding-assistant-dataset.tabulartext-generationn<1K0 likes95 downloads5mo agoHugging Face22ArkhAngelLifeJiggy /fable-5-coding-and-debugging-traces Claude Fable 5 Agent Traces 2,161 TRAJECTORIES · 11,235 TRAINING ROWS · 11 MB PARQUET · 656 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/fable-5-coding-and-debugging-traces.tabulartext-generation10K<n<100K0 likes87 downloads3mo agoHugging Face23siddharth0713 /fable-5-coding-and-debugging-traces Claude Fable 5 Agent Traces 2,161 TRAJECTORIES · 11,235 TRAINING ROWS · 11 MB PARQUET · 656 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task… See the full description on the dataset page: https://huggingface.co/datasets/siddharth0713/fable-5-coding-and-debugging-traces.tabulartext-generation10K<n<100K0 likes85 downloads3mo agoHugging Face24moehamid /fable-5-coding-and-debugging-traces Claude Fable 5 Agent Traces 2,374 TRAJECTORIES · 12,448 TRAINING ROWS · 14 MB PARQUET · 662 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Claude Fable 5 (anthropic/claude-fable-5). The category and row-share tables below describe the actual mix seen during training rather than assuming a particular task… See the full description on the dataset page: https://huggingface.co/datasets/moehamid/fable-5-coding-and-debugging-traces.tabulartext-generation10K<n<100K0 likes83 downloads2mo agoHugging Face25Kendamarron /magpie-python-coding-instruction-62k-qwen2.5-bakeneko-32b-instruct magpie-python-coding-instruction-62k-qwen2.5-bakeneko-32b-instruct rinna/qwen2.5-bakeneko-32b-instructを用いたMagpieで生成した合成Instructionデータセットです。 なお、計算リソースの問題からoutputの品質評価は行っていません。 ご利用の際はご注意ください。 作成手順 rinna/qwen2.5-bakeneko-32b-instruct-awqを用いたMagpieで"instruction"を生成(magpie_systemの値をシステムプロンプトとして使用) rinna/qwen2.5-bakeneko-32b-instruct-awqを用いて"instruction"の言語、タスクの種類、難易度、品質を評価 languageがja以外、もしくは品質がpoor/very poorのレコードを削除 rinna/qwen2.5-bakeneko-32b-instruct-awqを用いて応答を"output"として生成 tabular10K<n<100K0 likes66 downloads2y agoHugging Face26eitans-pu /cl-coding-inputtabular10K<n<100K0 likes60 downloads8mo agoHugging Face27weqweasdas /henry700k_rm_no_coding_and_mathtabular100K<n<1M0 likes58 downloads2y agoHugging Face28Distillio /kimi-k3-coding-and-debugging-traces Kimi K3 Coding, Tool Use & Instruction Following Traces 582 TRAJECTORIES · 3,956 TRAINING ROWS · 3 MB PARQUET · 72 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/Distillio/kimi-k3-coding-and-debugging-traces.tabulartext-generation1K<n<10K0 likes55 downloads2mo agoHugging Face29moehamid /kimi-k3-coding-and-debugging-traces Kimi K3 Coding, Tool Use & Instruction Following Traces 601 TRAJECTORIES · 4,089 TRAINING ROWS · 3 MB PARQUET · 73 MB JSONL Generated by moonshiner — an open harness for distilling verified instruction-following, tool-use, and agentic coding traces. Behavior-preserving instruction-following, tool-use, and agent trajectories from Kimi K3 (moonshotai/kimi-k3). The category and row-share tables below describe the actual mix seen during training rather than assuming a… See the full description on the dataset page: https://huggingface.co/datasets/moehamid/kimi-k3-coding-and-debugging-traces.tabulartext-generation1K<n<10K0 likes52 downloads2mo agoHugging Face30riddickz /multitask_v3_coding_11ktabular10K<n<100K0 likes48 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.