Team Ai
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SHSLab /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLab/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.tabulartext-generation10M<n<100M5 likes5.3k downloads1mo agoHugging Face02Manusagents /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🌌 Omni-Frontier Distillation SFT The Definitive Evolution of Open-Source Distillation & Human-Crafted Expertise Repository: Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection "The most comprehensive multi‑domain SFT corpus ever assembled — fusing 6.86 million cleaned distillation samples with 9.14 million human‑crafted expert examples across medical, cybersecurity, chemical, robotics, humanities, and more. 16 million… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.texttext-generation10M<n<100M7 likes1.8k downloads2mo agoHugging Face03Manusagents /Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.tabulartext-generation10M<n<100M1 likes440 downloads1mo agoHugging Face04harrytyp /ai-coding-plan-prices AI coding plan prices (tokens and requests per dollar) What the flat rate AI coding subscriptions cost, and what you actually get for the money: published quotas, per model spending caps, included models, and the resulting tokens and requests per dollar paid. This is the dataset behind vibeplan.cc. Snapshot of this revision: 2026-10-06 10:22 UTC, 80 plans from 26 providers, 53 of them directly comparable, built from 30 official sources. Disclosure split: 51 disclosed, 21… See the full description on the dataset page: https://huggingface.co/datasets/harrytyp/ai-coding-plan-prices.textquestion-answeringn<1K0 likes314 downloads4d agoHugging Face05ethanker /agentic_coding_dataset Agentic Coding Dataset This dataset is a compilation of various coding and instruction-following datasets, designed to train agentic coding models. Sources This dataset aggregates samples from the following sources: CodeAlpaca-20k Instruction-following coding tasks. Evol-CodeAlpaca-v1 Complex evolved coding instructions (WizardCoder style). Code Review Instruct Python code review, critique, and revision examples. APPS (Automated Programming Progress Standard)… See the full description on the dataset page: https://huggingface.co/datasets/ethanker/agentic_coding_dataset.texttext-generation100K<n<1M7 likes100 downloads11mo agoHugging Face06Kicaulah /dataset-coding Kicaulah AI — Dataset Coding (Generation, Debugging & Architecture) 📖 Description A production-grade Programming, Code Generation, Debugging, and Software Architecture dataset in natural English. Communicated through the voice of experienced senior software engineers—pragmatic, focused on root causes, and strictly free of generic AI boilerplate. Covers 10 essential software engineering domains: Python Debugging & Internal Gotchas (Mutable defaults, pandas, GIL)… See the full description on the dataset page: https://huggingface.co/datasets/Kicaulah/dataset-coding.texttext-generation1K<n<10K0 likes88 downloads11d agoHugging Face07himanshunakrani9 /mimo-coding-synthetic-5k MiMo Coding Synthetic 5.4K MiMo Coding Synthetic 5.4K is a purely synthetic coding instruction dataset generated with Xiaomi MiMo mimo-v2.5-pro. It contains 5,411 validated examples across programming languages, coding task types, and difficulty levels. The dataset is provided in two formats: A canonical rich JSONL format with metadata and labels. An OpenAI chat messages JSONL format for supervised fine-tuning pipelines. The generation run used 20 parallel workers for roughly… See the full description on the dataset page: https://huggingface.co/datasets/himanshunakrani9/mimo-coding-synthetic-5k.texttext-generation10K<n<100K0 likes72 downloads4mo agoHugging Face08Voidreaper2026 /coding-master-dataset Coding Master Dataset Overview A large-scale coding instruction-tuning dataset in ShareGPT conversational format, assembled from multiple open sources and deduplicated. Records: 766,987 Format: JSONL / ShareGPT License: Apache 2.0 Sources CodeX-2M-Thinking (430,542 records) python-code-dataset-500k (559,515 records) StackPulse high-quality subset (20,205 records) CodeFeedback-Filtered-Instruction (156,525 records) secure_programming_dpo (4,656… See the full description on the dataset page: https://huggingface.co/datasets/Voidreaper2026/coding-master-dataset.texttext-generation100K<n<1M3 likes61 downloads3mo agoHugging Face09smirki /Agentic-Coding-Tessa Agentic Coding Dataset for Tessa A comprehensive dataset for training coding agents with tool-use, reasoning, and software engineering capabilities. Dataset Composition This dataset combines multiple high-quality sources: hermes_reasoning (20.0%): Tool-use and reasoning dataset - interstellarninja/hermes_reasoning_tool_use search_arena (15.0%): Search and retrieval tasks - lmarena-ai/search-arena-24k arena_human_pref (15.0%): Human preference data for alignment -… See the full description on the dataset page: https://huggingface.co/datasets/smirki/Agentic-Coding-Tessa.texttext-generation10K<n<100K13 likes52 downloads1y agoHugging Face10cogbuji /MrGrammaticalOntology_clinical_coding Mr. Grammatical Ontology: Clinical Coding This dataset was created from a motivation to train Medical Large Language Models for improved fluency in clinical coding, as measurable by MedConceptsQA, an open-source medical coding evaluation benchmark designed to evaluate the understanding and reasoning capabilities of LLMs on medical concepts. It was extracted from the Centers for Medicare & Medicaid Services' International Classification of Diseases, Tenth Revision, Clinical… See the full description on the dataset page: https://huggingface.co/datasets/cogbuji/MrGrammaticalOntology_clinical_coding.textquestion-answering100K<n<1M2 likes49 downloads2y agoHugging Face11OpceanAI /sota-codingtexttext-generation100K<n<1M1 likes43 downloads5mo agoHugging Face12asnelt /visual-cortex-coding-qatextquestion-answering100K<n<1M0 likes41 downloads10mo agoHugging Face13chongpangnasilemak /icd10pcs-coding-mcq ICD-10-PCS Coding MCQ 405 multiple-choice items on ICD-10-PCS inpatient procedure coding — whether a model can build a seven-character procedure code from documentation it is handed: root operation selection, the seven character axes, Index→Tables verification, approach, device and qualifier values, and the Official Guidelines. Labels are what GPT-5.6-sol ruled they are. Items were written by Claude and adjudicated by GPT-5.6-sol; where the two disagreed, the adjudicator's… See the full description on the dataset page: https://huggingface.co/datasets/chongpangnasilemak/icd10pcs-coding-mcq.textquestion-answeringn<1K0 likes27 downloads2mo agoHugging Face14genbench-iitp /coding-variant GenBench CoCG QA Dataset Multi-hop genetic reasoning QA items generated from GenBench's knowledge graph (Ensembl, ClinVar, VEP, BioGRID, STRING, Reactome, UniProt, GO, SIGNOR, OmniPath, KEGG, DisGeNET, OpenTargets, PubTator3, GTEx, and more), built for CoCG (Co-Evolving Confidence Graph) agent training. 2513 items across 11 task types. Task types task_type count coding_variant 53 conservation_reasoning 246 counterfactual 246 disease_reasoning 246… See the full description on the dataset page: https://huggingface.co/datasets/genbench-iitp/coding-variant.tabularquestion-answering1K<n<10K1 likes26 downloads2mo agoHugging Face15chongpangnasilemak /icd10cm-coding-mcq ICD-10-CM Coding MCQ 403 multiple-choice items on ICD-10-CM diagnosis coding — whether a model can apply the classification's conventions to documentation it is handed: Excludes1 and Excludes2 notes, 7th-character selection, placeholder X, laterality, Index→Tabular verification, combination codes, specificity. Labels are what GPT-5.6-sol ruled they are. Items were written by Claude and adjudicated by GPT-5.6-sol; where the two disagreed, the adjudicator's ruling settled the… See the full description on the dataset page: https://huggingface.co/datasets/chongpangnasilemak/icd10cm-coding-mcq.textquestion-answeringn<1K0 likes22 downloads2mo agoHugging Face16mustavinsu /coding-model-rendered-qa Rendered QA Dataset: Code & Text (700K) Instruction-tuning dataset with optional rendered images for vision-language models. Sources Source Samples Has Context Image OpenCoder Stage 2 436K educational_instruct only InstructCoder 108K Yes (code input) OpenOrca 200K No (text-only) Schema Column Type Description prompt string Instruction/question prompt_image Image? Rendered prompt (optional) context string? Code context… See the full description on the dataset page: https://huggingface.co/datasets/mustavinsu/coding-model-rendered-qa.imagequestion-answering100K<n<1M0 likes18 downloads9mo agoHugging Face17GG13412 /CodingQuestionDatabaseCodeLlamaThe questions, responses, and topics were generated with the codellama 7b model. They may be empty data points due to the AI generation. Main.json: CodeLlama 7b moneywordmath.json: pplx-7b-chat (Math Word Problems) textquestion-answeringn<1K0 likes12 downloads2y agoHugging Face18thunder-research-group /SNU_Thunder-synthetic-codinggated Dataset Card for SNU Thunder Synthetic Coding Dataset Summary This dataset was used as part of the post-training corpus for SnuLLM(to_fill). This dataset consists of Korean and English question-answer pairs. Questions are sourced from publicly available datasets, and answers were generated using open large language models (Exaone 3.5, LLaMA 3.3, Qwen 2.5). It is intended for research and non-commercial use. Supported Tasks Tasks: Python coding Languages… See the full description on the dataset page: https://huggingface.co/datasets/thunder-research-group/SNU_Thunder-synthetic-coding.textquestion-answering1K<n<10K2 likes12 downloads1y agoHugging Face19genbench-iitp /genbench-coding-qagated GenBench CoCG QA Dataset Multi-hop genetic reasoning QA items generated from GenBench's knowledge graph (Ensembl, ClinVar, VEP, BioGRID, STRING, Reactome, UniProt, GO, SIGNOR, OmniPath, KEGG, DisGeNET, OpenTargets, PubTator3, GTEx, and more), built for CoCG (Co-Evolving Confidence Graph) agent training. 8159 items across 11 task types. Task types task_type count coding_variant 159 conservation_reasoning 800 counterfactual 800 disease_reasoning 800… See the full description on the dataset page: https://huggingface.co/datasets/genbench-iitp/genbench-coding-qa.tabularquestion-answering1K<n<10K1 likes7 downloads2mo agoHugging Face20SHSLabs /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLabs/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.tabulartext-generation10M<n<100M0 likes9h agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.