datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
context-ucurve-coding-agents
Context U-curve: 36 coding-agent runs under six context-clearing policies
How often should an LLM coding agent's context be cleared? This dataset holds every run behind the report
"Clear Every Third Task: A Measured U-Curve in the Context Economy of Coding Agents"
(Evgenii Arsentev, 2026; corrected version 1.2, DOI 10.5281/zenodo.22759217; version 1.0: DOI 10.5281/zenodo.22699668).
A fixed suite of twelve programming tasks was run under six session-length policies — a fresh… See the full description on the dataset page: https://huggingface.co/datasets/arsentev-ai/context-ucurve-coding-agents.tech-debt-ai-coding
Debt Behind the AI Boom — Replication Data
Data for the paper:
Debt Behind the AI Boom: A Large-Scale Empirical Study of AI-Generated Code in the Wild
Yue Liu, Ratnadira Widyasari, Yanjie Zhao, Ivana Clairine Irsan, Junkai Chen, David Lo
📄 arXiv:2603.28592 · 💻 Code: github.com/yueyueL/tech-debt-ai-coding
We mined 302.6K AI-authored commits from 6,299 GitHub repositories across five
AI coding assistants (GitHub Copilot, Claude, Cursor, Gemini, Devin), ran static
analysis… See the full description on the dataset page: https://huggingface.co/datasets/yueyuel/tech-debt-ai-coding.GenBench_non_coding
Task types by split
task_type
train
test
disease_reasoning
535
94
hallucination_detection
668
132
interaction_propagation
640
160
mechanistic_explanation
610
190
path_traversal
573
227
regulatory_reasoning
672
128
Schema
Each item has:
id, task_type, pipeline (coding_variant/noncoding_regulatory), difficulty
question, answer, choices (MCQ options, when applicable)
context -- either a templated chain narration, or (if llm_rewrite was… See the full description on the dataset page: https://huggingface.co/datasets/iit-patna-cse-ai/GenBench_non_coding.GenBench_coding
Task types by split
task_type
train
test
coding_variant
101
8
conservation_reasoning
624
132
counterfactual
598
202
disease_reasoning
523
93
hallucination_detection
627
173
interaction_propagation
651
149
path_traversal
263
151
structural_effect
677
112
Schema
Each item has:
id, task_type, pipeline (coding_variant/noncoding_regulatory), difficulty
question, answer, choices (MCQ options, when applicable)
context -- either a templated… See the full description on the dataset page: https://huggingface.co/datasets/iit-patna-cse-ai/GenBench_coding.ai-coding-agent-pricing-and-capability-dataset
AI Coding Agent Pricing and Capability Dataset
A source-backed market-intelligence dataset for comparing AI coding agents and developer workflow agents across pricing, workflow support, release signals, repository activity, integrations, and public capability claims.
Each row represents one observed market signal tied to an official product page, official documentation page, official pricing page, public GitHub repository, or public GitHub release note. The dataset is built for… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/ai-coding-agent-pricing-and-capability-dataset.anhem
