Team Ai
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Bc-AI /Vibe-coding-fable-5Basicsllycthe source dataset but i kept the fable only ones text1M<n<10M1 likes221 downloads12d agoHugging Face02harrytyp /ai-coding-plan-prices AI coding plan prices (tokens and requests per dollar) What the flat rate AI coding subscriptions cost, and what you actually get for the money: published quotas, per model spending caps, included models, and the resulting tokens and requests per dollar paid. This is the dataset behind vibeplan.cc. Snapshot of this revision: 2026-10-05 15:06 UTC, 80 plans from 26 providers, 53 of them directly comparable, built from 30 official sources. Disclosure split: 51 disclosed, 21… See the full description on the dataset page: https://huggingface.co/datasets/harrytyp/ai-coding-plan-prices.textquestion-answeringn<1K0 likes153 downloads20h agoHugging Face03arsentev-ai /context-ucurve-coding-agents Context U-curve: 36 coding-agent runs under six context-clearing policies How often should an LLM coding agent's context be cleared? This dataset holds every run behind the report "Clear Every Third Task: A Measured U-Curve in the Context Economy of Coding Agents" (Evgenii Arsentev, 2026; corrected version 1.2, DOI 10.5281/zenodo.22759217; version 1.0: DOI 10.5281/zenodo.22699668). A fixed suite of twelve programming tasks was run under six session-length policies — a fresh… See the full description on the dataset page: https://huggingface.co/datasets/arsentev-ai/context-ucurve-coding-agents.tabularn<1K0 likes106 downloads21d agoHugging Face04yueyuel /tech-debt-ai-coding Debt Behind the AI Boom — Replication Data Data for the paper: Debt Behind the AI Boom: A Large-Scale Empirical Study of AI-Generated Code in the Wild Yue Liu, Ratnadira Widyasari, Yanjie Zhao, Ivana Clairine Irsan, Junkai Chen, David Lo 📄 arXiv:2603.28592 · 💻 Code: github.com/yueyueL/tech-debt-ai-coding We mined 302.6K AI-authored commits from 6,299 GitHub repositories across five AI coding assistants (GitHub Copilot, Claude, Cursor, Gemini, Devin), ran static analysis… See the full description on the dataset page: https://huggingface.co/datasets/yueyuel/tech-debt-ai-coding.tabular1K<n<10K1 likes80 downloads3mo agoHugging Face05Aquiles-ai /Athenea-Coding-100k Athenea-Coding-100k A small dataset for code reasoning and solving code tasks. Dataset Details Size: 100,000 examples Format: Conversational (Hermes-style) Features: Chain-of-thought reasoning in <think> blocks Languages: English Use Case: Fine-tuning LLMs for code reasoning and solving code tasks. Contact More about Aquiles-ai. Aquiles-ai on GitHub. Our collections at HuggingFace. text100K<n<1M4 likes64 downloads10mo agoHugging Face06majeedkazemi /students-coding-questions-from-ai-assistant Dataset Documentation Overview This dataset contains 6776 questions asked by students from CodeAid, an AI coding assistant, during a C programming class over a 12-week semester from January to April 2023. The course did not allow the use of ChatGPT, but CodeAid was permitted. CodeAid, powered by GPT-3, did not directly disclose code solutions even when requested by students. Instead, it functioned like a teaching assistant, providing scaffolded responses in natural… See the full description on the dataset page: https://huggingface.co/datasets/majeedkazemi/students-coding-questions-from-ai-assistant.text1K<n<10K5 likes56 downloads3y agoHugging Face07iit-patna-cse-ai /GenBench_non_coding Task types by split task_type train test disease_reasoning 535 94 hallucination_detection 668 132 interaction_propagation 640 160 mechanistic_explanation 610 190 path_traversal 573 227 regulatory_reasoning 672 128 Schema Each item has: id, task_type, pipeline (coding_variant/noncoding_regulatory), difficulty question, answer, choices (MCQ options, when applicable) context -- either a templated chain narration, or (if llm_rewrite was… See the full description on the dataset page: https://huggingface.co/datasets/iit-patna-cse-ai/GenBench_non_coding.tabular1K<n<10K0 likes43 downloads1mo agoHugging Face08vnovaai /VNOVA_AI_CODING_LOGIC_TUTOR_DATASET_V1_JSONLVNOVA AI — Coding Logic Tutor Dataset (100 Scenarios) A high-quality, fully synthetic dataset designed to train LLMs that teach programming concepts, debugging logic, and problem-solving skills without executing code. Ideal for: 1-Coding tutors 2-Reasoning-focused LLMs 3-Debugging assistants 4-Educational chatbots 5-Beginner learning platforms This dataset focuses on conceptual understanding, not syntax or full solutions — making it safe and accessible for all audiences. Dataset Summary This… See the full description on the dataset page: https://huggingface.co/datasets/vnovaai/VNOVA_AI_CODING_LOGIC_TUTOR_DATASET_V1_JSONL.textn<1K0 likes41 downloads10mo agoHugging Face09kondasviktor /vcl-ai-coding-prompts VCL AI Coding Power Prompts 50 battle-tested prompts for Claude Code, Codex, Gemini CLI, and Cursor — by Vibe Coder's Life. Free catalog for vibe coders. Replace {{PLACEHOLDERS}} with your facts. Not the paid Apify Playbook prompt pack (those stay private). Load from datasets import load_dataset ds = load_dataset("kondasviktor/vcl-ai-coding-prompts", "prompts") print(ds["train"][0]["title"]) Columns Column Description id Stable id… See the full description on the dataset page: https://huggingface.co/datasets/kondasviktor/vcl-ai-coding-prompts.texttext-generationn<1K0 likes39 downloads23d agoHugging Face10joylarkin /AI-Coding-Tools Dataset Card for 2026 AI Coding Tools Last Updated: 24 May 2026 Curated By: Joy Larkin Language(s) (NLP): English License: MIT Repository: https://github.com/joylarkin/AI-Coding-Landscape Blog: https://cleverhack.com/ai-coding-landscape Dataset Description CSV file of AI Coding Tools released in 2026 & 2025. textn<1K0 likes38 downloads5mo agoHugging Face11iit-patna-cse-ai /GenBench_codinggated Task types by split task_type train test coding_variant 101 8 conservation_reasoning 624 132 counterfactual 598 202 disease_reasoning 523 93 hallucination_detection 627 173 interaction_propagation 651 149 path_traversal 263 151 structural_effect 677 112 Schema Each item has: id, task_type, pipeline (coding_variant/noncoding_regulatory), difficulty question, answer, choices (MCQ options, when applicable) context -- either a templated… See the full description on the dataset page: https://huggingface.co/datasets/iit-patna-cse-ai/GenBench_coding.tabular1K<n<10K0 likes35 downloads1mo agoHugging Face12joylarkin /AI-Coding-Models Dataset Card for 2026 AI Coding Models Last Updated: 24 May 2026 Curated By: Joy Larkin Language(s) (NLP): English License: MIT Repository: https://github.com/joylarkin/AI-Coding-Landscape Blog: https://cleverhack.com/ai-coding-landscape Dataset Description CSV file of AI Coding Models released in 2026 & 2025. textn<1K0 likes33 downloads5mo agoHugging Face13Bifrost-AI /Solana-blockchain-360-CodingThis dataset contains 360 coding and tech related samples for the Solana blockchain. Language: English Coding-languages: Rust, Typescript, & C# 214 general knowledge samples 146 coding knowledge samples Dataset Catalog: 201 Solana blockchain knowledge samples 49 Solana typescript coding samples 86 Solana rust coding samples 13 Solnet SDK knowledge samples 11 Solana c# coding samples textn<1K3 likes24 downloads1y agoHugging Face14gemmozero /ai-coding-tools-2026tabularn<1K0 likes24 downloads6d agoHugging Face15Karmane /ai-coding-agent-pricing-and-capability-dataset AI Coding Agent Pricing and Capability Dataset A source-backed market-intelligence dataset for comparing AI coding agents and developer workflow agents across pricing, workflow support, release signals, repository activity, integrations, and public capability claims. Each row represents one observed market signal tied to an official product page, official documentation page, official pricing page, public GitHub repository, or public GitHub release note. The dataset is built for… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/ai-coding-agent-pricing-and-capability-dataset.tabulartabular-classificationn<1K0 likes18 downloads4mo agoHugging Face16Ricco020 /ai-coding-assistants-benchmark-2026 AI Coding Assistants Benchmark 2026 — Methodology Dataset Independent benchmark methodology for evaluating AI coding assistants in 2026. Covers Claude Code (Anthropic), Cursor, GitHub Copilot, Windsurf (Codeium), Aider, Continue.dev, Cody (Sourcegraph), Tabnine, OpenAI Codex CLI, and Replit Agent. Methodology Test bench: 12 real-world coding tasks across Python, TypeScript, Rust, Go Benchmark: SWE-bench Verified scores per tool (cross-language) Performance:… See the full description on the dataset page: https://huggingface.co/datasets/Ricco020/ai-coding-assistants-benchmark-2026.n<1K0 likes16 downloads4mo agoHugging Face17melohomi /ai-coding-0-1text100K<n<1M0 likes5 downloads2y agoHugging Face18chronologies-ai /dolci-coding-sfttext100K<n<1M0 likes5 downloads9mo agoHugging Face19applied-ai-018 /Codinggatedtext10M<n<100M3 likes4 downloads2y agoHugging Face20AICoding91 /anhemtabularn<1K0 likes3 downloads2y agoHugging Face21Coding-With-Bashir /bwenge_ai_data1 likes3 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.