Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Suzhen /CodeChat CodeChat: Developer–LLM Conversations Dataset Paper: https://arxiv.org/abs/2509.10402 GitHub: https://github.com/Software-Evolution-Analytics-Lab-SEAL/CodeChat CodeChat is a large-scale dataset comprising 82,845 real-world developer–LLM conversations, containing 368,506 code snippets generated across more than 20 programming languages, derived from the WildChat (i.e., general Human-LLMs conversations dataset). The dataset enables empirical analysis of how developers… See the full description on the dataset page: https://huggingface.co/datasets/Suzhen/CodeChat.texttext-generation10K<n<100K3 likes1.2k downloads3mo agoHugging Face02BEE-spoke-data /code_contests_instruct Dataset Card for "code_contests_instruct" The deepmind/code_contests dataset formatted as markdown-instruct for text generation training. There are several different configs. Look at them. Comments: flesch_reading_ease is computed on the description col via textstat hq means that python2 (aka PYTHON in language column) is dropped, and keeps only rows with flesch_reading_ease 75 or greater min-cols drops all cols except language and text possible values for language are {'CPP'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code_contests_instruct.tabulartext-generation10M<n<100M7 likes1.1k downloads10mo agoHugging Face03Suzhen /CodeChat-V2.0 CodeChat: Developer–LLM Conversations Dataset Paper: https://arxiv.org/abs/2509.10402 GitHub: https://github.com/Software-Evolution-Analytics-Lab-SEAL/CodeChat CodeChat V2.0 is a large-scale dataset comprising 587,568 real-world developer–LLM conversations, derived from the WildChat dataset. The V2.0 statistics below match Table I of the CASCON 2026 camera-ready paper, Developer-LLM Conversations: An Empirical Study of Interactions and Generated Code Quality.… See the full description on the dataset page: https://huggingface.co/datasets/Suzhen/CodeChat-V2.0.texttext-generation100K<n<1M1 likes433 downloads11d agoHugging Face04voidful /agent-sft-stitch-zh-tts-taste-codec-chat-sample Gemma 4 E2B Taste-S multi-turn codec SFT This dataset contains 37,362 complete Traditional Chinese agent dialogues selected from voidful/agent-sft-stitch-zh-tts. It covers 229,434 synthesized speech segments, approximately 520.5 hours of audio before codec extraction. Every assistant speech segment is represented without Gemma native audio tags: <SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY> The first assistant output starts immediately with <SAY>. [SOPR]...[EOPR]… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-chat-sample.tabulartext-generation10K<n<100K0 likes380 downloads3mo agoHugging Face05louisbrulenaudet /code-commande-publique Code de la commande publique, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-commande-publique.tabulartext-generation1K<n<10K0 likes252 downloads1y agoHugging Face06codeparrot /codecomplex CodeComplex Dataset Dataset Description CodeComplex consists of 4,200 Java codes submitted to programming competitions by human programmers and their complexity labels annotated by a group of algorithm experts. How to use it You can load and iterate through the dataset with the following two lines of code: from datasets import load_dataset ds = load_dataset("codeparrot/codecomplex", split="train") print(next(iter(ds))) Data Structure… See the full description on the dataset page: https://huggingface.co/datasets/codeparrot/codecomplex.texttext-generation1K<n<10K30 likes216 downloads4y agoHugging Face07wttw /code_contest_instruct_cpptabulartext-generation1M<n<10M3 likes214 downloads2y agoHugging Face08CodedotAI /code_clippyThis dataset was generated by selecting GitHub repositories from a large collection of repositories. These repositories were collected from https://seart-ghs.si.usi.ch/ and Github portion of [The Pile](https://github.com/EleutherAI/github-downloader) (performed on July 7th, 2021). The goal of this dataset is to provide a training set for pretraining large language models on code data for helping software engineering researchers better understand their impacts on software related tasks such as autocompletion of code. The dataset is split into train, validation, and test splits. There is a version containing duplicates (209GBs compressed) and ones where exact duplicates (132GBs compressed) are removed. Contains mostly JavaScript and Python code, but other programming languages are included as well to various degrees.text-generation12 likes165 downloads4y agoHugging Face09iamtarun /code_contest_processed Dataset Card for Code Contest Processed Dataset Summary This dataset is created by processing code_contest dataset from Deepmind. It is a competitive programming dataset for machine-learning. Read more about dataset at original source. Columns Description id : unique string associated with a problem description : problem description code : one correct code for the problem language : programming language used for code test_samples : contains inputs and their… See the full description on the dataset page: https://huggingface.co/datasets/iamtarun/code_contest_processed.texttext-generation10K<n<100K3 likes165 downloads3y agoHugging Face10opencompass /CodeCompass CodeCompass: A Benchmark for Code Generation Paper: Rethinking Verification for LLM Code Generation: From Generation to Testing Description CodeCompass is a rigorous benchmark designed to evaluate the code generation capabilities of Large Language Models (LLMs). It comprises a comprehensive collection of programming problems sourced from competitive platforms, offering a standardized framework for assessing algorithmic reasoning, problem-solving, and code synthesis in a… See the full description on the dataset page: https://huggingface.co/datasets/opencompass/CodeCompass.text-generation1 likes161 downloads1y agoHugging Face11louisbrulenaudet /code-civil Code civil, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language models based… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-civil.tabulartext-generation1K<n<10K4 likes144 downloads1y agoHugging Face12louisbrulenaudet /code-collectivites-territoriales Code général des collectivités territoriales, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-collectivites-territoriales.tabulartext-generation1K<n<10K0 likes120 downloads1y agoHugging Face13louisbrulenaudet /code-consommation Code de la consommation, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-consommation.tabulartext-generation1K<n<10K0 likes99 downloads1y agoHugging Face14louisbrulenaudet /code-construction-habitation Code de la construction et de l'habitation, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-construction-habitation.tabulartext-generation1K<n<10K1 likes98 downloads1y agoHugging Face15tomhu /codecl Context Learning for Software Engineering Code • Paper Dataset Details Dataset Sources [optional] Code Generation: LeetCode Code Summary: Github Code Review: Github Patch Correctness Assessment: Defects4J Uses Evaulate LLMs on context learning ability under Software Engineering tasks. Citation @misc{hu2026cl4secontextlearningbenchmark, title={CL4SE: A Context Learning Benchmark For Software Engineering Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/tomhu/codecl.text-generation10K<n<100K1 likes84 downloads7mo agoHugging Face16kd13 /CodeChat-Instruct-v1 CodeChat-Instruct-v1 CodeChat-Instruct-v1 is a synthetic coding instruction dataset designed for supervised fine-tuning of language models on programming-related conversations. It includes diverse coding tasks such as code review, code improvement, complexity analysis, edge-case discussion, code explanation, library/API usage, refactoring guidance. The dataset is suitable for training coding assistants, educational programming tutors, and general-purpose code LLMs with strong… See the full description on the dataset page: https://huggingface.co/datasets/kd13/CodeChat-Instruct-v1.texttext-generation10K<n<100K1 likes78 downloads4mo agoHugging Face17code-critic-model /critic-sft-cwm-only critic-sft-cwm-only The CWM-only SFT corpus from Steer, Don't Solve: Training Small Critic Models for Large Code Agents. It trains Qwen3-8B-Critic-SFT-CWM-only, the CWM-only arm of the corpus ablation in Table 3. Each record is one critique point: a CWM-32B trajectory up to some step, followed by the critique that Claude Opus 4.6 wrote for it under the paper's high-level prompt. The main corpus, critic-sft-cwm-qwen, is this set plus 1,915 records from Qwen3-Next trajectories.… See the full description on the dataset page: https://huggingface.co/datasets/code-critic-model/critic-sft-cwm-only.texttext-generation1K<n<10K0 likes69 downloads1mo agoHugging Face18iamtarun /code_contest_python3_alpaca Dataset Card for Code Contest Processed Dataset Summary This dataset contains coding contest questions and their solution written in Python3. This dataset is created by processing code_contest dataset from Deepmind. It is a competitive programming dataset for machine-learning. Read more about dataset at original source. Columns Description id : unique string associated with a problem description : problem description code : one correct code for the problem… See the full description on the dataset page: https://huggingface.co/datasets/iamtarun/code_contest_python3_alpaca.textquestion-answering1K<n<10K8 likes66 downloads3y agoHugging Face19G2RPO-A /Code-CoT G2RPO-A Code-CoT Code-CoT contains 1,086 Python programming problems paired with generated chain-of-thought (CoT) trajectories and executable test cases. It is the coding guidance dataset released with G²RPO-A: Guided Group Relative Policy Optimization with Adaptive Guidance. Paper (ACL 2026) · Code · Source dataset · Math-Curriculum-1K Dataset at a glance Property Value Task Python code generation and reasoning Intended use Guidance for G²RPO-A and… See the full description on the dataset page: https://huggingface.co/datasets/G2RPO-A/Code-CoT.texttext-generation1K<n<10K0 likes66 downloads26d agoHugging Face20code-critic-model /critic-sft-cwm-only-detailed-prompt critic-sft-cwm-only-detailed-prompt The detailed-prompt SFT corpus from Steer, Don't Solve: Training Small Critic Models for Large Code Agents. It trains Qwen3-8B-Critic-SFT-Detailed-Prompt, the comparison arm of the prompt ablation in Table 4. Each record is one critique point: a CWM-32B trajectory up to some step, followed by the critique that Claude Opus 4.6 wrote for it. The difference from critic-sft-cwm-only is the teacher prompt. Here the teacher used the detailed prompt… See the full description on the dataset page: https://huggingface.co/datasets/code-critic-model/critic-sft-cwm-only-detailed-prompt.texttext-generation1K<n<10K0 likes61 downloads1mo agoHugging Face21code-critic-model /critic-sft-cwm-qwen critic-sft-cwm-qwen The main SFT corpus from Steer, Don't Solve: Training Small Critic Models for Large Code Agents. It trains both Qwen3-8B-Critic-SFT and Qwen3-4B-Critic-SFT. Each record is one critique point: a coding agent's trajectory up to some step, followed by the critique that Claude Opus 4.6 wrote for it. The teacher was prompted with the paper's high-level prompt, which asks for error detection and one or two sentences of guidance and forbids code and commands in the… See the full description on the dataset page: https://huggingface.co/datasets/code-critic-model/critic-sft-cwm-qwen.texttext-generation1K<n<10K0 likes61 downloads1mo agoHugging Face22codecainecowboy /Nemotron-Personas-Korea Nemotron-Personas-Korea 우리나라 실제 분포에 기반한 합성 페르소나를 위한 복합 AI 시스템 A compound AI approach to personas grounded in real-world distributions 데이터셋 개요 (Overview) Nemotron-Personas-Korea는 대한민국의 실제 인구통계학적·지리적·성격 특성 분포를 기반으로 합성된 오픈소스 페르소나 데이터셋(CC BY 4.0)으로, 우리나라 인구의 다양성과 특성을 폭넓게 반영하도록 설계되었습니다. 이는 최초의 대규모 우리말 페르소나 데이터셋이며, 이름, 성별, 나이, 혼인 상태, 교육 수준, 직업, 거주 지역 등의 속성을 실제 대한민국 통계청(KOSIS), 대법원, 국민건강보험공단, 농촌경제연구원, NAVER Cloud 통계 자료를 기반으로 합성하였습니다. Nemotron-Personas-Korea는… See the full description on the dataset page: https://huggingface.co/datasets/codecainecowboy/Nemotron-Personas-Korea.imagetext-generation1M<n<10M0 likes54 downloads5mo agoHugging Face23skonml /code-contract-repair APIContractRepair APIContractRepair is a provenance-tracked instruction-tuning dataset for software engineers and code-model researchers who need contract-faithful, minimal repairs with tests that distinguish a broken implementation from its fix. Magicoder-OSS-Instruct-75K supplies real function-identifier seeds, but it does not provide these documented contracts, deliberately buggy implementations, minimal corrected implementations, or paired regression tests. This release… See the full description on the dataset page: https://huggingface.co/datasets/skonml/code-contract-repair.texttext-generation1K<n<10K0 likes53 downloads2mo agoHugging Face24SSK-DNB /CodeContests_OnlyPythonDataThis dataset is based on the "code_contests" dataset by DeepMind, licensed under CC BY 4.0. Original dataset available at: https://huggingface.co/datasets/deepmind/code_contests このデータセットは、deepmind/code_contestsからAtCoderのデータのみを抽出し、Pythonで正解として提出された最初のコードを取得して格納したものです。これにより、教師あり学習に適した形に整えられています。 This dataset is created by extracting only the data from AtCoder within deepmind/code_contests and retrieving the first Python code submission marked as correct. As a result, it has been formatted to be… See the full description on the dataset page: https://huggingface.co/datasets/SSK-DNB/CodeContests_OnlyPythonData.texttext-generation1K<n<10K1 likes49 downloads2y agoHugging Face25DCAgent /code-contests-noblock code-contests-noblock Harbor-format competitive-programming RL tasks (8,728 tasks) derived from CodeContests. Each task is a Harbor task binary (gzip tar) stored in tasks.parquet with columns path (str, <task_id>.tar.gz) and task_binary (binary). Each tar contains instruction.md, task.toml, environment/Dockerfile, and a tests/ verifier (test.sh, test_state.py, test_data.json). v2 (current) — verifier reward-recording fix Resolves a silent reward-distribution… See the full description on the dataset page: https://huggingface.co/datasets/DCAgent/code-contests-noblock.texttext-generation1K<n<10K0 likes45 downloads4mo agoHugging Face26codecainecowboy /GLM-5.1-Reasoning-1M-Cleaned GLM-5.1-Reasoning-1M-Cleaned GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields. This release was prepared from the original dataset published by Kassadin88. Summary Teacher model in the data: GLM-5.1… See the full description on the dataset page: https://huggingface.co/datasets/codecainecowboy/GLM-5.1-Reasoning-1M-Cleaned.texttext-generation100K<n<1M0 likes45 downloads5mo agoHugging Face27HlebYakhnitski /codecomplex CodeComplex Dataset Dataset Description CodeComplex consists of 4,200 Java codes submitted to programming competitions by human programmers and their complexity labels annotated by a group of algorithm experts. How to use it You can load and iterate through the dataset with the following two lines of code: from datasets import load_dataset ds = load_dataset("codeparrot/codecomplex", split="train") print(next(iter(ds))) Data Structure… See the full description on the dataset page: https://huggingface.co/datasets/HlebYakhnitski/codecomplex.texttext-generation1K<n<10K0 likes39 downloads10mo agoHugging Face28eo-belyakin /codecontests-textbooks-dp-v1This dataset is a synthetic collection designed for algorithmic problem-solving, particularly in the dynamic programming domain. It is inspired by problems from the DeepMind/code_contests dataset, ensuring authenticity and relevance to competitive programming and algorithmic challenges. The dataset includes detailed problem statements, input-output specifications, constraints, and illustrative test cases. Each example mirrors real-world scenarios, providing not only the problem but also… See the full description on the dataset page: https://huggingface.co/datasets/eo-belyakin/codecontests-textbooks-dp-v1.texttext-generation1K<n<10K2 likes36 downloads2y agoHugging Face29code-critic-model /critic-sft-qwen-only critic-sft-qwen-only The Qwen-only SFT corpus from Steer, Don't Solve: Training Small Critic Models for Large Code Agents. It trains Qwen3-8B-Critic-SFT-Qwen-only and Qwen3-4B-Critic-SFT-Qwen-only, the Qwen-only arms of the corpus ablation in Table 3. Each record is one critique point: a Qwen3-Next-80B-A3B-Instruct trajectory up to some step, followed by the critique that Claude Opus 4.6 wrote for it under the paper's high-level prompt. The main corpus, critic-sft-cwm-qwen, is… See the full description on the dataset page: https://huggingface.co/datasets/code-critic-model/critic-sft-qwen-only.texttext-generation1K<n<10K0 likes36 downloads1mo agoHugging Face30louisbrulenaudet /code-communes Code des communes, non-instruct (2025-09-20) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language models… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-communes.tabulartext-generationn<1K0 likes35 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.