datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code_contests_dp_datasetst-code-datasetst_dataset_distillation_by_qwen2.5.jsonl 是由 stgenerate 蒸馏qwen2.514b生成的
/data/complier_improve 中的是在蒸馏的过程中增加编译器在环`
st_dataset_local.jsonl 是编译验证通过
st_dpo_dataset.jsonl 是交给ai改正后编译验证通过形成DPO
其他的都是各种原因为通过编译
st_dataset_distillation_by_st_coder_clean 这个适合做反面教材
StackPulse_778K_QnA_Code_dataset
💻 StackOverflow-778K: Multi-Year Developer Q&A Dataset
Dataset Summary
A large-scale Stack Overflow question dataset containing 778,929 unique
questions sampled across 7 years (2015–2022). Each question includes the
raw HTML body, plain-text version, tags, score, view count, answer count, and
a rich set of derived features for immediate ML use.
Collected across 8 sampling runs on Feb 27 2026, deduplicated to
778,929 unique questions with only 2 duplicates removed.… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/StackPulse_778K_QnA_Code_dataset.code_net_dev_datasetcode_net_datasetcode_net_test_final_datasetbackend-code-generator-dataset
Backend Code Generation Dataset
Dataset Description
This dataset contains examples for training AI models to generate backend application code. It includes descriptions of backend requirements paired with complete, functional code implementations across multiple frameworks and programming languages.
Dataset Summary
The Backend Code Generation Dataset is designed to train models that can generate complete backend applications from natural language descriptions.… See the full description on the dataset page: https://huggingface.co/datasets/Techta/backend-code-generator-dataset.ReForm-Python2Dafny-Dataset
Re:Form Datasets
This repository contains the datasets associated with the paper "Re:Form -- Reducing Human Priors in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny".
The paper introduces a framework that leverages Reinforcement Learning (RL) within Large Language Models (LLMs) to reduce reliance on human-annotated priors for formal software verification. The datasets provided are integral for training and evaluating models in formal language… See the full description on the dataset page: https://huggingface.co/datasets/Veri-Code/ReForm-Python2Dafny-Dataset.dl4se-code-review-datasetarduino-code-dataset
Arduino Programming Dataset
Overview
The Arduino Programming Dataset is a structured dataset designed for training and evaluating AI models that generate Arduino sketches for microcontroller projects.
The dataset contains natural language instructions paired with complete Arduino programs. Each entry describes a hardware task and provides a working Arduino implementation along with additional metadata such as libraries used, wiring information, and difficulty level.
This… See the full description on the dataset page: https://huggingface.co/datasets/suneeldk/arduino-code-dataset.FibonacciAi-CODE-Fine_Tuning-DatasetDocument Version: 1.0.2 | Last Updated: 07/17/2026
Code_Datasetgithub-source-code-dataset
Github Source Code Dataset
Complete source code from Agnuxo projects.
Part of the Agnuxo Ecosystem by Francisco Angulo de Lafuente.
st-code-datasetst_dataset_distillation_by_qwen2.5.jsonl 是由 stgenerate 蒸馏qwen2.514b生成的
/data/complier_improve 中的是在蒸馏的过程中增加编译器在环`
st_dataset_local.jsonl 是编译验证通过
st_dpo_dataset.jsonl 是交给ai改正后编译验证通过形成DPO
其他的都是各种原因为通过编译
st_dataset_distillation_by_st_coder_clean 这个适合做反面教材
PerturbReason_dataset_codecode-embedding-dataset
Code-to-Doc Embedding Dataset
AI-generated code documentation pairs for training code embedding / retrieval models.
Dataset Description
Each record contains a code anchor (real production code) paired with:
positive: A rich natural-language documentation of what the code does
queries: 4 natural-language search queries a developer might use to find this code
label: A short semantic label (3-8 words)
This dataset is designed for training bi-encoder embedding models (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/archit11/code-embedding-dataset.HTML_CSS_CodeDataSet_100kcode_datasetLegacy-Code-Dataset
Legacy Codebase Dataset
Dataset Description
The Legacy Codebase Dataset is a large-scale collection of enterprise software repositories designed for training next-generation Large Language Models (LLMs), AI coding assistants, software engineering copilots, automated refactoring systems, repository understanding models, and intelligent program analysis pipelines.
The complete collection contains 405 real-world legacy codebases spanning 23 major industries… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Legacy-Code-Dataset.code-review-dataset-ja
Japanese Code Review Dataset (500 Samples)
A dataset of 500 code review pairs (buggy code + fixed code) with Japanese review comments. Designed for training and evaluating code review assistance models.
Total samples: 500
Languages: Python (220), JavaScript (136), Go (54), Rust (52), TypeScript (38)
Difficulties: Easy (135), Medium (259), Hard (106)
Bug types: Logic Error (120), Null Pointer (81), Off-by-One (80), Edge Case (77), Type Error (60), Security (42), Performance (40)… See the full description on the dataset page: https://huggingface.co/datasets/gram-chan-jp/code-review-dataset-ja.lua_code_datasetpython-mal-code-dataset
Python Code Dataset
This dataset contains extracted Python code from various repositories for fine-tuning code-generation models.
pandas-plotly-code-datasetProduct-Source-Code-DatasetDataset Description:
This dataset is a large-scale collection of coding and data, designed to support the development of advanced AI systems for code generation, program understanding, software intelligence, debugging assistance, and next-generation developer AI applications.
Additionally, this dataset can be integrated into pipelines for Supervised Fine-Tuning (SFT) and Reinforcement Learning with Human Feedback (RLHF) workflows, helping improve AI performance in code completion, automated… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Product-Source-Code-Dataset.rahul-code-datasetServer_Text_Dataset_1p06-code-review-datasetmanim_code_grpo_dataset
Primary Maths Manim GRPO Dataset
This folder contains a deterministic 1,000-prompt GRPO training dataset for the
Manim model produced by this project's SFT stage.
The central idea is simple: each row gives the model a lesson prompt, the model
generates several candidate Manim programs, and a verifier scores those candidates.
There is deliberately no assistant answer in this dataset. Adding a reference
answer would turn it back into supervised fine-tuning rather than online GRPO.… See the full description on the dataset page: https://huggingface.co/datasets/Chimanwakis/manim_code_grpo_dataset.Nursing_dataset_Arabic
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Dev-Dr-Code/Nursing_dataset_Arabic.vtex-code-dataset
