datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
opc-fineweb-code-corpus
OpenCoder Dataset
The OpenCoder dataset is composed of the following datasets:
opc-sft-stage1: the sft data used for opencoder sft-stage1
opc-sft-stage2: the sft data used for opencoder sft-stage2
opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing
opc-fineweb-code-corpus: the code-related page recalled from fineweb <-- you are here
opc-fineweb-math-corpus: the math-related page recalled from finewebrefineCode-code-corpus-meta: the meta-data… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-fineweb-code-corpus.clean-code-corpus-v1RefineCode-code-corpus-metaThis dataset consists of meta information (including the repository name and file path) of the raw code data from RefineCode. You can collect those files referring to this metadata and reproduce RefineCode!
Note: Currently, we have uploaded the meta data covered by The Stack V2 (About 50% file volume). Due to complex legal considerations, we are unable to provide the complete source code currently. We are working hard to make the remaining part available.
RefineCode is a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/RefineCode-code-corpus-meta.High-Quality-Code
High-Quality-Code: Synthetic + Real (MAXIMUM CODE)
A massive, high-quality code dataset built with maximum code philosophy – as much code as possible.
Components
Synthetic syntax-correction dataset – 5M+ examples across 33 languages (original code_syntax_dataset_1GB.csv)
Real high-quality code from GitHub – 500 top-starred repositories – BOTH zips and extracted source
Current Status: IN PROGRESS
Target: 500 repos
Currently uploaded: 12 extracted… See the full description on the dataset page: https://huggingface.co/datasets/Corpus-NZ/High-Quality-Code.code_docstring_corpusHF version of Edinburgh-NLP's Code docstrings corpus
1gpu-llm-pretraining-corpus-15b-en-it-code
1GPU LLM Pretraining Corpus 15B EN-IT-CODE
1gpu-llm-pretraining-corpus-15b-en-it-code is the canonical document-level pretraining corpus used for the 1GPU LLM family.
It was built for training language models from scratch on a mixture of English, Italian and source code.
This Hugging Face release contains the clean, deduplicated, document-level corpus. It is intentionally published before tokenization and packing so that the training representation can be deterministically… See the full description on the dataset page: https://huggingface.co/datasets/nazdef/1gpu-llm-pretraining-corpus-15b-en-it-code.sage-code-corpus-v2verl-code-corpus-track-a-file-split
archit11/verl-code-corpus-track-a-file-split
Repository-specific code corpus extracted from the verl project and split by file for training/evaluation.
What is in this dataset
Source corpus: data/code_corpus_verl
Total files: 214
Train files: 172
Validation files: 21
Test files: 21
File type filter: .py
Split mode: file (file-level holdout)
Each row has:
file_name: flattened source file name
text: full file contents
Training context
This dataset was used… See the full description on the dataset page: https://huggingface.co/datasets/archit11/verl-code-corpus-track-a-file-split.PleIAs_common_corpus_code_classificationBurmese-English-Code-Mixed-Corpus
🇲🇲 Burmese-English Code-Mixed Corpus ꒰ 1,111 Rows ꒱
A high-quality, human-curated dataset of code-mixed Burmese and English sentences, specifically designed for Natural Language Processing (NLP) and Machine Learning (ML) research.
Dataset Details
Organization: DatarrX
Creator: Khant Sint Heinn (Kalix Louis)
Number of Rows: 1,111
Language: Burmese (Unicode) & English Mix
Dataset Format: .txt
License: Apache 2.0
Description
The Burmese-English Code-Mixed… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Burmese-English-Code-Mixed-Corpus.python_code_docstring_ast_corpus
Overview
This dataset contains 34,000+ rows of code-docstring-ast data along with additional metadata. Data was gathered from various Python libraries and frameworks and their
publicly available GitHub repos. This dataset was created for the purpose of training the CodeT5+ transformer on AST-enhanced code-to-doc tasks.
Sources
The dataset was gathered from various GitHub repos sampled from this repo by Vinta.
The 26 repos are:
matplotlib
pytorch
cryptography
django… See the full description on the dataset page: https://huggingface.co/datasets/Mir-2002/python_code_docstring_ast_corpus.Saudilang-Code-Switch-Corpus
SCC - Saudilang Code-Switch Corpus
The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), published the "SCC" dataset, which stands for "Saudilang Code-Switch Corpus”.
This dataset contains a transcription of general conversations taken from a YouTube podcast "Thmanyah" that has been transcribed by the National Center for Artificial Intelligence in SDAIA. The data features three episodes covering different domains: investment… See the full description on the dataset page: https://huggingface.co/datasets/SDAIANCAI/Saudilang-Code-Switch-Corpus.code-corpus-llm-training
Code Corpus for LLM Training
Manually collected from top open-source repositories across:
video/graphics editors, browsers, terminals, UI/UX, Qt/QML, Flutter, Rust, Python,
ethical hacking, system-level, game engines, web frameworks, and more.
Stats
Records: 240,378
Raw text: 2,156,908,643 chars (~2.01 GB)
Domains: 20
Domains
web_ui: 32,354 records
cpp: 29,792 records
kotlin_android: 19,476 records
ui_ux_design: 19,382 records
rust: 15,440 records
python:… See the full description on the dataset page: https://huggingface.co/datasets/krystv/code-corpus-llm-training.code-corpus
Code Video Text Data Notes
Dataset summary
Preparation notes and schema examples for Code tasks using Video Text data. Full source material is intentionally not bundled, so provenance and licensing remain explicit.
Included material
prepare.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card and usage… See the full description on the dataset page: https://huggingface.co/datasets/ashleymendoza/code-corpus.1gpu-llm-validation-corpus-en-it-code
1GPU LLM Validation Corpus (EN/IT/Code, document-level)
Validation-only corpus for the 1GPU LLM pretraining project (nazdef/1gpu-llm Collection).
This repository contains the frozen document-level validation master (T050-SUB002, COMPLETE 2026-09-24), not the packed 48k/ctx2500 training artifact. The packed validation set (~10M tokens) is derived downstream from this master and stays local.
Relationship to the training corpus
Training corpus:… See the full description on the dataset page: https://huggingface.co/datasets/nazdef/1gpu-llm-validation-corpus-en-it-code.arch-resilience-lpi-260903T1715-code-corpusSDAIANCAI-Saudilang-Code-Switch-Corpuspowershell-code-corpus
PowerShell Code Corpus
A dataset of PowerShell scripts collected from public GitHub repositories.
Content
4,632 files from popular PowerShell repositories on GitHub
Sourced from repos with the highest star counts (quality signal)
Each record contains the raw script text plus metadata
Fields
Field
Description
source
Always github
repo
owner/repo slug
repo_url
Full GitHub URL
path
File path within the repo
language
Always PowerShell
license… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/powershell-code-corpus.code-cpt-corpus
Dataset Card for code-cpt-corpus
Dataset Summary
ဒီ dataset က code-cpt-corpus အတွက် ဖန်တီးထားတာပါ။
Languages
Myanmar (my) / English (en)
Dataset Structure
Data Instances
{ "text": "နမူနာ စာသား", "label": "အညွှန်း" }
Data Fields
text: main content, label: optional.
Data Splits
Split
Files
train
data/train.jsonl
Licensing Information
ဒီ dataset က CC BY-NC… See the full description on the dataset page: https://huggingface.co/datasets/kkomyoeminaung/code-cpt-corpus.synthetic-code-corpushyperswitch-code-corpus-track-a
archit11/hyperswitch-code-corpus-track-a
Repository-specific code corpus extracted from hyperswitch and split by file for training/evaluation.
What is in this dataset
Source corpus: data/code_corpus_hyperswitch
Total files: 300
Train files: 270
Validation files: 30
Test files: 0
File type filter: .rs
Split mode: file (file-level holdout)
Each row has:
file_name: flattened source file name
text: full file contents
Training context
This dataset was used for… See the full description on the dataset page: https://huggingface.co/datasets/archit11/hyperswitch-code-corpus-track-a.construction-code-corpus-v1
Construction Code-Citation Corpus v1
Open dataset of construction-site incident narratives paired with OIICS
hazard codes (event, source, nature, body) and OSHA 29 CFR 1926 citation
candidates. Built for the Adaption Labs AutoScientist Challenge
("All Other Domains" category).
Sources
OSHA Severe Injury Reports (DOL, public domain): 2015-01 → 2025-08,
103,750 records. Each row has Final Narrative (incident text) plus
OIICS classification codes for the event… See the full description on the dataset page: https://huggingface.co/datasets/rigidhat/construction-code-corpus-v1.sage-code-corpus-v1code-agent-corpus
Synthetic Code-Agent Run Corpus (Projected)
⚠️ SYNTHETIC DATA — NOT REAL PRODUCTION TELEMETRY.
This dataset is synthetically generated for capacity planning / illustration of a
code-fixing agent platform's data pipeline. It does not contain real user data, real
repositories, real credentials, or real production captures. Identifiers and entity names
are anonymized (repo-A…repo-E) and run IDs/timestamps are fabricated.
What this is
A representative, labeled… See the full description on the dataset page: https://huggingface.co/datasets/Collaops/code-agent-corpus.mini-code-corpus
Dataset Card for "mini-code-corpus"
More Information needed
code-quality-corpus
CatQualia code-quality corpus — semantic smell classes with before/after fixes
39,383 rows · 25,541,792 bytes · JSON Lines, one object per line.
What this is
Real code smells paired with the fix: a smell_class that names the semantic problem (not just the syntax), the original lines, the corrected lines, the file and line it came from, and a rationale explaining why the original was wrong. Useful for code-review or repair training where the label has to say what… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/code-quality-corpus.gleam-code-corpus
Gleam Code Corpus
A structured, attributed corpus of 22,581 Gleam source files collected from 1,528 GitHub repositories for continued pre-training and code generation research.
Overview
This dataset contains .gleam source files from the Gleam programming language ecosystem. Every file includes full attribution — repo owner, name, URL, license, star count, and Hex.pm package metadata where available.
What this is for:
Continued pre-training (CPT) of code-focused… See the full description on the dataset page: https://huggingface.co/datasets/kasuboski/gleam-code-corpus.Code-Syntax-Expanded
Code-Syntax-Expanded
A massive, high-quality synthetic dataset for training LLMs to identify and correct syntax errors across 33 programming languages. Contains 5+ million unique examples (~1.1 GB) with English explanations – no artificial padding, no duplicate rows.
📊 Dataset Overview
Property
Value
Total rows
5,000,000+
File size
~1.1 GB (uncompressed CSV)
Languages
33
Unique templates
160+ error patterns
Format
CSV (4 columns)
License… See the full description on the dataset page: https://huggingface.co/datasets/Corpus-NZ/Code-Syntax-Expanded.verl-code-corpusopc-annealing-corpus-synth-qa-code_python_js_ts
