datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
opc-fineweb-code-corpus
OpenCoder Dataset
The OpenCoder dataset is composed of the following datasets:
opc-sft-stage1: the sft data used for opencoder sft-stage1
opc-sft-stage2: the sft data used for opencoder sft-stage2
opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing
opc-fineweb-code-corpus: the code-related page recalled from fineweb <-- you are here
opc-fineweb-math-corpus: the math-related page recalled from finewebrefineCode-code-corpus-meta: the meta-data… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-fineweb-code-corpus.clean-code-corpus-v1RefineCode-code-corpus-metaThis dataset consists of meta information (including the repository name and file path) of the raw code data from RefineCode. You can collect those files referring to this metadata and reproduce RefineCode!
Note: Currently, we have uploaded the meta data covered by The Stack V2 (About 50% file volume). Due to complex legal considerations, we are unable to provide the complete source code currently. We are working hard to make the remaining part available.
RefineCode is a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/RefineCode-code-corpus-meta.code_docstring_corpusHF version of Edinburgh-NLP's Code docstrings corpus
verl-code-corpus-track-a-file-split
archit11/verl-code-corpus-track-a-file-split
Repository-specific code corpus extracted from the verl project and split by file for training/evaluation.
What is in this dataset
Source corpus: data/code_corpus_verl
Total files: 214
Train files: 172
Validation files: 21
Test files: 21
File type filter: .py
Split mode: file (file-level holdout)
Each row has:
file_name: flattened source file name
text: full file contents
Training context
This dataset was used… See the full description on the dataset page: https://huggingface.co/datasets/archit11/verl-code-corpus-track-a-file-split.PleIAs_common_corpus_code_classificationSDAIANCAI-Saudilang-Code-Switch-Corpuscode-corpus-llm-training
Code Corpus for LLM Training
Manually collected from top open-source repositories across:
video/graphics editors, browsers, terminals, UI/UX, Qt/QML, Flutter, Rust, Python,
ethical hacking, system-level, game engines, web frameworks, and more.
Stats
Records: 240,378
Raw text: 2,156,908,643 chars (~2.01 GB)
Domains: 20
Domains
web_ui: 32,354 records
cpp: 29,792 records
kotlin_android: 19,476 records
ui_ux_design: 19,382 records
rust: 15,440 records
python:… See the full description on the dataset page: https://huggingface.co/datasets/krystv/code-corpus-llm-training.hyperswitch-code-corpus-track-a
archit11/hyperswitch-code-corpus-track-a
Repository-specific code corpus extracted from hyperswitch and split by file for training/evaluation.
What is in this dataset
Source corpus: data/code_corpus_hyperswitch
Total files: 300
Train files: 270
Validation files: 30
Test files: 0
File type filter: .rs
Split mode: file (file-level holdout)
Each row has:
file_name: flattened source file name
text: full file contents
Training context
This dataset was used for… See the full description on the dataset page: https://huggingface.co/datasets/archit11/hyperswitch-code-corpus-track-a.sage-code-corpus-v1verl-code-corpusgleam-code-corpus
Gleam Code Corpus
A structured, attributed corpus of 22,581 Gleam source files collected from 1,528 GitHub repositories for continued pre-training and code generation research.
Overview
This dataset contains .gleam source files from the Gleam programming language ecosystem. Every file includes full attribution — repo owner, name, URL, license, star count, and Hex.pm package metadata where available.
What this is for:
Continued pre-training (CPT) of code-focused… See the full description on the dataset page: https://huggingface.co/datasets/kasuboski/gleam-code-corpus.mini-code-corpus
Dataset Card for "mini-code-corpus"
More Information needed
opc-annealing-corpus-synth-qa-code_python_js_tsCode_corpus
