datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
opc-fineweb-code-corpus
OpenCoder Dataset
The OpenCoder dataset is composed of the following datasets:
opc-sft-stage1: the sft data used for opencoder sft-stage1
opc-sft-stage2: the sft data used for opencoder sft-stage2
opc-annealing-corpus: the synthetic data & algorithmic corpus used for opencoder annealing
opc-fineweb-code-corpus: the code-related page recalled from fineweb <-- you are here
opc-fineweb-math-corpus: the math-related page recalled from finewebrefineCode-code-corpus-meta: the meta-data… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/opc-fineweb-code-corpus.RefineCode-code-corpus-metaThis dataset consists of meta information (including the repository name and file path) of the raw code data from RefineCode. You can collect those files referring to this metadata and reproduce RefineCode!
Note: Currently, we have uploaded the meta data covered by The Stack V2 (About 50% file volume). Due to complex legal considerations, we are unable to provide the complete source code currently. We are working hard to make the remaining part available.
RefineCode is a high-quality… See the full description on the dataset page: https://huggingface.co/datasets/OpenCoder-LLM/RefineCode-code-corpus-meta.Saudilang-Code-Switch-Corpus
SCC - Saudilang Code-Switch Corpus
The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), published the "SCC" dataset, which stands for "Saudilang Code-Switch Corpus”.
This dataset contains a transcription of general conversations taken from a YouTube podcast "Thmanyah" that has been transcribed by the National Center for Artificial Intelligence in SDAIA. The data features three episodes covering different domains: investment… See the full description on the dataset page: https://huggingface.co/datasets/SDAIANCAI/Saudilang-Code-Switch-Corpus.powershell-code-corpus
PowerShell Code Corpus
A dataset of PowerShell scripts collected from public GitHub repositories.
Content
4,632 files from popular PowerShell repositories on GitHub
Sourced from repos with the highest star counts (quality signal)
Each record contains the raw script text plus metadata
Fields
Field
Description
source
Always github
repo
owner/repo slug
repo_url
Full GitHub URL
path
File path within the repo
language
Always PowerShell
license… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/powershell-code-corpus.code-quality-corpus
CatQualia code-quality corpus — semantic smell classes with before/after fixes
39,383 rows · 25,541,792 bytes · JSON Lines, one object per line.
What this is
Real code smells paired with the fix: a smell_class that names the semantic problem (not just the syntax), the original lines, the corrected lines, the file and line it came from, and a rationale explaining why the original was wrong. Useful for code-review or repair training where the label has to say what… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/code-quality-corpus.gleam-code-corpus
Gleam Code Corpus
A structured, attributed corpus of 22,581 Gleam source files collected from 1,528 GitHub repositories for continued pre-training and code generation research.
Overview
This dataset contains .gleam source files from the Gleam programming language ecosystem. Every file includes full attribution — repo owner, name, URL, license, star count, and Hex.pm package metadata where available.
What this is for:
Continued pre-training (CPT) of code-focused… See the full description on the dataset page: https://huggingface.co/datasets/kasuboski/gleam-code-corpus.ruby-code-corpus
Ruby Code Corpus
A large corpus of Ruby source code collected from public GitHub repositories.
Content
294,074 files from 3,920+ Ruby repositories on GitHub
Sourced from repos ranked by star count (quality signal)
Filtered: removed files under 200 bytes (trivial/empty files)
Only permissive licenses: MIT, Apache-2.0, BSD, ISC, Ruby
Fields
Field
Description
source
Always github
repo
owner/repo slug
repo_url
Full GitHub URL
path
File path… See the full description on the dataset page: https://huggingface.co/datasets/NickIBrody/ruby-code-corpus.
