datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
github-codeThe GitHub Code dataest consists of 115M code files from GitHub in 32 programming languages with 60 extensions totalling in 1TB of text data. The dataset was created from the GitHub dataset on BiqQuery.github-code-cleanThe GitHub Code clean dataset in a more filtered version of codeparrot/github-code dataset, it consists of 115M code files from GitHub in 32 programming languages with 60 extensions totaling in almost 1TB of text data.open-github-major-reposπ AdhyanshVerma's Open GitHub Major Repos
An elite, curated collection of GitHub commit metadata from the world's most influential technology companies: Microsoft, Google, Meta, and Intel.
π Introduction
Welcome to AdhyanshVerma's Open GitHub Major Repos dataset. This dataset focuses exclusively on high-impact, industry-standard repositories maintained by the world's leading technology giants.
It utilizes the Lazy Pointer Pattern: instead of bloating your storage⦠See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/open-github-major-repos.github-code-2025-language-split
π Source Data & Attribution
This dataset is a processed derivative of nick007x/github-code-2025.
Origination
The original data was aggregated by nick007x from public GitHub repositories. We have retained the original content, file paths, and metadata while restructuring the format for easier consumption by language-specific models.
Processing Steps
To create this dataset, we performed the following processing on the source data:
Language⦠See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/github-code-2025-language-split.gitee-code
Gitee Code Dataset
Dataset Description
This dataset was compiled from code repositories hosted on Gitee, China's largest code hosting platform and a leading alternative to GitHub in the Chinese developer community. Gitee is widely used by Chinese developers, enterprises, and open-source projects, making this dataset particularly valuable for training code models with strong Chinese language understanding and Chinese coding conventions.
Dataset Summary⦠See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/gitee-code.D-GITT-RTE7000-2023
README
See the common Readme
code_clippy_githubThe Code Clippy dataset consists of various public codebases from GitHub in 22 programming languages with 23 extensions totalling about 16 TB of data when uncompressed. The dataset was created from the public GitHub dataset on Google BiqQuery.codeparrot-github-code-10GThis is data is derived from the Codeparrot Dataset by taking the first 10GB of text from each language, and splitting it into individual configs. This results in a download size of about 3GB per language.
Sample usage:
from datasets import load_dataset
dataset = load_dataset("ruediste/codeparrot-github-code-10G", "java")
List of Languages:
languages = {
'HTML': 'html',
'Java': 'java',
'JavaScript': 'js',
'CSS': 'css',
'C#': 'cs',
'TypeScript': 'ts',
"Batchfile":β¦ See the full description on the dataset page: https://huggingface.co/datasets/ruediste/codeparrot-github-code-10G.github_archive
GitHub Archive
Description
According to GitHubβs terms of service, issues and pull request descriptionsβalong with the their commentsβinherit the license of their associated repository.
To collect this data, we used the GitHub Archiveβs public BigQuery table of events to extracted all issue, pull request, and comment events since 2011 and aggregated them into threads.
The table appeared to be missing βeditβ events so the text from each comment is the original from whenβ¦ See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive.llama.cpp_AlgMor24_github
Ξ©FFFΞ£LLIa β’ llama.cpp β’ AlgMor24
βββββββ βββββββββββββββββββββββββββ βββ βββ ββββββ
ββββββββββββββββββββββββββββββββββββ βββ βββββββββββ
βββ βββββββββ ββββββ ββββββ βββ βββ βββββββββββ
βββ βββββββββ ββββββ ββββββ βββ βββ βββββββββββ
ββββββββββββ βββ ββββββββββββββββββββββββββββββ βββ
βββββββ βββ βββ ββββββββββββββββββββββββββββββ βββ
High-Performance LLM / VLM Inference & Autonomous Agentic Ecosystem⦠See the full description on the dataset page: https://huggingface.co/datasets/Brunobkr/llama.cpp_AlgMor24_github.D-GITT-RTE7000-2022
README
See the common Readme
github-top-code
GitHub Top Developer Source Code
A curated dataset of 1.3M+ source code files from GitHub's top ranked developers (2015-2025).
This dataset is based on the top ranked developers from this dataset: https://huggingface.co/datasets/ronantakizawa/github-top-developers
Dataset Summary
1.3M+ source code files from repositories across ~4,700 unique developers
80+ programming languages included (Python, JavaScript, TypeScript, Rust, Go, C/C++, Java, and more)
Source code only ββ¦ See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-top-code.open-github
OpenGitHub
What is it?
This dataset contains every public event on GitHub: every push, pull request, issue, star, fork, code review, release, and discussion across all public repositories. GitHub is the world's largest software development platform, home to over 200 million repositories and the daily work of tens of millions of developers, from individual open-source contributors to the engineering teams behind the most widely used software on earth.
The archive currently⦠See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-github.github-code-2025StockChina-Minute
A-Share Minute-Level Historical Data
Dataset Description
This dataset contains minute-level trading data for Chinese A-share stocks from 2005 to 2023, covering 5267 stocks with complete historical trading records.
Data Format
Each CSV file corresponds to one stock and contains the following fields:
Field
Description
open
Opening price
close
Closing price
high
Highest price
low
Lowest price
volume
Trading volume
money
Trading amount
avg⦠See the full description on the dataset page: https://huggingface.co/datasets/jobs-git/StockChina-Minute.gitskills
GitSkills: A Dataset of Agent Skills on GitHub
Paper (arXiv:2608.10906) Β·
Sample repository Β·
Zenodo DOI: 10.5281/zenodo.21875637
An agent skill is a folder containing a SKILL.md file with instructions for
a language-model agent, optionally accompanied by scripts and reference
files. The agent loads the skill when it judges that a task matches the
skill description. Anthropic introduced the format in October 2025 as an
open specification. Nine months later, skill files in the⦠See the full description on the dataset page: https://huggingface.co/datasets/mvaccargiu/gitskills.python-github-codegithub-code-2025
π GitHub Code 2025: The Clean Code Manifesto
A meticulously curated dataset of 1.5M+ repositories representing both quality and innovation in 2025's code ecosystem
π The Philosophy
Quality Over Quantity, Purpose Over Volume
In an era of data abundance, we present a dataset built on radical curation. Every file, every repository, every byte has been carefully selected to represent the signal in the noise of open-source development.
π― What This Dataset Isβ¦ See the full description on the dataset page: https://huggingface.co/datasets/hybridfree/github-code-2025.javascript-github-codetrellis500k-github-archives-7D-GITT-RTE7000-2021
D-GITT_RTE 7000 Nodes Dataset
We, at OpenSynth/D-GITT, are proud to announce that RTE released open-source the complete French electrical network data. D-GITT stands for Detailed Grid Inner Topology Time-series
This first datasets provides a series of snapshots of the French transmission electricity network in node-breaker topology, with a temporal granularity of 5 minutes, covering the three-year period from January 2021 to December 2023.
You'll find three datasets per year due to⦠See the full description on the dataset page: https://huggingface.co/datasets/OpenSynth/D-GITT-RTE7000-2021.github-repos-pythonThe Github repository retrieval source for [code-rag-bench], containing all Python files from the entire GitHub dump (in github-repos)
GIT-SCRAPED
π SKT-NRS / GIT-SCRAPED
This repository is dedicated to hosting structural, curated, and diverse datasetsβincluding GitHub roadmaps,and Roadmaps.sh Sites system architectures, technical diagrams, and mass scraped assets.
Our ultimate mission is to fuel the development of next-generation Sovereign Indian Intelligence base models with high-fidelity, production-grade text-image structures.
π Repository Structure
All the raw and structured crawled data is⦠See the full description on the dataset page: https://huggingface.co/datasets/SKT-NRS/GIT-SCRAPED.git-commits
Themis-Git-Commits
Overview
Themis-Git-Commits is a large-scale dataset of single-file code commits mined from permissively licensed GitHub repositories via the BigQuery GitHub public dataset. The SQL query restricts to repositories under permissive open-source licenses only (MIT, Apache-2.0, BSD-2/3-Clause, ISC, CC0-1.0, EPL-1.0, MPL-2.0, Unlicense, AGPL-3.0, LGPL-2.1, Artistic-2.0). The BigQuery snapshot used contains commits up to early 2022 ββ¦ See the full description on the dataset page: https://huggingface.co/datasets/project-themis/git-commits.GITQA-Aug-Legacygithub_archive_filtered
GitHub Archive
Description
According to GitHubβs terms of service, issues and pull request descriptionsβalong with their commentsβinherit the license of their associated repository.
To collect this data, we used the GitHub Archiveβs public BigQuery table of events to extract all issue, pull request, and comment events since 2011 and aggregated them into threads.
The table appeared to be missing βeditβ events so the text from each comment is the original from when it wasβ¦ See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive_filtered.github-commits-diff-dedup-pjjs-april
Deduplicated Commits
Deduplicated based on diff:
content = '\n'.join(difflib.unified_diff(
old_content.splitlines(keepends=True),
new_content.splitlines(keepends=True),
n=5
))
Parameters:
Minimum ngram size: 5
MinHash ngram size: 5
MinHash threshold: 0.8
github_all_lang_filteredspark-hermes-rounds
Spark-Hermes rounds
Training data from a competition (Bittensor SN74) in which miners submit prose strategies and a validator
runs one pinned agent (Hermes) and model (Qwen3.8-27B through r0055, SparkHermes-27B-v1 from era e1) against
every strategy inside a sealed sandbox. The tasks
are bug fixes drawn from SWE-bench/SWE-smith (MIT).
Only the crowned strategy of each round is exported. SFT rows are its episodes that a withheld test suite
verified fully. DPO pairs a⦠See the full description on the dataset page: https://huggingface.co/datasets/gittensor-model-hub/spark-hermes-rounds.github-codereview
Code Review Dataset
A large-scale dataset of the best human-written code reviews from top GitHub repositories.
Each row captures a moment where a human code reviewer left an inline comment on a pull request, and the author subsequently modified the code in response.
The dataset also includes negative examples β code from the same PRs that passed review without comments β to help models learn when code is acceptable.
This provides a natural signal for training models to:
Generate⦠See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-codereview.
