datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
github-codeThe GitHub Code dataest consists of 115M code files from GitHub in 32 programming languages with 60 extensions totalling in 1TB of text data. The dataset was created from the GitHub dataset on BiqQuery.github-code-cleanThe GitHub Code clean dataset in a more filtered version of codeparrot/github-code dataset, it consists of 115M code files from GitHub in 32 programming languages with 60 extensions totaling in almost 1TB of text data.github-code-2025-language-split
๐ Source Data & Attribution
This dataset is a processed derivative of nick007x/github-code-2025.
Origination
The original data was aggregated by nick007x from public GitHub repositories. We have retained the original content, file paths, and metadata while restructuring the format for easier consumption by language-specific models.
Processing Steps
To create this dataset, we performed the following processing on the source data:
Languageโฆ See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/github-code-2025-language-split.open-github-major-repos๐ AdhyanshVerma's Open GitHub Major Repos
An elite, curated collection of GitHub commit metadata from the world's most influential technology companies: Microsoft, Google, Meta, and Intel.
๐ Introduction
Welcome to AdhyanshVerma's Open GitHub Major Repos dataset. This dataset focuses exclusively on high-impact, industry-standard repositories maintained by the world's leading technology giants.
It utilizes the Lazy Pointer Pattern: instead of bloating your storageโฆ See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/open-github-major-repos.codeparrot-github-code-10GThis is data is derived from the Codeparrot Dataset by taking the first 10GB of text from each language, and splitting it into individual configs. This results in a download size of about 3GB per language.
Sample usage:
from datasets import load_dataset
dataset = load_dataset("ruediste/codeparrot-github-code-10G", "java")
List of Languages:
languages = {
'HTML': 'html',
'Java': 'java',
'JavaScript': 'js',
'CSS': 'css',
'C#': 'cs',
'TypeScript': 'ts',
"Batchfile":โฆ See the full description on the dataset page: https://huggingface.co/datasets/ruediste/codeparrot-github-code-10G.code_clippy_githubThe Code Clippy dataset consists of various public codebases from GitHub in 22 programming languages with 23 extensions totalling about 16 TB of data when uncompressed. The dataset was created from the public GitHub dataset on Google BiqQuery.github_archive
GitHub Archive
Description
According to GitHubโs terms of service, issues and pull request descriptionsโalong with the their commentsโinherit the license of their associated repository.
To collect this data, we used the GitHub Archiveโs public BigQuery table of events to extracted all issue, pull request, and comment events since 2011 and aggregated them into threads.
The table appeared to be missing โeditโ events so the text from each comment is the original from whenโฆ See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive.llama.cpp_AlgMor24_github
ฮฉFFFฮฃLLIa โข llama.cpp โข AlgMor24
โโโโโโโ โโโโโโโโโโโโโโโโโโโโโโโโโโโ โโโ โโโ โโโโโโ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โโโ โโโโโโโโโโโ
โโโ โโโโโโโโโ โโโโโโ โโโโโโ โโโ โโโ โโโโโโโโโโโ
โโโ โโโโโโโโโ โโโโโโ โโโโโโ โโโ โโโ โโโโโโโโโโโ
โโโโโโโโโโโโ โโโ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โโโ
โโโโโโโ โโโ โโโ โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ โโโ
High-Performance LLM / VLM Inference & Autonomous Agentic Ecosystemโฆ See the full description on the dataset page: https://huggingface.co/datasets/Brunobkr/llama.cpp_AlgMor24_github.github-code-2025
๐ GitHub Code 2025: The Clean Code Manifesto
A meticulously curated dataset of 1.5M+ repositories representing both quality and innovation in 2025's code ecosystem
๐ The Philosophy
Quality Over Quantity, Purpose Over Volume
In an era of data abundance, we present a dataset built on radical curation. Every file, every repository, every byte has been carefully selected to represent the signal in the noise of open-source development.
๐ฏ What This Dataset Isโฆ See the full description on the dataset page: https://huggingface.co/datasets/hybridfree/github-code-2025.github-top-code
GitHub Top Developer Source Code
A curated dataset of 1.3M+ source code files from GitHub's top ranked developers (2015-2025).
This dataset is based on the top ranked developers from this dataset: https://huggingface.co/datasets/ronantakizawa/github-top-developers
Dataset Summary
1.3M+ source code files from repositories across ~4,700 unique developers
80+ programming languages included (Python, JavaScript, TypeScript, Rust, Go, C/C++, Java, and more)
Source code only โโฆ See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-top-code.open-github
OpenGitHub
What is it?
This dataset contains every public event on GitHub: every push, pull request, issue, star, fork, code review, release, and discussion across all public repositories. GitHub is the world's largest software development platform, home to over 200 million repositories and the daily work of tens of millions of developers, from individual open-source contributors to the engineering teams behind the most widely used software on earth.
The archive currentlyโฆ See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-github.python-github-codegithub-code-2025trellis500k-github-archives-7javascript-github-codegithub-repos-pythonThe Github repository retrieval source for [code-rag-bench], containing all Python files from the entire GitHub dump (in github-repos)
github-commits-diff-dedup-pjjs-april
Deduplicated Commits
Deduplicated based on diff:
content = '\n'.join(difflib.unified_diff(
old_content.splitlines(keepends=True),
new_content.splitlines(keepends=True),
n=5
))
Parameters:
Minimum ngram size: 5
MinHash ngram size: 5
MinHash threshold: 0.8
github_all_lang_filteredgithub_archive_filtered
GitHub Archive
Description
According to GitHubโs terms of service, issues and pull request descriptionsโalong with their commentsโinherit the license of their associated repository.
To collect this data, we used the GitHub Archiveโs public BigQuery table of events to extract all issue, pull request, and comment events since 2011 and aggregated them into threads.
The table appeared to be missing โeditโ events so the text from each comment is the original from when it wasโฆ See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive_filtered.github-codereview
Code Review Dataset
A large-scale dataset of the best human-written code reviews from top GitHub repositories.
Each row captures a moment where a human code reviewer left an inline comment on a pull request, and the author subsequently modified the code in response.
The dataset also includes negative examples โ code from the same PRs that passed review without comments โ to help models learn when code is acceptable.
This provides a natural signal for training models to:
Generateโฆ See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-codereview.github-jupyter
GitHub Jupyter Dataset
Dataset Description
The dataset was extracted from Jupyter Notebooks on BigQuery.
Licenses
Each example has the license of its associated repository. There are in total 15 licenses:
[
'mit',
'apache-2.0',
'gpl-3.0',
'gpl-2.0',
'bsd-3-clause',
'agpl-3.0',
'lgpl-3.0',
'lgpl-2.1',
'bsd-2-clause',
'cc0-1.0',
'epl-1.0',
'mpl-2.0',
'unlicense',
'isc',
'artistic-2.0'
]
the_pile_github
Dataset Card for The Pile GitHub
Dataset Summary
This is the GitHub subset of EleutherAi/The Pile dataset and contains GitHub repositories. The programming languages are identified using the guesslang library. A total of 54 programming languages are included in the dataset.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
The following languages are covered by the dataset:
'Assembly', 'Batchfile', 'C', 'C#', 'C++', 'CMake'โฆ See the full description on the dataset page: https://huggingface.co/datasets/andstor/the_pile_github.github-readme-retrieval-multilingual_beirThis is a copy of https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us atโฆ See the full description on the dataset page: https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual_beir.trellis500k-github-archives-10github-code-html-css-2github-code-html-css-1trellis500k-github-archives-9github-diffs-dedupedgithub-code-fontend-lang
github-code fontend code
Dwonload
ๆนๅผไธ
huggingface-cli download --resume-download LiXiang12/github-code-fontend-lang --include "*/*.zip" --repo-type dataset --local-dir github_code
ๆนๅผไบ
่ฟๅ
ฅFiles and versions/data็ดๆฅไธ่ฝฝzipๆไปถ
ๆฐๆฎ็ป่ฎก
github-code-permissive-sampleSampling from codeparrot/github-code under more permissive license ['mit', 'apache-2.0', 'bsd-3-clause', 'bsd-2-clause', 'cc0-1.0'].It is intended to be used for training code language classifier.
