datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
github_archive
GitHub Archive
Description
According to GitHub’s terms of service, issues and pull request descriptions—along with the their comments—inherit the license of their associated repository.
To collect this data, we used the GitHub Archive’s public BigQuery table of events to extracted all issue, pull request, and comment events since 2011 and aggregated them into threads.
The table appeared to be missing “edit” events so the text from each comment is the original from when… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive.spark-hermes-rounds
Spark-Hermes rounds
Training data from a competition (Bittensor SN74) in which miners submit prose strategies and a validator
runs one pinned agent (Hermes) and model (Qwen3.8-27B through r0055, SparkHermes-27B-v1 from era e1) against
every strategy inside a sealed sandbox. The tasks
are bug fixes drawn from SWE-bench/SWE-smith (MIT).
Only the crowned strategy of each round is exported. SFT rows are its episodes that a withheld test suite
verified fully. DPO pairs a… See the full description on the dataset page: https://huggingface.co/datasets/gittensor-model-hub/spark-hermes-rounds.github_archive_filtered
GitHub Archive
Description
According to GitHub’s terms of service, issues and pull request descriptions—along with their comments—inherit the license of their associated repository.
To collect this data, we used the GitHub Archive’s public BigQuery table of events to extract all issue, pull request, and comment events since 2011 and aggregated them into threads.
The table appeared to be missing “edit” events so the text from each comment is the original from when it was… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive_filtered.github-jupyter
GitHub Jupyter Dataset
Dataset Description
The dataset was extracted from Jupyter Notebooks on BigQuery.
Licenses
Each example has the license of its associated repository. There are in total 15 licenses:
[
'mit',
'apache-2.0',
'gpl-3.0',
'gpl-2.0',
'bsd-3-clause',
'agpl-3.0',
'lgpl-3.0',
'lgpl-2.1',
'bsd-2-clause',
'cc0-1.0',
'epl-1.0',
'mpl-2.0',
'unlicense',
'isc',
'artistic-2.0'
]
github-code-fontend-lang
github-code fontend code
Dwonload
方式一
huggingface-cli download --resume-download LiXiang12/github-code-fontend-lang --include "*/*.zip" --repo-type dataset --local-dir github_code
方式二
进入Files and versions/data直接下载zip文件
数据统计
github-repos-pythonThe Github repository retrieval source for [code-rag-bench], containing all Python files from the entire GitHub dump (in github-repos)
github-reposThe entire dump of GitHub repositories.
github-issues
Dataset Card for GitHub Issues
Dataset Summary
GitHub Issues is a dataset consisting of GitHub issues and pull requests associated with the 🤗 Datasets repository. It is intended for educational purposes and can be used for semantic search or multilabel text classification. The contents of each GitHub issue are in English and concern the domain of datasets for NLP, computer vision, and beyond.
Supported Tasks and Leaderboards
For each of the tasks tagged… See the full description on the dataset page: https://huggingface.co/datasets/lewtun/github-issues.github-actionsGITQA-Aug-Legacygit-ops-recovery-trajectories
Git Ops Recovery Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/git-ops-recovery-trajectories.GitHub-CC0
Public Domain GitHub Repositories Dataset
This dataset contains metadata and source code of 9,000 public domain (cc0 or unlicense) licensed GitHub repositories that have more than 25 stars.
The dataset was created by scraping the GitHub API and downloading the repositories, so long as they are under 100mb.
The dataset can be used for various natural language processing and software engineering tasks, such as code summarization, code generation, code search, code analysis, etc.… See the full description on the dataset page: https://huggingface.co/datasets/KoalaAI/GitHub-CC0.1c_githubGitChameleon-2.0
GitChameleon 2.0
GitChameleon 2.0 is an AI coding benchmark comprising 328 Python-based problems conditioned on specific versions of popular libraries for scientific computing and web development. It evaluates whether AI code generation models can correctly use library APIs as they existed at a particular version — a challenging test of version-specific knowledge.
Note: This is GitChameleon 2.0, a distinct and newer work from the original GitChameleon benchmark. Please do not… See the full description on the dataset page: https://huggingface.co/datasets/cabbage972/GitChameleon-2.0.github-issuesgithub-issuesGitHub-code-dialogs-1.2K-v0.1
Github Codes
This is first version of dataset.
All the "user" rows were synthetically generated by Mistral-Large-Instruct-2407
github-issuesgit-commit-message-dtgithub-embeddings-doyPyTorchConference2025_GithubRepos
PyTorch Conference 2025 GitHub Repos
I created a list of every GitHub repo mentioned during PyTorch Conference 2025 and Open Source AI Week.
github-issuesgithub-issuesai-github-ai-2026
Ai Github Ai 2026
Part of the LEGION Intelligence dataset collection.
Provider: LEGION Systems
Access: Requires approval — submit request below
Usage
from datasets import load_dataset
dataset = load_dataset("gemmozero/ai-github-ai-2026")
API Access
Real-time access via LEGION API:
curl https://api.legion-api.com/incidents
API Docs · Pro Access €29/mo
License
CC BY-NC 4.0 — Research and non-commercial use only.
Commercial use… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-github-ai-2026.lm_code_github-eval_subsetgithub-actions-provenance-signed-untampered
GitHub Actions Provenance (Signed, Untampered)
This dataset contains signed provenance metadata in JSONL format generated from successful and untampered CI/CD builds using GitHub Actions. The data represents clean, valid examples of what software artifact provenance should look like in secure, uncompromised environments.
📂 Dataset Structure
Format: .jsonl files (each line is a JSON object)
Source: Generated by GitHub Actions CI/CD pipelines
Signature: Signed using… See the full description on the dataset page: https://huggingface.co/datasets/vchirrav/github-actions-provenance-signed-untampered.github-pythongit-history-mcq-ru
git-history-mcq-ru
805 вопросов с вариантами ответа по истории трёх открытых репозиториев
(digitable-lol/digit, digitable-lol/digitwm, digitable-lol/flang), плюс
8 672 ответа пяти моделей и 4 878 разборов этих ответов.
Вопросы на русском. Ключ каждого выведен из вывода git-команды, и сама команда
и её вывод лежат в записи — задачу можно перепроверить, не доверяя составителю.
Набор собран для одной проверки: меняют ли что-нибудь приёмы промптинга. Девять
вариантов оформления… See the full description on the dataset page: https://huggingface.co/datasets/the-homeless-god/git-history-mcq-ru.github-pr-history
What is this dataset?
This dataset is a collection of Pull Requests that contain comments from the Accelerate.
It contains the full contextual comments as well as code suggestions that exist inside of a code review
github-code-nanochatbpe-1B
github-code-nanochatbpe-1B
GitHub Code (all-all) (from codeparrot/github-code), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536) and packaged as flat uint16 token-id .bin files for fast memmap training.
file
split
tokens
train.bin
train
1,000,000,000
val.bin
val
10,000,000
train and val are disjoint held-out partitions. Each .bin is a raw little-endian uint16 stream (no header); token count = filesize / 2, and train.meta.json / val.meta.json carry… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/github-code-nanochatbpe-1B.
