datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
github_archive
GitHub Archive
Description
According to GitHub’s terms of service, issues and pull request descriptions—along with the their comments—inherit the license of their associated repository.
To collect this data, we used the GitHub Archive’s public BigQuery table of events to extracted all issue, pull request, and comment events since 2011 and aggregated them into threads.
The table appeared to be missing “edit” events so the text from each comment is the original from when… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive.github-repos-pythonThe Github repository retrieval source for [code-rag-bench], containing all Python files from the entire GitHub dump (in github-repos)
github_archive_filtered
GitHub Archive
Description
According to GitHub’s terms of service, issues and pull request descriptions—along with their comments—inherit the license of their associated repository.
To collect this data, we used the GitHub Archive’s public BigQuery table of events to extract all issue, pull request, and comment events since 2011 and aggregated them into threads.
The table appeared to be missing “edit” events so the text from each comment is the original from when it was… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive_filtered.github-jupyter
GitHub Jupyter Dataset
Dataset Description
The dataset was extracted from Jupyter Notebooks on BigQuery.
Licenses
Each example has the license of its associated repository. There are in total 15 licenses:
[
'mit',
'apache-2.0',
'gpl-3.0',
'gpl-2.0',
'bsd-3-clause',
'agpl-3.0',
'lgpl-3.0',
'lgpl-2.1',
'bsd-2-clause',
'cc0-1.0',
'epl-1.0',
'mpl-2.0',
'unlicense',
'isc',
'artistic-2.0'
]
github-code-fontend-lang
github-code fontend code
Dwonload
方式一
huggingface-cli download --resume-download LiXiang12/github-code-fontend-lang --include "*/*.zip" --repo-type dataset --local-dir github_code
方式二
进入Files and versions/data直接下载zip文件
数据统计
github-reposThe entire dump of GitHub repositories.
github-actionsgithub-issues
Dataset Card for GitHub Issues
Dataset Summary
GitHub Issues is a dataset consisting of GitHub issues and pull requests associated with the 🤗 Datasets repository. It is intended for educational purposes and can be used for semantic search or multilabel text classification. The contents of each GitHub issue are in English and concern the domain of datasets for NLP, computer vision, and beyond.
Supported Tasks and Leaderboards
For each of the tasks tagged… See the full description on the dataset page: https://huggingface.co/datasets/lewtun/github-issues.GitHub-code-dialogs-1.2K-v0.1
Github Codes
This is first version of dataset.
All the "user" rows were synthetically generated by Mistral-Large-Instruct-2407
GitHub-CC0
Public Domain GitHub Repositories Dataset
This dataset contains metadata and source code of 9,000 public domain (cc0 or unlicense) licensed GitHub repositories that have more than 25 stars.
The dataset was created by scraping the GitHub API and downloading the repositories, so long as they are under 100mb.
The dataset can be used for various natural language processing and software engineering tasks, such as code summarization, code generation, code search, code analysis, etc.… See the full description on the dataset page: https://huggingface.co/datasets/KoalaAI/GitHub-CC0.1c_githubgithub-issuesgithub-issuesgithub-issuesgithub-embeddings-doygithub-issuesgithub-issueslm_code_github-eval_subsetPyTorchConference2025_GithubRepos
PyTorch Conference 2025 GitHub Repos
I created a list of every GitHub repo mentioned during PyTorch Conference 2025 and Open Source AI Week.
ai-github-ai-2026
Ai Github Ai 2026
Part of the LEGION Intelligence dataset collection.
Provider: LEGION Systems
Access: Requires approval — submit request below
Usage
from datasets import load_dataset
dataset = load_dataset("gemmozero/ai-github-ai-2026")
API Access
Real-time access via LEGION API:
curl https://api.legion-api.com/incidents
API Docs · Pro Access €29/mo
License
CC BY-NC 4.0 — Research and non-commercial use only.
Commercial use… See the full description on the dataset page: https://huggingface.co/datasets/gemmozero/ai-github-ai-2026.github-pythongithub-actions-provenance-signed-untampered
GitHub Actions Provenance (Signed, Untampered)
This dataset contains signed provenance metadata in JSONL format generated from successful and untampered CI/CD builds using GitHub Actions. The data represents clean, valid examples of what software artifact provenance should look like in secure, uncompromised environments.
📂 Dataset Structure
Format: .jsonl files (each line is a JSON object)
Source: Generated by GitHub Actions CI/CD pipelines
Signature: Signed using… See the full description on the dataset page: https://huggingface.co/datasets/vchirrav/github-actions-provenance-signed-untampered.github-pr-history
What is this dataset?
This dataset is a collection of Pull Requests that contain comments from the Accelerate.
It contains the full contextual comments as well as code suggestions that exist inside of a code review
pile_githubCoding-Agent-Github-2025-Feb
Coding Agent AI Agent Directory to Host All Coding Agent related AI Agents Web Traffic Data, Search Ranking, Community, Reviews and More.
This is the Coding Agent Dataset from pypi package "coding_agent" https://pypi.org/project/coding_agent. You can use this package to download and get statistics (forks/stars/website traffic) of AI agents on website from AI Agent Marketplace AI Agent Directory (http://www.deepnlp.org/store/ai-agent) and AI Agent Search Portal… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/Coding-Agent-Github-2025-Feb.github-code-nanochatbpe-1B
github-code-nanochatbpe-1B
GitHub Code (all-all) (from codeparrot/github-code), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536) and packaged as flat uint16 token-id .bin files for fast memmap training.
file
split
tokens
train.bin
train
1,000,000,000
val.bin
val
10,000,000
train and val are disjoint held-out partitions. Each .bin is a raw little-endian uint16 stream (no header); token count = filesize / 2, and train.meta.json / val.meta.json carry… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/github-code-nanochatbpe-1B.github-issuesannotations_creators:
other
language_creators:
crowdsourced
languages:
en-US
licenses:
other-my-license
multilinguality:
monolingual
pretty_name: HuggingFace Github Issues
size_categories:
unknown
source_datasets:
original
task_categories:
text-classification
text-retrieval
task_ids:
multi-class-classification
multi-label-classification
document-retrieval
korean_code_reviews_from_githubgithub-source-code-dataset
Github Source Code Dataset
Complete source code from Agnuxo projects.
Part of the Agnuxo Ecosystem by Francisco Angulo de Lafuente.
github-trending
GitHub Trending: топ репозиториев за неделю (2026-09-28)
Открытые данные GitHub API: топ-10 репозиториев за 7 дней.
