datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-github
OpenGitHub
What is it?
This dataset contains every public event on GitHub: every push, pull request, issue, star, fork, code review, release, and discussion across all public repositories. GitHub is the world's largest software development platform, home to over 200 million repositories and the daily work of tens of millions of developers, from individual open-source contributors to the engineering teams behind the most widely used software on earth.
The archive currently… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-github.gitskills
GitSkills: A Dataset of Agent Skills on GitHub
Paper (arXiv:2608.10906) ·
Sample repository ·
Zenodo DOI: 10.5281/zenodo.21875637
An agent skill is a folder containing a SKILL.md file with instructions for
a language-model agent, optionally accompanied by scripts and reference
files. The agent loads the skill when it judges that a task matches the
skill description. Anthropic introduced the format in October 2025 as an
open specification. Nine months later, skill files in the… See the full description on the dataset page: https://huggingface.co/datasets/mvaccargiu/gitskills.spark-hermes-rounds
Spark-Hermes rounds
Training data from a competition (Bittensor SN74) in which miners submit prose strategies and a validator
runs one pinned agent (Hermes) and model (Qwen3.8-27B through r0055, SparkHermes-27B-v1 from era e1) against
every strategy inside a sealed sandbox. The tasks
are bug fixes drawn from SWE-bench/SWE-smith (MIT).
Only the crowned strategy of each round is exported. SFT rows are its episodes that a withheld test suite
verified fully. DPO pairs a… See the full description on the dataset page: https://huggingface.co/datasets/gittensor-model-hub/spark-hermes-rounds.github-codereview
Code Review Dataset
A large-scale dataset of the best human-written code reviews from top GitHub repositories.
Each row captures a moment where a human code reviewer left an inline comment on a pull request, and the author subsequently modified the code in response.
The dataset also includes negative examples — code from the same PRs that passed review without comments — to help models learn when code is acceptable.
This provides a natural signal for training models to:
Generate… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-codereview.github-readme-retrieval-multilingual_beirThis is a copy of https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual_beir.GitScholar
GitScholar: arXiv AI papers, their citations, and their GitHub footprint
GitScholar is a relational, fully timestamped dataset linking AI arXiv papers
to their citation history on Semantic Scholar and to the GitHub repositories that
reference them. It is built for studying, and predicting, how research papers gain
attention over time: every row carries the date on which it became true, so the state
of the whole graph can be reconstructed as of any day between 1991 and… See the full description on the dataset page: https://huggingface.co/datasets/huawei-csl/GitScholar.github-diffs-dedupedopen-github-issues
OpenGitHub Issues
What is it?
The full development metadata of 7 public GitHub repositories, fetched from the GitHub REST API and GraphQL API, converted to Parquet and hosted here for easy access.
Right now the archive has 6.0M rows across 8 tables in 699.6 MB of Zstd-compressed Parquet. Every issue, pull request, comment, code review, timeline event, file change, and CI status check is stored as a separate table you can load individually or query together.
This… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-github-issues.github-issuesgithub-issues
Dataset Card for GitHub Issues
Dataset Summary
GitHub Issues is a dataset consisting of GitHub issues and pull requests associated with the 🤗 Datasets repository. It is intended for educational purposes and can be used for semantic search or multilabel text classification. The contents of each GitHub issue are in English and concern the domain of datasets for NLP, computer vision, and beyond.
Supported Tasks and Leaderboards
For each of the tasks tagged… See the full description on the dataset page: https://huggingface.co/datasets/lewtun/github-issues.github-actionsgithub-repos-metadata-40M
📊 Metadata for 40 million GitHub repositories
A cleaned, analysis-ready dataset with per-repository statistics aggregated from GH Archive events: stars, forks, pull requests, open issues, visibility, language signals, and more. Column names mirror the GH Archive / GitHub API semantics where possible.
GitHub repo: https://github.com/ibragim-bad/github-repos-metadata-40M
Source: GH Archive (public GitHub event stream).
✅ Projects with complementary ideas
GitHub Repo… See the full description on the dataset page: https://huggingface.co/datasets/ibragim-bad/github-repos-metadata-40M.trellis500k-github-archives-9GitHub-Agentic-PR-Dataset
GitHub Agentic PR Dataset
A large-scale dataset of ~2 million GitHub Pull Requests authored by AI coding agents (Claude Code, Cursor, GitHub Copilot, Devin) and human developers, complete with commits, file-level diffs, patches, bug-fix classification, review-depth metrics, and test-quality measures.
The GitHub Agentic PR Dataset is a research-grade corpus for studying how AI coding agents contribute to real-world open-source software, and how their pull requests compare to… See the full description on the dataset page: https://huggingface.co/datasets/mabujadallah/GitHub-Agentic-PR-Dataset.the-stack-github-issues
Dataset Description
This dataset contains conversations from GitHub issues and Pull Requests. Each conversation is comprised of a series of events, such as opening an issue, creating a comment,
or closing the issue, and includes the author's username, text, action, and identifiers such as the issue ID and number.
The dataset, which is mostly in English, has a total size of 54GB and 30.9M files.
Dataset Structure
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-github-issues.github-code-haskell-file
Dataset Card for "github-code-haskell-file"
Rows: 339k
Download Size: 806M
This dataset is extracted from github-code-clean.
Each row also contains attribute values for my personal analysis project.
12.6% (43k) of the rows have cyclomatic complexity and LOC valued at -1 because homplexity failed in parsing the row's uncommented_code.
trellis500k-github-archives-10random-small-github-repositories
random-small-github-repositories
A collection of 5,613 small-to-medium open-source GitHub repositories, packaged as zipped archives alongside a metadata CSV. Intended as a seed dataset for code retrieval, context engineering, and SWE-bench-style dataset construction tasks.
Contents
seed_small_repos.csv — metadata for each repo (owner, repo_name, stars, license, repo_hash)
repos-zipped/ — one .zip per repo, named {repo_hash}.zip
unzipper.py - unzipping python… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/random-small-github-repositories.random-python-github-repositories
random-python-github-repositories
A collection of 1650 open-source Python GitHub repositories, packaged as zipped archives alongside a metadata CSV. Intended as a seed dataset for code retrieval, context engineering, and SWE-bench-style dataset construction tasks. All repos contain 250+ .py files.
Contents
repos_meta_data.csv — metadata for each repo (owner, repo_name, stars, license, py_file_count, alpha_hash)
repos-zipped/ — one .zip per repo, named… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/random-python-github-repositories.Axis_Mundu_BDBV_2026
Axis Mũndũ: Bundibugyo Ebolavirus (BDBV) Structural Research 2026
🕊️ Project Overview
This dataset contains the formal archival assets of the Axis Mũndũ research initiative, conducted under the Honia-Heal Afrika Initiative (HAI). The project focuses on the biomathematical auditing of the Bundibugyo Ebolavirus (BDBV) Glycoprotein using the ℤ₉ Parity Law.
Core Research Objectives
Structural Enforcement: Identifying mathematical 'Locked' states in… See the full description on the dataset page: https://huggingface.co/datasets/gitauwakungu/Axis_Mundu_BDBV_2026.github-issuespython-github-codegithub-repo-enumerationThis dataset was generated from GHArchive's Google BigQuery table.
It contains a list of every public repo (~380,000,000) committed to from January 2016 up to August 2024, as well as the number of unique contributors and
totals of the amounts of various events on those repositories in that time period.
This is useless on its own, but represents more than a few hours of effort and roughly $8 worth of cloud processing,
so I figured I would save the next person to try this some effort.
github-issuesgithub_2025_highgithub-pr-agent-trajectories
Description
Reconstructed multi-step agent trajectories for resolving public GitHub PRs: a plan of subtasks and per-subtask subagents with scoped files and solution diffs.
Derived dataset. Source material retains its original per-item licence (see source/repo columns); treat as other / mixed. Provided as-is.
Usage
from datasets import load_dataset
ds = load_dataset("PotatoHD/github-pr-agent-trajectories")
Gitruck-LUT-15K
Gitruck LUT 15K
Gitruck LUT 15K is an authorized collection of .cube color lookup tables
prepared for LUT retrieval, similarity, classification, color-science analysis,
and generative modeling. The release contains viewable Parquet metadata,
diagnostic before/after previews, normalized tensors, and lossless raw archives.
Release summary
Item
Count
Source records
15,792
Unique raw byte streams
14,021
Unique normalized LUTs
13,365
Canonical train… See the full description on the dataset page: https://huggingface.co/datasets/Hocassian/Gitruck-LUT-15K.github-top-projects
GitHub Trending Projects (2013-2025)
A comprehensive dataset of 423,098 GitHub trending repository entries spanning 12+ years (August 2013 - November 2025), scraped from Wayback Machine snapshots of GitHub's trending page.
🎯 Dataset Overview
This dataset captures the evolution of GitHub's trending repositories over time, providing insights into:
Software development trends across programming languages and domains
Popular open-source projects and their trending patterns… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-top-projects.requests-github-issuesBhagavad-Gita_Dataset
Srimad Bhagavad Gita Dataset
A parallel corpus of the Srimad Bhagavad Gita containing original verses in Sanskrit (sa), with Hindi (hi) and English (en) translations.This dataset is suitable for translation, text generation, and feature extraction tasks.
Dataset Details
Based on: This dataset is based on a book called 'Srimad Bhagavad Gita' kept at Central Archaelogical Library, New Delhi
Original Book Source (IGNCA): https://ignca.gov.in/Asi_data/279.pdf… See the full description on the dataset page: https://huggingface.co/datasets/JDhruv14/Bhagavad-Gita_Dataset.
