datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-github
OpenGitHub
What is it?
This dataset contains every public event on GitHub: every push, pull request, issue, star, fork, code review, release, and discussion across all public repositories. GitHub is the world's largest software development platform, home to over 200 million repositories and the daily work of tens of millions of developers, from individual open-source contributors to the engineering teams behind the most widely used software on earth.
The archive currently… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-github.github-codereview
Code Review Dataset
A large-scale dataset of the best human-written code reviews from top GitHub repositories.
Each row captures a moment where a human code reviewer left an inline comment on a pull request, and the author subsequently modified the code in response.
The dataset also includes negative examples — code from the same PRs that passed review without comments — to help models learn when code is acceptable.
This provides a natural signal for training models to:
Generate… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-codereview.github-readme-retrieval-multilingual_beirThis is a copy of https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual_beir.github-issuesgithub-diffs-dedupedgithub-repos-metadata-40M
📊 Metadata for 40 million GitHub repositories
A cleaned, analysis-ready dataset with per-repository statistics aggregated from GH Archive events: stars, forks, pull requests, open issues, visibility, language signals, and more. Column names mirror the GH Archive / GitHub API semantics where possible.
GitHub repo: https://github.com/ibragim-bad/github-repos-metadata-40M
Source: GH Archive (public GitHub event stream).
✅ Projects with complementary ideas
GitHub Repo… See the full description on the dataset page: https://huggingface.co/datasets/ibragim-bad/github-repos-metadata-40M.GitHub-Agentic-PR-Dataset
GitHub Agentic PR Dataset
A large-scale dataset of ~2 million GitHub Pull Requests authored by AI coding agents (Claude Code, Cursor, GitHub Copilot, Devin) and human developers, complete with commits, file-level diffs, patches, bug-fix classification, review-depth metrics, and test-quality measures.
The GitHub Agentic PR Dataset is a research-grade corpus for studying how AI coding agents contribute to real-world open-source software, and how their pull requests compare to… See the full description on the dataset page: https://huggingface.co/datasets/mabujadallah/GitHub-Agentic-PR-Dataset.github-code-haskell-file
Dataset Card for "github-code-haskell-file"
Rows: 339k
Download Size: 806M
This dataset is extracted from github-code-clean.
Each row also contains attribute values for my personal analysis project.
12.6% (43k) of the rows have cyclomatic complexity and LOC valued at -1 because homplexity failed in parsing the row's uncommented_code.
the-stack-github-issues
Dataset Description
This dataset contains conversations from GitHub issues and Pull Requests. Each conversation is comprised of a series of events, such as opening an issue, creating a comment,
or closing the issue, and includes the author's username, text, action, and identifiers such as the issue ID and number.
The dataset, which is mostly in English, has a total size of 54GB and 30.9M files.
Dataset Structure
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-github-issues.github-pr-agent-trajectories
Description
Reconstructed multi-step agent trajectories for resolving public GitHub PRs: a plan of subtasks and per-subtask subagents with scoped files and solution diffs.
Derived dataset. Source material retains its original per-item licence (see source/repo columns); treat as other / mixed. Provided as-is.
Usage
from datasets import load_dataset
ds = load_dataset("PotatoHD/github-pr-agent-trajectories")
requests-github-issuespython-github-codegithub-embedded
github-embedded
A bunch of python code, gotten from Github, along with a json dump of their ast's and a vector embedding of the code using openai's text-embedding-3-small.
deprecated-github-code-haskell-function
Dataset Card for "github-code-haskell-function"
Rows: 3.26M
Download Size: 1.17GB
This dataset is extracted from github-code-haskell-file.
Each row has 3 flavors of the same function:
uncommented_code: Includes the function and its closest signature.
function_only_code: Includes the function only.
full_code: Includes the function and its closest signature and comment.
The heuristic for finding the closest signature and comment follows: If the immediate previous neighbor of the… See the full description on the dataset page: https://huggingface.co/datasets/blastwind/deprecated-github-code-haskell-function.github-issues-negatives-maxgithub-codereview-dataset
Github-Codereview-Dataset
Made with ❤️ using 🦥 Unsloth Studio
github-codereview-dataset was generated with Unsloth Recipe Studio. It contains 10,000 generated records.
🚀 Quick Start
from datasets import load_dataset
# Load the main dataset
dataset = load_dataset("manishsaini1/github-codereview-dataset", "data", split="train")
df = dataset.to_pandas()
📊 Dataset Summary
📈 Records: 10,000
📋 Columns: 23
📋 Schema & Statistics… See the full description on the dataset page: https://huggingface.co/datasets/manishsaini1/github-codereview-dataset.github-pullrequestsgithub-issuesPulled from GitHub Issues respository using the REST API, the dataset consists of 3,384 Hugging Face GitHub issues (excluded the pull requests) with comments. It is intended for evaluation purpose and free to use if you're following the Chapter 5 - The Dataset Library of Hugging Face LLM Course.
github-issues
Dataset for GitHub Issues
Based on tutorial Creating your own dataset of the LLM course.
Differences from lewtun/github-issues:
Data for the last 5 years has been added.
Duplicates have been removed.
Actividad-Repositorios-GitHub-2025
Dataset Card: Actividad-Repositorios-GitHub-2025
DOI: 10.57967/hf/10760 · Código y datos crudos: github.com/ymyxm/Actividad-Repositorios-Github-2025
Autores
Shengkai Zhu
Kaihao Wang
Yixuan Miren Mei
Yixuan Lu Guo
Dataset Summary
Dataset sobre la actividad diaria de 46 repositorios open-source populares de GitHub durante 2025, que combina texto y temporalidad: para cada repositorio y día UTC reúne los recuentos de actividad (pushes, commits, issues… See the full description on the dataset page: https://huggingface.co/datasets/Kaihao06/Actividad-Repositorios-GitHub-2025.github-java-corpus
github-java-corpus
Summary
This dataset contains Java source-code text samples prepared for pretraining.
Repository
TheFinAI/github-java-corpus
Required Columns
Source: dataset name
Date: year
Text: the pure text of each sample
Token_count: the token count computed with tiktoken
Schema
Source (string)
Date (int32)
Text (string)
Token_count (int32)
Construction
The dataset was built from streamed… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/github-java-corpus.rust-github-issues
Dataset Card for "rust-github-issues"
More Information needed
github-pytorch-issues
Dataset Card for github-pytorch-issues
Dataset Summary
This dataset is a curated collection of GitHub issues from the PyTorch repository. Each entry includes the issue title, body, user, state, labels, comments, and other relevant fields that are useful for tasks such as text classification, semantic search, and question answering.
Supported Tasks and Leaderboards
The dataset supports the following tasks:
Open-domain Question Answering: Given a user query… See the full description on the dataset page: https://huggingface.co/datasets/mayankpuvvala/github-pytorch-issues.github-issuesgithub-repos-metadata-ge3
GitHub All Repositories Metadata Dataset (>= 3 Stars)
A comprehensive metadata dataset covering 5,289,726 public GitHub repositories with 3 or more stars (>= 3) spanning the history of GitHub from 2008 to 2026.
Data Recency & Snapshot Notice
[!NOTE]
Snapshot Methodology: This dataset combines a comprehensive historical base archive (up to mid-2022) with continuous periodic crawler snapshots (2023 through 2026).
Star Counts & Metrics: Star counts and repository… See the full description on the dataset page: https://huggingface.co/datasets/Mieaz/github-repos-metadata-ge3.Transformers-Github-Issuesgithub_latest
Latest GitHub Repositories
You could always access the latest Github repos via this dataset.
We update the dataset weekly, on every Sunday. So the dataset always provides the latest Github repos from the last week.
The current dataset on main branch contains the latest Github Repos submitted from 2024-08-26 to 2024-09-02.
The data collection is conducted on 2024-09-09.
Use the dataset via:
ds = datasets.load_dataset('RealTimeData/github_latest')
Previsou versions
You… See the full description on the dataset page: https://huggingface.co/datasets/RealTimeData/github_latest.lean-github-biggithub-issues-vul-detection-post🛡️ GitHub Issues Vulnerability Detection
A benchmark dataset for automating the detection of code vulnerabilities by analyzing GitHub Issues.
Curated to validate the findings in Automating the Detection of Code Vulnerabilities by Analyzing GitHub Issues (2025).
📖 Dataset Description
GitHub Issues Vulnerability Detection is a specialized dataset designed to evaluate the feasibility of identifying software vulnerabilities early by analyzing textual discussions in GitHub… See the full description on the dataset page: https://huggingface.co/datasets/DanCip/github-issues-vul-detection-post.github-repo-embeddings
GitHub Repo Embeddings (Dataset)
This dataset contains:
GitHub repository embeddings learned from star co-occurrence.
Raw data for training such embeddings (2016 - 2025 years)
It is generated by the same pipeline as this repo and is intended for offline
analysis, research, and downstream search/indexing.
See Demo which uses trained embeddings
Summary
Source: GitHub Archive (BigQuery) WatchEvent + repo metadata.
Signal: repositories starred together by the same… See the full description on the dataset page: https://huggingface.co/datasets/Puzer/github-repo-embeddings.
