ronantakizawa/github-codereview
Code Review Dataset A large-scale dataset of the best human-written code reviews from top GitHub repositories. Each row captures a moment where a human code reviewer left an inline comment on a pull request, and the author subsequently modified the code in response. The dataset also includes negative examples — code from the same PRs that passed review without comments — to help models learn when code is acceptable. This provides a natural signal for training models to:… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-codereview.
Code Review Dataset
A large-scale dataset of the best human-written code reviews from top GitHub repositories.
Each row captures a moment where a human code reviewer left an inline comment on a pull request, and the author subsequently modified the code in response.
The dataset also includes negative examples — code from the same PRs that passed review without comments — to help models learn when code is acceptable.
This provides a natural signal for training models to:
- Generate code review comments given a code diff
- Apply review feedback by modifying code based on reviewer suggestions
- Understand code quality patterns across languages and projects
- Know when not to comment — recognizing clean code that needs no changes
Key Features
- 167K+ positive triplets from 725 top GitHub repositories
- 51K+ negative examples (~23% of dataset) of clean code labeled "No issues found."
- 37 programming languages (Python, TypeScript, Go, Rust, C++, JavaScript, C#, Java, Kotlin, Swift, and more)
- Human-only reviews: AI/bot reviewers (Copilot, linter bots, etc.) are excluded
- Quality-filtered: noise and auto-generated content removed
- Chunk-focused: ~50 lines of context around the reviewed code, not entire files
- Permissive licenses only: all source repos use MIT, Apache-2.0, BSD, or similar licenses
- Verified changes: only includes triplets where the code chunk actually changed after the review
Collection Methodology
- Repo selection: Top GitHub repos by stars with permissive licenses, sourced from ronantakizawa/github-top-projects and curated additions
- PR discovery: Paginate merged PRs, filter bot authors, fetch inline review comments
- Comment filtering: Remove bots, noise patterns, auto-generated comments, non-English text, non-code files, reply comments
- Triplet extraction: Fetch file contents at the review commit (before) and PR head (after), extract focused chunks around the comment line
- Change verification: Only keep triplets where the code chunk around the comment actually changed
- Negative extraction: For each reviewed PR, identify source code files that were changed but received no review comments; extract a ~50-line chunk as a negative example labeled "No issues found."
Splits
Splits are deterministic by repository — all examples from the same repo appear in the same split.
Schema
Usage
from datasets import load_dataset
ds = load_dataset("ronantakizawa/github-codereview")
# Get a training example
example = ds["train"][0]
print(f"Review comment: {example['reviewer_comment']}")
print(f"Language: {example['language']}")
print(f"Before:\n{example['before_code'][:200]}")
print(f"After:\n{example['after_code'][:200]}")Filter by language
python_reviews = ds["train"].filter(lambda x: x["language"] == "Python")Filter by quality
high_quality = ds["train"].filter(lambda x: x["quality_score"] >= 0.5)Positive examples only
positives = ds["train"].filter(lambda x: not x["is_negative"])Negative examples only
negatives = ds["train"].filter(lambda x: x["is_negative"])Citation
If you use this dataset, please cite:
@dataset{takizawa2026codereviewdiffs,
title={Code Review Diffs: A Large-Scale Dataset of Review-Driven Code Changes},
author={Takizawa, Ronan},
year={2026},
publisher={Hugging Face},
url={https://huggingface.co/datasets/ronantakizawa/github-codereview}
}