Quad4/commit-messages-high-quality
Commit Messages from High-Quality Repositories 292,269 cleaned git commit messages scraped from the full histories of 15 well-regarded open-source projects, balanced across two styles: normal (196,372) and conventional commits (95,897). Dataset Summary Each record contains the commit subject, body, plus metadata: repo, sha, date, author, and labels: style (normal/conventional), type (fix, feat, docs, ...), scope, breaking. Heavy cleaning: GitHub squash suffixes… See the full description on the dataset page: https://huggingface.co/datasets/Quad4/commit-messages-high-quality.
Commit Messages from High-Quality Repositories
292,269 cleaned git commit messages scraped from the full histories of 15 well-regarded open-source projects, balanced across two styles: normal (196,372) and conventional commits (95,897).
Dataset Summary
- Each record contains the commit
subject,body, plus metadata:repo,sha,date,author, and labels:style(normal/conventional),type(fix, feat, docs, ...),scope,breaking. - Heavy cleaning: GitHub squash suffixes
(#123)stripped, bot/automated commits removed, canned boilerplate removed, exact duplicates removed, git trailers (Signed-off-by etc.) stripped from bodies, junk subjects ("wip", "update", ...) removed. - Temporal split: train is the oldest ~98% of commits, validation the next ~1%, test the newest ~1%. No train/eval temporal leakage: every eval commit is newer than every train commit.
Supported Tasks
text-generation: commit message generation, style modeling, conventional-commit classification, software-engineering NLP.
Usage
from datasets import load_dataset
ds = load_dataset("USERNAME/REPO", "default") # full metadata records
view = load_dataset("USERNAME/REPO", "training_view") # prompt/completion pairs (body -> subject)training_view maps a change description (the commit body) to the commit subject line and is formatted for fine-tuning tools such as Unsloth, with explicit prompt and completion columns.
Source Repositories
Upstream Licenses
Commit messages are inherited from their source projects. Repositories under copyleft or source-available licenses (git: GPL-2.0; redis: BSD-3 / RSALv2 / SSPLv1) were excluded as a precaution: software licenses technically cover the licensed program rather than VCS metadata, but commit messages are therefore not openly licensed either, and longer commit bodies can constitute original expression of their authors. Only permissively-licensed projects are included:
Users should still respect upstream license terms when redistributing or using this dataset commercially.
Cleaning Statistics
Splits
Privacy
- No author names or email addresses are included. Records are identified only by content
sha+ repository, which preserves verifiable provenance without identifying individuals. - Note: commit bodies may reference contributor names in prose (e.g. credit lines), as these are part of the historical commit text itself.
Limitations
- English-dominant; commit messages from a single ecosystem of C/C++ and JavaScript/TypeScript projects.
- No diffs attached: the dataset describes what was written, not the code change itself.
- Bodies are truncated at 2,000 characters.
- Human authors made the source commits; quality varies by repository and era.
Reproduction
Built with a two-stage pipeline (blobless git clone --filter=blob:none, then git log export and rule-based cleaning). See stats.json for exact counts.
