Team Ai
Datasetpublic

Quad4/commit-messages-high-quality

Commit Messages from High-Quality Repositories 292,269 cleaned git commit messages scraped from the full histories of 15 well-regarded open-source projects, balanced across two styles: normal (196,372) and conventional commits (95,897). Dataset Summary Each record contains the commit subject, body, plus metadata: repo, sha, date, author, and labels: style (normal/conventional), type (fix, feat, docs, ...), scope, breaking. Heavy cleaning: GitHub squash suffixes… See the full description on the dataset page: https://huggingface.co/datasets/Quad4/commit-messages-high-quality.

sourceHugging Faceotherupdated 15d agoView on Hugging Face
0likes84downloads
Dataset Card

Commit Messages from High-Quality Repositories

292,269 cleaned git commit messages scraped from the full histories of 15 well-regarded open-source projects, balanced across two styles: normal (196,372) and conventional commits (95,897).

Dataset Summary

  • —Each record contains the commit subject, body, plus metadata: repo, sha, date, author, and labels: style (normal/conventional), type (fix, feat, docs, ...), scope, breaking.
  • —Heavy cleaning: GitHub squash suffixes (#123) stripped, bot/automated commits removed, canned boilerplate removed, exact duplicates removed, git trailers (Signed-off-by etc.) stripped from bodies, junk subjects ("wip", "update", ...) removed.
  • —Temporal split: train is the oldest ~98% of commits, validation the next ~1%, test the newest ~1%. No train/eval temporal leakage: every eval commit is newer than every train commit.

Supported Tasks

  • —text-generation: commit message generation, style modeling, conventional-commit classification, software-engineering NLP.

Usage

python
from datasets import load_dataset

ds = load_dataset("USERNAME/REPO", "default")          # full metadata records
view = load_dataset("USERNAME/REPO", "training_view")  # prompt/completion pairs (body -> subject)

training_view maps a change description (the commit body) to the commit subject line and is formatted for fine-tuning tools such as Unsloth, with explicit prompt and completion columns.

Source Repositories

RepositoryURLCommits in dataset
postgreshttps://github.com/postgres/postgres57,971
opensslhttps://github.com/openssl/openssl53,169
angularhttps://github.com/angular/angular39,465
curlhttps://github.com/curl/curl37,130
ant-designhttps://github.com/ant-design/ant-design22,853
neovimhttps://github.com/neovim/neovim22,096
babelhttps://github.com/babel/babel13,719
vitehttps://github.com/vitejs/vite8,209
eslinthttps://github.com/eslint/eslint7,837
tmuxhttps://github.com/tmux/tmux7,618
vue-corehttps://github.com/vuejs/core7,540
vitesthttps://github.com/vitest-dev/vitest5,120
nestjshttps://github.com/nestjs/nest4,147
chart-jshttps://github.com/chartjs/Chart.js3,718
axioshttps://github.com/axios/axios1,677

Upstream Licenses

Commit messages are inherited from their source projects. Repositories under copyleft or source-available licenses (git: GPL-2.0; redis: BSD-3 / RSALv2 / SSPLv1) were excluded as a precaution: software licenses technically cover the licensed program rather than VCS metadata, but commit messages are therefore not openly licensed either, and longer commit bodies can constitute original expression of their authors. Only permissively-licensed projects are included:

RepositoryLicense
angular/angularMIT
vuejs/coreMIT
vitejs/viteMIT
babel/babelMIT
eslint/eslintMIT
ant-design/ant-designMIT
axios/axiosMIT
nestjs/nestMIT
chartjs/Chart.jsMIT
vitest-dev/vitestMIT
curl/curlcurl (MIT-style)
neovim/neovimApache-2.0 / VIM
tmux/tmuxISC
openssl/opensslApache-2.0
postgres/postgresPostgreSQL License

Users should still respect upstream license terms when redistributing or using this dataset commercially.

Cleaning Statistics

FilterRecords dropped
duplicate77,038
bad_subject16,190
botorautomated26,432
license_excluded83,369
canned_message9,943

Splits

SplitRecordsDate range
train286,4251996-07-09 → 2026-07-14
validation2,9222026-07-14 → 2026-08-20
test2,9222026-08-20 → 2026-09-23

Privacy

  • —No author names or email addresses are included. Records are identified only by content sha + repository, which preserves verifiable provenance without identifying individuals.
  • —Note: commit bodies may reference contributor names in prose (e.g. credit lines), as these are part of the historical commit text itself.

Limitations

  • —English-dominant; commit messages from a single ecosystem of C/C++ and JavaScript/TypeScript projects.
  • —No diffs attached: the dataset describes what was written, not the code change itself.
  • —Bodies are truncated at 2,000 characters.
  • —Human authors made the source commits; quality varies by repository and era.

Reproduction

Built with a two-stage pipeline (blobless git clone --filter=blob:none, then git log export and rule-based cleaning). See stats.json for exact counts.