datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
conventional-commits
Git Diff → Conventional Commit Messages
A dataset of 12,433 (git diff, commit message) pairs scraped from real open-source TypeScript repositories, filtered for quality and formatted for fine-tuning small language models to generate Conventional Commits.
Built as part of a project to fine-tune a local LLM to write commit messages as a prepare-commit-msg git hook. Full write-up: eliotbas.com/projects/commits-fine-tuning
Dataset details
Source… See the full description on the dataset page: https://huggingface.co/datasets/Elib27/conventional-commits.commitpack-subset-cfA subset of CommitPack used for pretraining SantaCoderPack from the OctoPack paper.
It focuses on data where the code before + special token + code after fits into 8192 tokens and on 6 languages. The data is in commit format (cf): <commit_before>code_before<commit_message>commit_message<commit_after>code_after.
commit-msg-edits
✍️ Commit Message Edits Dataset
This dataset is a collection of expert-labeled commit message edits contributed via Commit Message Editing app presented in Towards Realistic Evaluation of Commit Message Generation by Matching Online and Offline Settings.
Labelers were presented with GPT-4 generated messages for 15 commits from CMG benchmark from Long Code Arena and asked to manually edit them to be of good enough quality to submit to VCS.
You can check Manual tab in our… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/commit-msg-edits.commits-8192git-commit-message-dtcommit-message-quality
Commit Message Quality dataset
This is the dataset for commit message quality classification, used during processing of Commit Message Generation dataset from
🏟️ Long Code Arena benchmark.
This is a cleaned and relabeled version of the dataset from 📜 "Commit Message Matters: Investigating Impact and Evolution of Commit Message Quality", ICSE'23. We drop "Neither Why nor What" examples, clean all the external references (URLs, issues/PR references) from messages and manually label… See the full description on the dataset page: https://huggingface.co/datasets/saridormi/commit-message-quality.vaani-gujarati-sft-data
Vaani — Gujarati SFT & Eval Data
Final datasets for the Vaani 110M Gujarati medical SLM.
File
Rows
Purpose
sft_v4.jsonl
~150k
Instruction SFT: general instructions + format skills (JSON, extraction, exact-count lists)
medical_sft_v3.jsonl
~114k
Medical SFT: closed-book MedMCQA-gu + raw-grounded + redacted-grounded + 30% general mix
medmcqa_gu_val.jsonl
4,183
Held-out MedMCQA-gu validation split (eval only; disjoint from SFT)
The pretraining corpus is… See the full description on the dataset page: https://huggingface.co/datasets/pratham-commits/vaani-gujarati-sft-data.commit-messages-high-quality
Commit Messages from High-Quality Repositories
292,269 cleaned git commit messages scraped from the full histories of 15 well-regarded open-source
projects, balanced across two styles: normal (196,372) and
conventional commits (95,897).
Dataset Summary
Each record contains the commit subject, body, plus metadata: repo, sha, date,
author, and labels: style (normal/conventional), type (fix, feat, docs, ...),
scope, breaking.
Heavy cleaning: GitHub squash suffixes… See the full description on the dataset page: https://huggingface.co/datasets/Quad4/commit-messages-high-quality.git-commits
Dataset: dataset.jsonl
Auto-labeled commit dataset scraped from GitHub repositories. Each line is a JSON object representing one commit with extracted features and an inferred label.
Features
Field
Type
Description
Stats
text
string
Commit message first line, conventional prefix stripped
—
files_count
int
Number of files changed
mean 4.3, median 1, max 300
additions
int
Lines added
mean 88, median 6, max 187K
deletions
int
Lines deleted
mean 172… See the full description on the dataset page: https://huggingface.co/datasets/akaruineko/git-commits.synthetic-commit-msg-edits
✍️ Commit Message Edits Dataset - 🤖Synthetic
This dataset is a synthetic extension of our expert-labeled commit message edits dataset presented in Towards Realistic Evaluation of Commit Message Generation by Matching Online and Offline Settings.
You can check Synthetic tab in our visualization app to browse through the datapoints!
Dataset Structure
Default
Default split contains the synthetic messages generated from expert-labeled dataset by an LLM.… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/synthetic-commit-msg-edits.commitpackft
Dataset Card for CommitPackFT
Dataset Summary
CommitPackFT is a 2GB filtered version of CommitPack to contain only high-quality commit messages that resemble natural language instructions.
Creation: The dataset can be recreated using instructions available here.
Languages: 277
OctoPack🐙🎒:
Data
CommitPack
4TB of GitHub commits across 350 programming languages
CommitPackFT
Filtered version of CommitPack for high-quality commit messages that resemble… See the full description on the dataset page: https://huggingface.co/datasets/vnixxa31/commitpackft.openhands-commit-noise-databases
OpenHands Commit Noise Databases
This dataset contains commit-retrieval databases for 12 SWE-bench repositories at five noise ratios: 0%, 25%, 50%, 75%, and 100%.
Each archive expands to noise_NNN/<repository>/ directories containing:
commits.db: SQLite commit records
commits.faiss: normalized inner-product FAISS index
commits.meta.jsonl: FAISS row-to-commit metadata
commits.index_meta.json: embedding and index configuration
The 0% archive is an exact file-level copy of the… See the full description on the dataset page: https://huggingface.co/datasets/dengyixuan/openhands-commit-noise-databases.linux-kernel-commits-qwen35
Linux Kernel Commit Reasoning
Description
Teaches domain-specific instruction following and code generation for this expert.
Source
ewedubs/linux-kernel-commits-aireason-instruct - Apache-2.0
Formatted for the MoE-orchestrator project
(https://github.com/michaelowusuntim6/MoE-orchestrator). Expert target:
linux_kernel.
Format
Each record is a JSON object with a messages field formatted for Qwen3.5's
native chat template:… See the full description on the dataset page: https://huggingface.co/datasets/michaelowusuntim6/linux-kernel-commits-qwen35.code-text-galeras-commit-generation-3k-dedupedopenbsd-commits-alpaca
OpenBSD Commit History — Alpaca Format (v1)
Fine-tuning dataset derived from the full commit history of the
OpenBSD src repository, structured
for instruction fine-tuning in Alpaca format.
Task: given a unified diff, generate the commit message.
Dataset Summary
Field
Value
Examples
103,383
Size
~202 MB
Format
Alpaca JSONL
Date range
2000-01-01 → present
Source
openbsd/src (GitHub mirror)
License
ISC
Format
Each line is a… See the full description on the dataset page: https://huggingface.co/datasets/ajsbsd/openbsd-commits-alpaca.git-commit-message-splittercommitment-statements-pilot
Commitment statements — synthetic pilot
Small assistant-authored English statements for an experimental binary
commitment/decision classifier. Released for reproducibility, not as an
independently collected benchmark.
Training: 118 rows, 59 per label.
Validation/development: 14 rows, 7 per label.
The original 16-row test set is deliberately not distributed or evaluated.
Related contrastive families were kept together during the original split;
all 48 extension rows were… See the full description on the dataset page: https://huggingface.co/datasets/TwilightTechie/commitment-statements-pilot.bill_committees_us
Dataset Card for "bill_committees_us"
Dataset Summary
Dataset for US Congressional bills with committees information (bill_committees_us). Contains data for bills from the 108th to the 118th Congress, approximately 132,000 documents.
Supported Tasks and Leaderboards
More Information Needed
Languages
English
Dataset Structure
Data Instances
default
Data Fields
id: id of the bill in format(congress number +… See the full description on the dataset page: https://huggingface.co/datasets/dreamproit/bill_committees_us.commitpack50kgujarati-autoscientist-datasetmlab-synopsys-commitpack-sample5k-sample-commitpacktw-ly-committee
Taiwan Legislative Yuan Committee Data(ly-tw-committee)
您也可以透過網頁介面瀏覽 committee 資料集:https://dataly.openfun.app/collection/list/committee
委員會基本資料比較少會變動,可視為參考資料(Reference Data)
Data Fields
資料欄位
說明
委員會代號
為 committee 的 id
委員會名稱
如欄位名稱所述
委員會職掌
為一段短文敘述該委員會的工作職責
委員會類別
int 總共三個類別: 1:常設委員會 2:特種委員會 3:國會改革前舊委員會名稱
委員會類別:str
類別的中文名稱
pr-poet-commits
PR Poet Commit Poems
One thousand public Git repository commit messages paired with synthetic four-line poems generated for PR Poet.
Each record contains:
repo: the source GitHub repository
commit: the commit message
target: the generated poem as line_one through line_four
This is the original generated source set used by the PR Poet training curriculum. It intentionally includes poems that were later filtered out for rhyme or style, so not every example satisfies the final… See the full description on the dataset page: https://huggingface.co/datasets/mkly/pr-poet-commits.google-translate-commitpackft5k-sample-commitpack-editzephyr-commitsRoutine-Commitment-Datacommitpack_original_cm_old_code_to_diffcommitpack10k
