Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Elib27 /conventional-commits Git Diff → Conventional Commit Messages A dataset of 12,433 (git diff, commit message) pairs scraped from real open-source TypeScript repositories, filtered for quality and formatted for fine-tuning small language models to generate Conventional Commits. Built as part of a project to fine-tune a local LLM to write commit messages as a prepare-commit-msg git hook. Full write-up: eliotbas.com/projects/commits-fine-tuning Dataset details Source… See the full description on the dataset page: https://huggingface.co/datasets/Elib27/conventional-commits.texttext-generation10K<n<100K0 likes2.4k downloads3mo agoHugging Face02bigcode /commitpack-subset-cfA subset of CommitPack used for pretraining SantaCoderPack from the OctoPack paper. It focuses on data where the code before + special token + code after fits into 8192 tokens and on 6 languages. The data is in commit format (cf): <commit_before>code_before<commit_message>commit_message<commit_after>code_after. text100K<n<1M2 likes2.4k downloads3y agoHugging Face03JetBrains-Research /commit-msg-edits ✍️ Commit Message Edits Dataset This dataset is a collection of expert-labeled commit message edits contributed via Commit Message Editing app presented in Towards Realistic Evaluation of Commit Message Generation by Matching Online and Offline Settings. Labelers were presented with GPT-4 generated messages for 15 commits from CMG benchmark from Long Code Arena and asked to manually edit them to be of good enough quality to submit to VCS. You can check Manual tab in our… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/commit-msg-edits.textn<1K1 likes734 downloads2y agoHugging Face04bigcode /commits-8192text100K<n<1M3 likes270 downloads3y agoHugging Face05Tavernari /git-commit-message-dttext1K<n<10K5 likes162 downloads2y agoHugging Face06saridormi /commit-message-quality Commit Message Quality dataset This is the dataset for commit message quality classification, used during processing of Commit Message Generation dataset from 🏟️ Long Code Arena benchmark. This is a cleaned and relabeled version of the dataset from 📜 "Commit Message Matters: Investigating Impact and Evolution of Commit Message Quality", ICSE'23. We drop "Neither Why nor What" examples, clean all the external references (URLs, issues/PR references) from messages and manually label… See the full description on the dataset page: https://huggingface.co/datasets/saridormi/commit-message-quality.texttext-classification1K<n<10K0 likes122 downloads3y agoHugging Face07pratham-commits /vaani-gujarati-sft-data Vaani — Gujarati SFT & Eval Data Final datasets for the Vaani 110M Gujarati medical SLM. File Rows Purpose sft_v4.jsonl ~150k Instruction SFT: general instructions + format skills (JSON, extraction, exact-count lists) medical_sft_v3.jsonl ~114k Medical SFT: closed-book MedMCQA-gu + raw-grounded + redacted-grounded + 30% general mix medmcqa_gu_val.jsonl 4,183 Held-out MedMCQA-gu validation split (eval only; disjoint from SFT) The pretraining corpus is… See the full description on the dataset page: https://huggingface.co/datasets/pratham-commits/vaani-gujarati-sft-data.texttext-generation1K<n<10K1 likes120 downloads27d agoHugging Face08Quad4 /commit-messages-high-quality Commit Messages from High-Quality Repositories 292,269 cleaned git commit messages scraped from the full histories of 15 well-regarded open-source projects, balanced across two styles: normal (196,372) and conventional commits (95,897). Dataset Summary Each record contains the commit subject, body, plus metadata: repo, sha, date, author, and labels: style (normal/conventional), type (fix, feat, docs, ...), scope, breaking. Heavy cleaning: GitHub squash suffixes… See the full description on the dataset page: https://huggingface.co/datasets/Quad4/commit-messages-high-quality.texttext-generation100K<n<1M0 likes86 downloads15d agoHugging Face09akaruineko /git-commits Dataset: dataset.jsonl Auto-labeled commit dataset scraped from GitHub repositories. Each line is a JSON object representing one commit with extracted features and an inferred label. Features Field Type Description Stats text string Commit message first line, conventional prefix stripped — files_count int Number of files changed mean 4.3, median 1, max 300 additions int Lines added mean 88, median 6, max 187K deletions int Lines deleted mean 172… See the full description on the dataset page: https://huggingface.co/datasets/akaruineko/git-commits.tabular10K<n<100K0 likes72 downloads3mo agoHugging Face10JetBrains-Research /synthetic-commit-msg-edits ✍️ Commit Message Edits Dataset - 🤖Synthetic This dataset is a synthetic extension of our expert-labeled commit message edits dataset presented in Towards Realistic Evaluation of Commit Message Generation by Matching Online and Offline Settings. You can check Synthetic tab in our visualization app to browse through the datapoints! Dataset Structure Default Default split contains the synthetic messages generated from expert-labeled dataset by an LLM.… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/synthetic-commit-msg-edits.tabular10K<n<100K0 likes68 downloads2y agoHugging Face11vnixxa31 /commitpackft Dataset Card for CommitPackFT Dataset Summary CommitPackFT is a 2GB filtered version of CommitPack to contain only high-quality commit messages that resemble natural language instructions. Creation: The dataset can be recreated using instructions available here. Languages: 277 OctoPack🐙🎒: Data CommitPack 4TB of GitHub commits across 350 programming languages CommitPackFT Filtered version of CommitPack for high-quality commit messages that resemble… See the full description on the dataset page: https://huggingface.co/datasets/vnixxa31/commitpackft.text100K<n<1M1 likes68 downloads7mo agoHugging Face12dengyixuan /openhands-commit-noise-databases OpenHands Commit Noise Databases This dataset contains commit-retrieval databases for 12 SWE-bench repositories at five noise ratios: 0%, 25%, 50%, 75%, and 100%. Each archive expands to noise_NNN/<repository>/ directories containing: commits.db: SQLite commit records commits.faiss: normalized inner-product FAISS index commits.meta.jsonl: FAISS row-to-commit metadata commits.index_meta.json: embedding and index configuration The 0% archive is an exact file-level copy of the… See the full description on the dataset page: https://huggingface.co/datasets/dengyixuan/openhands-commit-noise-databases.tabularn<1K0 likes37 downloads3mo agoHugging Face13michaelowusuntim6 /linux-kernel-commits-qwen35 Linux Kernel Commit Reasoning Description Teaches domain-specific instruction following and code generation for this expert. Source ewedubs/linux-kernel-commits-aireason-instruct - Apache-2.0 Formatted for the MoE-orchestrator project (https://github.com/michaelowusuntim6/MoE-orchestrator). Expert target: linux_kernel. Format Each record is a JSON object with a messages field formatted for Qwen3.5's native chat template:… See the full description on the dataset page: https://huggingface.co/datasets/michaelowusuntim6/linux-kernel-commits-qwen35.texttext-generation10K<n<100K0 likes32 downloads3d agoHugging Face14semeru /code-text-galeras-commit-generation-3k-dedupedtabular1K<n<10K0 likes30 downloads3y agoHugging Face15ajsbsd /openbsd-commits-alpaca OpenBSD Commit History — Alpaca Format (v1) Fine-tuning dataset derived from the full commit history of the OpenBSD src repository, structured for instruction fine-tuning in Alpaca format. Task: given a unified diff, generate the commit message. Dataset Summary Field Value Examples 103,383 Size ~202 MB Format Alpaca JSONL Date range 2000-01-01 → present Source openbsd/src (GitHub mirror) License ISC Format Each line is a… See the full description on the dataset page: https://huggingface.co/datasets/ajsbsd/openbsd-commits-alpaca.texttext-generation100K<n<1M1 likes30 downloads4mo agoHugging Face16Tavernari /git-commit-message-splittertext1K<n<10K0 likes26 downloads1y agoHugging Face17TwilightTechie /commitment-statements-pilot Commitment statements — synthetic pilot Small assistant-authored English statements for an experimental binary commitment/decision classifier. Released for reproducibility, not as an independently collected benchmark. Training: 118 rows, 59 per label. Validation/development: 14 rows, 7 per label. The original 16-row test set is deliberately not distributed or evaluated. Related contrastive families were kept together during the original split; all 48 extension rows were… See the full description on the dataset page: https://huggingface.co/datasets/TwilightTechie/commitment-statements-pilot.texttext-classificationn<1K0 likes23 downloads2d agoHugging Face18dreamproit /bill_committees_us Dataset Card for "bill_committees_us" Dataset Summary Dataset for US Congressional bills with committees information (bill_committees_us). Contains data for bills from the 108th to the 118th Congress, approximately 132,000 documents. Supported Tasks and Leaderboards More Information Needed Languages English Dataset Structure Data Instances default Data Fields id: id of the bill in format(congress number +… See the full description on the dataset page: https://huggingface.co/datasets/dreamproit/bill_committees_us.tabulartext-generation100K<n<1M5 likes19 downloads2y agoHugging Face19rqadri /commitpack50ktext10K<n<100K0 likes19 downloads5mo agoHugging Face20pratham-commits /gujarati-autoscientist-datasettext100K<n<1M1 likes18 downloads4mo agoHugging Face21LeonJia /mlab-synopsys-commitpack-sampletextn<1K0 likes16 downloads3y agoHugging Face22JiyangZhang /5k-sample-commitpacktext1K<n<10K0 likes14 downloads3y agoHugging Face23openfun /tw-ly-committee Taiwan Legislative Yuan Committee Data(ly-tw-committee) 您也可以透過網頁介面瀏覽 committee 資料集:https://dataly.openfun.app/collection/list/committee 委員會基本資料比較少會變動,可視為參考資料(Reference Data) Data Fields 資料欄位 說明 委員會代號 為 committee 的 id 委員會名稱 如欄位名稱所述 委員會職掌 為一段短文敘述該委員會的工作職責 委員會類別 int 總共三個類別: 1:常設委員會 2:特種委員會 3:國會改革前舊委員會名稱 委員會類別:str 類別的中文名稱 tabularn<1K0 likes12 downloads2y agoHugging Face24mkly /pr-poet-commits PR Poet Commit Poems One thousand public Git repository commit messages paired with synthetic four-line poems generated for PR Poet. Each record contains: repo: the source GitHub repository commit: the commit message target: the generated poem as line_one through line_four This is the original generated source set used by the PR Poet training curriculum. It intentionally includes poems that were later filtered out for rhyme or style, so not every example satisfies the final… See the full description on the dataset page: https://huggingface.co/datasets/mkly/pr-poet-commits.texttext-generation1K<n<10K0 likes10 downloads1mo agoHugging Face25mesolitica /google-translate-commitpackfttext100K<n<1M1 likes9 downloads3y agoHugging Face26JiyangZhang /5k-sample-commitpack-edittext1K<n<10K0 likes9 downloads3y agoHugging Face27konovaai /zephyr-commitstextn<1K0 likes7 downloads2y agoHugging Face28adiez85 /Routine-Commitment-Datatextn<1K0 likes7 downloads9mo agoHugging Face29rqadri /commitpack_original_cm_old_code_to_difftext10K<n<100K0 likes6 downloads4mo agoHugging Face30rqadri /commitpack10ktext10K<n<100K0 likes4 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.