Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Maxscha /commitbench CommitBench: A Benchmark for Commit Message Generation EXECUTIVE SUMMARY We provide CommitBench as an open-source, reproducible and privacy- and license-aware benchmark for commit message generation. The dataset is gathered from GitHub repositories with licenses that permit redistribution. We provide six programming languages, Java, Python, Go, JavaScript, PHP, and Ruby. The commit messages in natural language are restricted to English, as it is the working language in… See the full description on the dataset page: https://huggingface.co/datasets/Maxscha/commitbench.text1M<n<10M13 likes6k downloads3y agoHugging Face02ASSERT-KTH /agent-commits-raw AI Coding-Agent Commits on GitHub This dataset documents commits associated with four AI coding agents: Claude, OpenAI Codex, GitHub Copilot and Cursor. It contains 1,853,915 commit records across 444,055 GitHub repositories, with 220,753 identifiable GitHub user accounts recorded as commit authors. Messages and author identities are in commits; repository metadata, file changes and patch text are available in separate tables. Dataset Agent Commit records… See the full description on the dataset page: https://huggingface.co/datasets/ASSERT-KTH/agent-commits-raw.tabular10M<n<100M0 likes4.9k downloads13d agoHugging Face03commit0 /commit0textn<1K3 likes3k downloads2y agoHugging Face04djtereano /commitbench_conventionaltext10K<n<100K0 likes2.8k downloads1y agoHugging Face05Elib27 /conventional-commits Git Diff → Conventional Commit Messages A dataset of 12,433 (git diff, commit message) pairs scraped from real open-source TypeScript repositories, filtered for quality and formatted for fine-tuning small language models to generate Conventional Commits. Built as part of a project to fine-tune a local LLM to write commit messages as a prepare-commit-msg git hook. Full write-up: eliotbas.com/projects/commits-fine-tuning Dataset details Source… See the full description on the dataset page: https://huggingface.co/datasets/Elib27/conventional-commits.texttext-generation10K<n<100K0 likes2.4k downloads3mo agoHugging Face06bigcode /commitpack-subset-cfA subset of CommitPack used for pretraining SantaCoderPack from the OctoPack paper. It focuses on data where the code before + special token + code after fits into 8192 tokens and on 6 languages. The data is in commit format (cf): <commit_before>code_before<commit_message>commit_message<commit_after>code_after. text100K<n<1M2 likes2.4k downloads3y agoHugging Face07JetBrains-Research /commit-chronicle 📜 CommitChronicle 🔮 This is the dataset for commit message generation (and/or completion), introduced in the paper "From Commit Message Generation to History-Aware Commit Message Completion", ASE 2023. Its key features: large-scale and multilingual: contains 10.7M commits from 11.9k GitHub repositories in 20 programming languages; diverse: avoids restrictive filtering on commit messages or commit diffs structure; suitable for experiments with commit history: provides metadata… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/commit-chronicle.tabulartext-generation10M<n<100M13 likes1.7k downloads3y agoHugging Face08meriemm6 /commit-classification-dataset Commit Classification Dataset This dataset is designed for multi-label classification of Git commit messages into predefined categories. Dataset Summary This dataset contains: Training data: Commit messages and their corresponding labels for training the model. Validation data: A separate set of messages for tuning and evaluation. Testing data: Unlabeled commit messages for testing the model’s performance. The goal of the dataset is to classify each commit message into… See the full description on the dataset page: https://huggingface.co/datasets/meriemm6/commit-classification-dataset.texttext-classification1K<n<10K0 likes1.5k downloads2y agoHugging Face09bigcode /github-commits-diff-dedup-pjjs-april Deduplicated Commits Deduplicated based on diff: content = '\n'.join(difflib.unified_diff( old_content.splitlines(keepends=True), new_content.splitlines(keepends=True), n=5 )) Parameters: Minimum ngram size: 5 MinHash ngram size: 5 MinHash threshold: 0.8 text100K<n<1M4 likes1.3k downloads3y agoHugging Face10rsh-raj /commit-classification-17ktext10K<n<100K0 likes1.2k downloads2y agoHugging Face11knesset-asr /knesset-committees-chunkstabular1M<n<10M0 likes1.1k downloads27d agoHugging Face12wentingzhao /commit0_combinedtextn<1K0 likes1k downloads2y agoHugging Face13bigcode /commitpackmetaGitHub metadata for https://huggingface.co/datasets/bigcode/commitpack text10M<n<100M4 likes998 downloads3y agoHugging Face14ZipLime /commitments-of-traders Commitments of Traders Who was long and who was short in every US futures market — and, for once, when anyone could actually see it. 421 223 market-weeks · 211 071 point-in-time rows · 2 585 weekly releases · 762 markets · 2010-01-05 to 2026-09-08 The pipeline lives in recipe/ at the same revision as the data. See PIPELINE.md for the method. The data is from Tuesday. It comes out on Friday. A COT report is taken as of the close on Tuesday and published at 3:30… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/commitments-of-traders.tabulartabular-regression100K<n<1M0 likes938 downloads2d agoHugging Face15placeholderlabs /pretrain-commits-v2-mixNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 79,318,557,379 (79.3B) Trainable tokens 28,772,968,648 (28.8B) Documents 31,431,846 Shards 696 UTF-8 bytes 310,412,170,445 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-commits-v2-mix.tabular10M<n<100M0 likes793 downloads22d agoHugging Face16Hirunima /commit-chronicle 📜 CommitChronicle 🔮 This is the dataset for commit message generation (and/or completion), introduced in the paper "From Commit Message Generation to History-Aware Commit Message Completion", ASE 2023. Its key features: large-scale and multilingual: contains 10.7M commits from 11.9k GitHub repositories in 20 programming languages; diverse: avoids restrictive filtering on commit messages or commit diffs structure; suitable for experiments with commit history: provides metadata… See the full description on the dataset page: https://huggingface.co/datasets/Hirunima/commit-chronicle.tabulartext-generation10M<n<100M0 likes780 downloads7mo agoHugging Face17project-themis /git-commits-merged Themis-Git-Commits-Merged Overview Themis-Git-Commits-Merged is a large-scale dataset of ~3.98M single-file code commits from permissively licensed GitHub repositories that have been cross-referenced with GHTorrent pull request data to retain only commits that are part of successfully merged, non-reverted pull requests. This provides implicit human validation of each code change — a merge decision by project maintainers confirms the intent and quality of… See the full description on the dataset page: https://huggingface.co/datasets/project-themis/git-commits-merged.texttext-generation1M<n<10M1 likes758 downloads2mo agoHugging Face18bigcode /commits-codegeex Dataset Card for "commits-codegeex" More Information needed text1M<n<10M6 likes738 downloads3y agoHugging Face19JetBrains-Research /commit-msg-edits ✍️ Commit Message Edits Dataset This dataset is a collection of expert-labeled commit message edits contributed via Commit Message Editing app presented in Towards Realistic Evaluation of Commit Message Generation by Matching Online and Offline Settings. Labelers were presented with GPT-4 generated messages for 15 commits from CMG benchmark from Long Code Arena and asked to manually edit them to be of good enough quality to submit to VCS. You can check Manual tab in our… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/commit-msg-edits.textn<1K1 likes734 downloads2y agoHugging Face20andstor /cvevc_commits Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/andstor/cvevc_commits.text1M<n<10M0 likes664 downloads8mo agoHugging Face21placeholderlabs /pretrain-commits-v2-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 12,571,681,749 (12.6B) Trainable tokens 4,460,160,435 (4.5B) Documents 992,475 Shards 327 UTF-8 bytes 49,288,867,997 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-commits-v2-mix-long-context.tabular1M<n<10M0 likes589 downloads22d agoHugging Face22bigcode /commits-pjj-diff Dataset Card for "commits-pjj-diff" More Information needed text1M<n<10M2 likes438 downloads4y agoHugging Face23committa /serena-synthetic-it-28h Qwen3-TTS Italian Synthetic Speech (27h) Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV, Piper-ready metadata. Dataset summary Property Value Clips (train / val) 26,523 / 2,947 Total duration ~27.3 h (98,099 s) Sample rate 22,050 Hz mono, 16-bit WAV Loudness Normalized to -23 LUFS, silence-trimmed Language Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-28h.audiotext-to-speech10K<n<100K1 likes423 downloads2mo agoHugging Face24Maxscha /commitbench_longCommitBench: A Benchmark for Commit Message Generation We provide CommitBench as an open-source, reproducible and privacy- and license-aware benchmark for commit message generation. The dataset is gathered from GitHub repositories with licenses that permit redistribution. We provide six programming languages, Java, Python, Go, JavaScript, PHP, and Ruby. The commit messages in natural language are restricted to English, as it is the working language in many software development projects. The… See the full description on the dataset page: https://huggingface.co/datasets/Maxscha/commitbench_long.texttranslation1M<n<10M0 likes358 downloads3y agoHugging Face25JetBrains-Research /lca-commit-message-generation 🏟️ Long Code Arena (Commit message generation) This is the benchmark for the Commit message generation task as part of the 🏟️ Long Code Arena benchmark. The dataset is a manually curated subset of the Python test set from the 🤗 CommitChronicle dataset, tailored for larger commits. All the repositories are published under permissive licenses (MIT, Apache-2.0, and BSD-3-Clause). The datapoints can be removed upon request. How-to from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-commit-message-generation.text1K<n<10K0 likes357 downloads2y agoHugging Face26ObscuraCoder /commit-chronicleThis is a filtered version of the JetBrains-Research/commit-chronicle dataset. It has been subsetted for the following languages: [ "C", "C++", "Go", "Java", "Python", "Rust", "TypeScript" ] Further filtering steps undertaken are: Useless features have been removed and only the message and diff retained. Only commits that modify a single file have been chosen. Samples containing diffs longer than 1024 tokens (by the ObscuraCoder/Tokenizer tokenizer estimate) have been discarded.text1M<n<10M4 likes352 downloads2y agoHugging Face27Berom0227 /tangled-ccs-commits Detecting Multiple Semantic Concerns in Tangled Code Commits using Small Language Models This dataset contains commit data for training and evaluating models on software engineering tasks, specifically focusing on identifying and separating concerns in multi-concern commits. Every tangled (multi-concern) commit in this dataset is composed exclusively of atomic commits from a single repository — resolving a structural weakness in earlier cross-repo tangles (which were trivially… See the full description on the dataset page: https://huggingface.co/datasets/Berom0227/tangled-ccs-commits.texttext-generation1K<n<10K1 likes286 downloads2mo agoHugging Face28RASMUS /commitmoe-qwen35-fp8-expert-routing-tracestext0 likes284 downloads2mo agoHugging Face29bigcode /commits-8192text100K<n<1M3 likes270 downloads3y agoHugging Face30rsh-raj /appsmith-commitstext1K<n<10K0 likes270 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.