Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ahnaftaz /locus-commit-pool-v1 Locus Commit Pool v1 Native Git history, preserved as replayable software changes Commit message · complete selected before-state · unified patches · native object IDs · provenance · experimental labels Locus Commit Pool v1 is a large evidence pool for studying and training on how real software changes. Each document represents one surviving single-parent, multi-file Git commit. It keeps the commit message, the selected files as they existed before the… See the full description on the dataset page: https://huggingface.co/datasets/ahnaftaz/locus-commit-pool-v1.text-generation0 likes91k downloads2mo agoHugging Face02bigcode /commitpackftCommitPackFT is is a 2GB filtered version of CommitPack to contain only high-quality commit messages that resemble natural language instructions.114 likes34k downloads3y agoHugging Face03bigcode /commitpackCommitPack is is a 4TB dataset of commits scraped from GitHub repositories that are permissively licensed.79 likes15k downloads2y agoHugging Face04Maxscha /commitbench CommitBench: A Benchmark for Commit Message Generation EXECUTIVE SUMMARY We provide CommitBench as an open-source, reproducible and privacy- and license-aware benchmark for commit message generation. The dataset is gathered from GitHub repositories with licenses that permit redistribution. We provide six programming languages, Java, Python, Go, JavaScript, PHP, and Ruby. The commit messages in natural language are restricted to English, as it is the working language in… See the full description on the dataset page: https://huggingface.co/datasets/Maxscha/commitbench.text1M<n<10M13 likes6k downloads3y agoHugging Face05ASSERT-KTH /agent-commits-raw AI Coding-Agent Commits on GitHub This dataset documents commits associated with four AI coding agents: Claude, OpenAI Codex, GitHub Copilot and Cursor. It contains 1,853,915 commit records across 444,055 GitHub repositories, with 220,753 identifiable GitHub user accounts recorded as commit authors. Messages and author identities are in commits; repository metadata, file changes and patch text are available in separate tables. Dataset Agent Commit records… See the full description on the dataset page: https://huggingface.co/datasets/ASSERT-KTH/agent-commits-raw.tabular10M<n<100M0 likes4.9k downloads13d agoHugging Face06ivrit-ai /knesset-committeesgated About This dataset is derived from raw a/v recordings and human-generated protocols of the Knesset (the Israeli house of representatives) committee sessions as part of the ivrit.ai project. Consider visiting the preview space for this dataset here Method Data dumps from the Knesset contain A/V recordings of committee sessions, alongside human-generated protocols. We extract the audio stream, abd produce weakly time stamped segmentation of the protocol text (we… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-committees.automatic-speech-recognition3 likes4.3k downloads5mo agoHugging Face07commit0 /commit0textn<1K3 likes3k downloads2y agoHugging Face08FineEnvs /repo2rlenv-commit-runtime repo2rlenv-commit-runtime-v2 Generated by Repo2RLEnv — turning real GitHub repositories into verifiable RL environments. 💡 Browse this dataset in your browser — click the badge above or open HuggingFaceH4/harbor-visualiser to inspect every task's spec, instruction, oracle patch, test script, and Dockerfile. Source repos (22): encode/httpx encode/starlette gin-gonic/gin gofiber/fiber golang-jwt/jwt google/uuid gorilla/mux gorilla/websocket labstack/echo pallets/click… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/repo2rlenv-commit-runtime.n<1K0 likes3k downloads14d agoHugging Face09djtereano /commitbench_conventionaltext10K<n<100K0 likes2.8k downloads1y agoHugging Face10Elib27 /conventional-commits Git Diff → Conventional Commit Messages A dataset of 12,433 (git diff, commit message) pairs scraped from real open-source TypeScript repositories, filtered for quality and formatted for fine-tuning small language models to generate Conventional Commits. Built as part of a project to fine-tune a local LLM to write commit messages as a prepare-commit-msg git hook. Full write-up: eliotbas.com/projects/commits-fine-tuning Dataset details Source… See the full description on the dataset page: https://huggingface.co/datasets/Elib27/conventional-commits.texttext-generation10K<n<100K0 likes2.4k downloads3mo agoHugging Face11bigcode /commitpack-subset-cfA subset of CommitPack used for pretraining SantaCoderPack from the OctoPack paper. It focuses on data where the code before + special token + code after fits into 8192 tokens and on 6 languages. The data is in commit format (cf): <commit_before>code_before<commit_message>commit_message<commit_after>code_after. text100K<n<1M2 likes2.4k downloads3y agoHugging Face12JetBrains-Research /commit-chronicle 📜 CommitChronicle 🔮 This is the dataset for commit message generation (and/or completion), introduced in the paper "From Commit Message Generation to History-Aware Commit Message Completion", ASE 2023. Its key features: large-scale and multilingual: contains 10.7M commits from 11.9k GitHub repositories in 20 programming languages; diverse: avoids restrictive filtering on commit messages or commit diffs structure; suitable for experiments with commit history: provides metadata… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/commit-chronicle.tabulartext-generation10M<n<100M13 likes1.7k downloads3y agoHugging Face13meriemm6 /commit-classification-dataset Commit Classification Dataset This dataset is designed for multi-label classification of Git commit messages into predefined categories. Dataset Summary This dataset contains: Training data: Commit messages and their corresponding labels for training the model. Validation data: A separate set of messages for tuning and evaluation. Testing data: Unlabeled commit messages for testing the model’s performance. The goal of the dataset is to classify each commit message into… See the full description on the dataset page: https://huggingface.co/datasets/meriemm6/commit-classification-dataset.texttext-classification1K<n<10K0 likes1.5k downloads2y agoHugging Face14project-themis /git-commits Themis-Git-Commits Overview Themis-Git-Commits is a large-scale dataset of single-file code commits mined from permissively licensed GitHub repositories via the BigQuery GitHub public dataset. The SQL query restricts to repositories under permissive open-source licenses only (MIT, Apache-2.0, BSD-2/3-Clause, ISC, CC0-1.0, EPL-1.0, MPL-2.0, Unlicense, AGPL-3.0, LGPL-2.1, Artistic-2.0). The BigQuery snapshot used contains commits up to early 2022 —… See the full description on the dataset page: https://huggingface.co/datasets/project-themis/git-commits.text-generation10M<n<100M1 likes1.4k downloads2mo agoHugging Face15bigcode /github-commits-diff-dedup-pjjs-april Deduplicated Commits Deduplicated based on diff: content = '\n'.join(difflib.unified_diff( old_content.splitlines(keepends=True), new_content.splitlines(keepends=True), n=5 )) Parameters: Minimum ngram size: 5 MinHash ngram size: 5 MinHash threshold: 0.8 text100K<n<1M4 likes1.3k downloads3y agoHugging Face16rsh-raj /commit-classification-17ktext10K<n<100K0 likes1.2k downloads2y agoHugging Face17knesset-asr /knesset-committees-chunkstabular1M<n<10M0 likes1.1k downloads27d agoHugging Face18wentingzhao /commit0_combinedtextn<1K0 likes1k downloads2y agoHugging Face19bigcode /commitpackmetaGitHub metadata for https://huggingface.co/datasets/bigcode/commitpack text10M<n<100M4 likes998 downloads3y agoHugging Face20ZipLime /commitments-of-traders Commitments of Traders Who was long and who was short in every US futures market — and, for once, when anyone could actually see it. 421 223 market-weeks · 211 071 point-in-time rows · 2 585 weekly releases · 762 markets · 2010-01-05 to 2026-09-08 The pipeline lives in recipe/ at the same revision as the data. See PIPELINE.md for the method. The data is from Tuesday. It comes out on Friday. A COT report is taken as of the close on Tuesday and published at 3:30… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/commitments-of-traders.tabulartabular-regression100K<n<1M0 likes938 downloads2d agoHugging Face21fulldecent /one-million-commits One million commits A large variety of git commits pulled from across GitHub. Created by William Entriken, released 2023-09-26, version 1. This composition is licensed under the MIT license. Intended use This dataset could be used to train a model concerned with programming tasks: Summarize some programming work Perform work given a description of the work to do Learn-by-example the syntax for all active programming languages and structured data formats This… See the full description on the dataset page: https://huggingface.co/datasets/fulldecent/one-million-commits.text-classification1M<n<10M4 likes895 downloads1y agoHugging Face22placeholderlabs /pretrain-commits-v2-mixNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 79,318,557,379 (79.3B) Trainable tokens 28,772,968,648 (28.8B) Documents 31,431,846 Shards 696 UTF-8 bytes 310,412,170,445 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-commits-v2-mix.tabular10M<n<100M0 likes793 downloads21d agoHugging Face23Hirunima /commit-chronicle 📜 CommitChronicle 🔮 This is the dataset for commit message generation (and/or completion), introduced in the paper "From Commit Message Generation to History-Aware Commit Message Completion", ASE 2023. Its key features: large-scale and multilingual: contains 10.7M commits from 11.9k GitHub repositories in 20 programming languages; diverse: avoids restrictive filtering on commit messages or commit diffs structure; suitable for experiments with commit history: provides metadata… See the full description on the dataset page: https://huggingface.co/datasets/Hirunima/commit-chronicle.tabulartext-generation10M<n<100M0 likes780 downloads7mo agoHugging Face24project-themis /git-commits-merged Themis-Git-Commits-Merged Overview Themis-Git-Commits-Merged is a large-scale dataset of ~3.98M single-file code commits from permissively licensed GitHub repositories that have been cross-referenced with GHTorrent pull request data to retain only commits that are part of successfully merged, non-reverted pull requests. This provides implicit human validation of each code change — a merge decision by project maintainers confirms the intent and quality of… See the full description on the dataset page: https://huggingface.co/datasets/project-themis/git-commits-merged.texttext-generation1M<n<10M1 likes758 downloads2mo agoHugging Face25bigcode /commits-codegeex Dataset Card for "commits-codegeex" More Information needed text1M<n<10M6 likes738 downloads3y agoHugging Face26JetBrains-Research /commit-msg-edits ✍️ Commit Message Edits Dataset This dataset is a collection of expert-labeled commit message edits contributed via Commit Message Editing app presented in Towards Realistic Evaluation of Commit Message Generation by Matching Online and Offline Settings. Labelers were presented with GPT-4 generated messages for 15 commits from CMG benchmark from Long Code Arena and asked to manually edit them to be of good enough quality to submit to VCS. You can check Manual tab in our… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/commit-msg-edits.textn<1K1 likes734 downloads2y agoHugging Face27bigcode /commits_ftCode Commits for Instruction Tuning0 likes695 downloads3y agoHugging Face28AdithyaSK /repo2rlenv-commit-runtime-test repo2rlenv-commit-runtime Generated by Repo2RLEnv — turning real GitHub repositories into verifiable RL environments. 💡 Browse this dataset in your browser — click the badge above or open HuggingFaceH4/harbor-visualiser to inspect every task's spec, instruction, oracle patch, test script, and Dockerfile. Source repos (12): encode/httpx gin-gonic/gin gorilla/mux pallets/click pallets/werkzeug pocketbase/pocketbase psf/requests python-attrs/attrs sirupsen/logrus spf13/cobra… See the full description on the dataset page: https://huggingface.co/datasets/AdithyaSK/repo2rlenv-commit-runtime-test.n<1K0 likes688 downloads3mo agoHugging Face29andstor /cvevc_commits Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/andstor/cvevc_commits.text1M<n<10M0 likes664 downloads8mo agoHugging Face302doo /conventional-commit-messages0 likes610 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.