Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ASSERT-KTH /agent-commits-raw AI Coding-Agent Commits on GitHub This dataset documents commits associated with four AI coding agents: Claude, OpenAI Codex, GitHub Copilot and Cursor. It contains 1,853,915 commit records across 444,055 GitHub repositories, with 220,753 identifiable GitHub user accounts recorded as commit authors. Messages and author identities are in commits; repository metadata, file changes and patch text are available in separate tables. Dataset Agent Commit records… See the full description on the dataset page: https://huggingface.co/datasets/ASSERT-KTH/agent-commits-raw.tabular10M<n<100M0 likes4.9k downloads15d agoHugging Face02djtereano /commitbench_conventionaltext10K<n<100K0 likes2.6k downloads1y agoHugging Face03commit0 /commit0textn<1K3 likes2k downloads2y agoHugging Face04JetBrains-Research /commit-chronicle 📜 CommitChronicle 🔮 This is the dataset for commit message generation (and/or completion), introduced in the paper "From Commit Message Generation to History-Aware Commit Message Completion", ASE 2023. Its key features: large-scale and multilingual: contains 10.7M commits from 11.9k GitHub repositories in 20 programming languages; diverse: avoids restrictive filtering on commit messages or commit diffs structure; suitable for experiments with commit history: provides metadata… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/commit-chronicle.tabulartext-generation10M<n<100M13 likes1.6k downloads3y agoHugging Face05bigcode /github-commits-diff-dedup-pjjs-april Deduplicated Commits Deduplicated based on diff: content = '\n'.join(difflib.unified_diff( old_content.splitlines(keepends=True), new_content.splitlines(keepends=True), n=5 )) Parameters: Minimum ngram size: 5 MinHash ngram size: 5 MinHash threshold: 0.8 text100K<n<1M4 likes1k downloads3y agoHugging Face06ZipLime /commitments-of-traders Commitments of Traders Who was long and who was short in every US futures market — and, for once, when anyone could actually see it. 421 223 market-weeks · 211 071 point-in-time rows · 2 585 weekly releases · 762 markets · 2010-01-05 to 2026-09-08 The pipeline lives in recipe/ at the same revision as the data. See PIPELINE.md for the method. The data is from Tuesday. It comes out on Friday. A COT report is taken as of the close on Tuesday and published at 3:30… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/commitments-of-traders.tabulartabular-regression100K<n<1M0 likes982 downloads4h agoHugging Face07bigcode /commitpackmetaGitHub metadata for https://huggingface.co/datasets/bigcode/commitpack text10M<n<100M4 likes967 downloads3y agoHugging Face08wentingzhao /commit0_combinedtextn<1K0 likes817 downloads2y agoHugging Face09placeholderlabs /pretrain-commits-v2-mixNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 79,318,557,379 (79.3B) Trainable tokens 28,772,968,648 (28.8B) Documents 31,431,846 Shards 696 UTF-8 bytes 310,412,170,445 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-commits-v2-mix.tabular10M<n<100M0 likes795 downloads24d agoHugging Face10Hirunima /commit-chronicle 📜 CommitChronicle 🔮 This is the dataset for commit message generation (and/or completion), introduced in the paper "From Commit Message Generation to History-Aware Commit Message Completion", ASE 2023. Its key features: large-scale and multilingual: contains 10.7M commits from 11.9k GitHub repositories in 20 programming languages; diverse: avoids restrictive filtering on commit messages or commit diffs structure; suitable for experiments with commit history: provides metadata… See the full description on the dataset page: https://huggingface.co/datasets/Hirunima/commit-chronicle.tabulartext-generation10M<n<100M0 likes750 downloads7mo agoHugging Face11andstor /cvevc_commits Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/andstor/cvevc_commits.text1M<n<10M0 likes631 downloads8mo agoHugging Face12bigcode /commits-codegeex Dataset Card for "commits-codegeex" More Information needed text1M<n<10M6 likes616 downloads3y agoHugging Face13project-themis /git-commits-merged Themis-Git-Commits-Merged Overview Themis-Git-Commits-Merged is a large-scale dataset of ~3.98M single-file code commits from permissively licensed GitHub repositories that have been cross-referenced with GHTorrent pull request data to retain only commits that are part of successfully merged, non-reverted pull requests. This provides implicit human validation of each code change — a merge decision by project maintainers confirms the intent and quality of… See the full description on the dataset page: https://huggingface.co/datasets/project-themis/git-commits-merged.texttext-generation1M<n<10M1 likes599 downloads2mo agoHugging Face14placeholderlabs /pretrain-commits-v2-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 12,571,681,749 (12.6B) Trainable tokens 4,460,160,435 (4.5B) Documents 992,475 Shards 327 UTF-8 bytes 49,288,867,997 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-commits-v2-mix-long-context.tabular1M<n<10M0 likes589 downloads24d agoHugging Face15bigcode /commits-pjj-diff Dataset Card for "commits-pjj-diff" More Information needed text1M<n<10M2 likes369 downloads4y agoHugging Face16ObscuraCoder /commit-chronicleThis is a filtered version of the JetBrains-Research/commit-chronicle dataset. It has been subsetted for the following languages: [ "C", "C++", "Go", "Java", "Python", "Rust", "TypeScript" ] Further filtering steps undertaken are: Useless features have been removed and only the message and diff retained. Only commits that modify a single file have been chosen. Samples containing diffs longer than 1024 tokens (by the ObscuraCoder/Tokenizer tokenizer estimate) have been discarded.text1M<n<10M4 likes337 downloads2y agoHugging Face17JetBrains-Research /lca-commit-message-generation 🏟️ Long Code Arena (Commit message generation) This is the benchmark for the Commit message generation task as part of the 🏟️ Long Code Arena benchmark. The dataset is a manually curated subset of the Python test set from the 🤗 CommitChronicle dataset, tailored for larger commits. All the repositories are published under permissive licenses (MIT, Apache-2.0, and BSD-3-Clause). The datapoints can be removed upon request. How-to from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-commit-message-generation.text1K<n<10K0 likes330 downloads2y agoHugging Face18knesset-asr /knesset-committees-chunkstabular1M<n<10M0 likes324 downloads1mo agoHugging Face19PiotrSty /sejm-committee-transcripts Polish Sejm committee transcripts — full API coverage (terms 9 and 10) Official committee transcripts ("pełny zapis przebiegu posiedzenia") from the Sejm of the Republic of Poland, parsed into individually attributed speaker turns. Scope Committees: all standing committees with zapis PDFs in the Sejm API. Terms: 9 (2019-11-14 → 2023-11-09) and 10 (2023-11-14 → 2026-09-17). Provider and primary source: Kancelaria Sejmu RP, https://api.sejm.gov.pl/… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/sejm-committee-transcripts.tabulartext-generation100K<n<1M1 likes236 downloads22d agoHugging Face20chargoddard /commitpack-ft-instruct-ratedThis is commitpack-ft-instruct, derived from Octocode's CommitPackFT, augmented with a quality analysis of the instruction-response pair by a local model. This did a pretty decent job of identifying pairs that obviously don't have enough context to know what change is being requested, or where the commit message does not match with the changes made. Data files (yaml, plain text, json, etc.) were heavily downsampled in preparing this dataset to skew it more towards actual code work. All entries… See the full description on the dataset page: https://huggingface.co/datasets/chargoddard/commitpack-ft-instruct-rated.text100K<n<1M4 likes188 downloads3y agoHugging Face21mamiksik /processed-commit-diffs List of repositories included in the dataset Project Language Fetched Count Url Moby Go 5 943 https://github.com/moby/moby Rxjava Java 516 https://github.com/RxJava/ReactiveX Spring-framework Java 2 529 https://github.com/spring-framework/spring-project Chart.js Javascript 641 https://github.com/Chart.js/chartjs Three.js Javascript 1 512https://github.com/three.js/mrdoob Redux Javascript 592 https://github.com/redux/reduxjs React-native Javascript 2 901… See the full description on the dataset page: https://huggingface.co/datasets/mamiksik/processed-commit-diffs.text10K<n<100K5 likes181 downloads4y agoHugging Face22Farmaanaa /cftc_commitments_of_traders_positioning موقعیت سفته‌بازان در بازار آتی کالا (گزارش COT) — هفتگی موقعیت خرید و فروش صندوق‌های سفته‌باز در بازار آتی نفت، طلا، نقره، مس، گاز، گندم و ذرت — هفتگی از ۲۰۱۵. در کنار سری قیمت همان کالاها، نشان می‌دهد حرکت قیمت را انتظارات می‌سازد یا واقعیت بازار. پوشش: 1393-10-16 → 1405-06-24 · تناوب: هفتگی · سطح: بازارهای آتی آمریکا تعداد مشاهده: 5,499 · تعداد مکان: 1 منبع: کمیسیون معاملات آتی کالای آمریکا (CFTC) — https://www.cftc.gov/MarketReports/CommitmentsofTraders… See the full description on the dataset page: https://huggingface.co/datasets/Farmaanaa/cftc_commitments_of_traders_positioning.tabular1K<n<10K0 likes173 downloads17d agoHugging Face23chargoddard /commitpack-ft-instructOctocode's CommitPackFT in Alpaca instruction format, with several randomly selected natural language preludes to the commit messages to make them better resemble a user request. When the instruction, old code, and new code combined are small enough to fit within 4096 Llama tokens the output is usually the full contents of the file after a commit. Otherwise, the output will be a sequence of ndiff chunks with up to five lines of context each. An example: ```ndiff from… See the full description on the dataset page: https://huggingface.co/datasets/chargoddard/commitpack-ft-instruct.text100K<n<1M3 likes164 downloads3y agoHugging Face24bigcode /commits_sample_filesThis is a sample of GitHub commits and files reconstructed from the Software Heritage dataset. It contains the latest 1024 commits in all of pytorch/* and huggingface/* repos. The tables are split to avoid an explosion of rows (lots of repeated files between commits), so you will need to pre-filter the commits before adding the file contents. Table descriptions: 1. commits The commit message table. Join it with commit_filepath on commits.directory_id == commit_filepath.directory_id… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/commits_sample_files.text10M<n<100M1 likes163 downloads3y agoHugging Face25balrampandey /qmmit-open-source-agent-commit-index Repository-Level Measurement of Self-Declared Coding-Agent Commit Signatures Dataset release: 2026-09-18-v3.0Schema: 3.0.0 Release stamp: dataset 2026-09-18-v3.0 · ruleset sha256:b2e8889c66f72c18a61839f0bf1a9f77b5481ba2def044dba33c797b2f2bdcae · scanned 2026-09-16 Abstract This dataset contains 2000 repository-level observations from public Git repositories. Each observation estimates a lower bound on the proportion of non-merge, non-infrastructure-bot commits… See the full description on the dataset page: https://huggingface.co/datasets/balrampandey/qmmit-open-source-agent-commit-index.tabulartabular-classification1K<n<10K0 likes163 downloads22d agoHugging Face26bigcode /commits-pjj-2048 Dataset Card for "commits-pjj-2048" More Information needed text1M<n<10M0 likes156 downloads4y agoHugging Face27andstor /cvevc_cve_commit_mappings CVEVC Commit CVE Mappings Mapping between commmits and cves. Example usage: from datasets import load_dataset cve_ds = load_dataset('fals3/cvevc_cve', split='test') commit_ds = load_dataset('fals3/cvevc_commits', 'patches', split='test') commit_cve_mapping = load_dataset('fals3/cvevc_cve_commit_mappings', split='test') cve_df = cve_ds.to_polars() commit_df = commit_ds.to_polars() commit_cve_mapping_df = commit_cve_mapping.to_polars() combined_df =… See the full description on the dataset page: https://huggingface.co/datasets/andstor/cvevc_cve_commit_mappings.text10M<n<100M0 likes144 downloads8mo agoHugging Face28stallone /CommitPackFTtext1M<n<10M1 likes141 downloads2y agoHugging Face29placeholderlabs /exp-pool-commit-code-raw Locus EXP Commit Code - shuffled raw proxy pool Deterministically shuffled commit-message and unified-diff documents with complete source metadata. MANIFEST.json pins source identity, sampling policy, token budgets, and per-file checksums. The paired Dolma-2-tokenized repository preserves prompt masking for reproducible proxy training. tabular1M<n<10M0 likes134 downloads2mo agoHugging Face30climatebert /climate_commitments_actions Dataset Card for climate_commitments_actions Dataset Summary We introduce an expert-annotated dataset for identifying climate-related paragraphs about climate commitments and actions in corporate disclosures. Supported Tasks and Leaderboards The dataset supports a binary classification task of whether a given climate-related paragraph is about climate commitments and actions or not. Languages The text in the dataset is in English. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/climatebert/climate_commitments_actions.texttext-classification1K<n<10K9 likes133 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.