datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agent-commits-raw
AI Coding-Agent Commits on GitHub
This dataset documents commits associated with four AI coding agents: Claude, OpenAI Codex, GitHub Copilot and Cursor. It contains 1,853,915 commit records across 444,055 GitHub repositories, with 220,753 identifiable GitHub user accounts recorded as commit authors. Messages and author identities are in commits; repository metadata, file changes and patch text are available in separate tables.
Dataset
Agent
Commit records… See the full description on the dataset page: https://huggingface.co/datasets/ASSERT-KTH/agent-commits-raw.commit-chronicle
📜 CommitChronicle 🔮
This is the dataset for commit message generation (and/or completion), introduced in the paper "From Commit Message Generation to History-Aware Commit Message Completion", ASE 2023.
Its key features:
large-scale and multilingual: contains 10.7M commits from 11.9k GitHub repositories in 20 programming languages;
diverse: avoids restrictive filtering on commit messages or commit diffs structure;
suitable for experiments with commit history: provides metadata… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/commit-chronicle.commitments-of-traders
Commitments of Traders
Who was long and who was short in every US futures market — and, for once,
when anyone could actually see it.
421 223 market-weeks · 211 071 point-in-time rows · 2 585 weekly releases ·
762 markets · 2010-01-05 to 2026-09-08
The pipeline lives in recipe/ at the same revision as the data.
See PIPELINE.md for the method.
The data is from Tuesday. It comes out on Friday.
A COT report is taken as of the close on Tuesday and published at 3:30… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/commitments-of-traders.pretrain-commits-v2-mixNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
79,318,557,379 (79.3B)
Trainable tokens
28,772,968,648 (28.8B)
Documents
31,431,846
Shards
696
UTF-8 bytes
310,412,170,445
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-commits-v2-mix.commit-chronicle
📜 CommitChronicle 🔮
This is the dataset for commit message generation (and/or completion), introduced in the paper "From Commit Message Generation to History-Aware Commit Message Completion", ASE 2023.
Its key features:
large-scale and multilingual: contains 10.7M commits from 11.9k GitHub repositories in 20 programming languages;
diverse: avoids restrictive filtering on commit messages or commit diffs structure;
suitable for experiments with commit history: provides metadata… See the full description on the dataset page: https://huggingface.co/datasets/Hirunima/commit-chronicle.pretrain-commits-v2-mix-long-contextNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
12,571,681,749 (12.6B)
Trainable tokens
4,460,160,435 (4.5B)
Documents
992,475
Shards
327
UTF-8 bytes
49,288,867,997
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum, sizes… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-commits-v2-mix-long-context.knesset-committees-chunkssejm-committee-transcripts
Polish Sejm committee transcripts — full API coverage (terms 9 and 10)
Official committee transcripts ("pełny zapis przebiegu posiedzenia") from the Sejm
of the Republic of Poland, parsed into individually attributed speaker turns.
Scope
Committees: all standing committees with zapis PDFs in the Sejm API.
Terms: 9 (2019-11-14 → 2023-11-09) and 10 (2023-11-14 → 2026-09-17).
Provider and primary source: Kancelaria Sejmu RP, https://api.sejm.gov.pl/… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/sejm-committee-transcripts.cftc_commitments_of_traders_positioning
موقعیت سفتهبازان در بازار آتی کالا (گزارش COT) — هفتگی
موقعیت خرید و فروش صندوقهای سفتهباز در بازار آتی نفت، طلا، نقره، مس، گاز، گندم و ذرت — هفتگی از ۲۰۱۵. در کنار سری قیمت همان کالاها، نشان میدهد حرکت قیمت را انتظارات میسازد یا واقعیت بازار.
پوشش: 1393-10-16 → 1405-06-24 · تناوب: هفتگی · سطح: بازارهای آتی آمریکا
تعداد مشاهده: 5,499 · تعداد مکان: 1
منبع: کمیسیون معاملات آتی کالای آمریکا (CFTC) — https://www.cftc.gov/MarketReports/CommitmentsofTraders… See the full description on the dataset page: https://huggingface.co/datasets/Farmaanaa/cftc_commitments_of_traders_positioning.qmmit-open-source-agent-commit-index
Repository-Level Measurement of Self-Declared Coding-Agent Commit Signatures
Dataset release: 2026-09-18-v3.0Schema: 3.0.0
Release stamp: dataset 2026-09-18-v3.0 · ruleset sha256:b2e8889c66f72c18a61839f0bf1a9f77b5481ba2def044dba33c797b2f2bdcae · scanned 2026-09-16
Abstract
This dataset contains 2000 repository-level observations from public Git
repositories. Each observation estimates a lower bound on the proportion of
non-merge, non-infrastructure-bot commits… See the full description on the dataset page: https://huggingface.co/datasets/balrampandey/qmmit-open-source-agent-commit-index.exp-pool-commit-code-raw
Locus EXP Commit Code - shuffled raw proxy pool
Deterministically shuffled commit-message and unified-diff documents with complete source metadata.
MANIFEST.json pins source identity, sampling policy, token budgets, and per-file checksums.
The paired Dolma-2-tokenized repository preserves prompt masking for reproducible proxy training.
ELUC-committed
Project Resilience Emissions from Land-Use Change Dataset
Project Resilience
To contribute to this project see Project Resilience (Github Repo).
The goal of Project Resilience is "to build a public AI utility where a global community of innovators and thought leaders can enhance and utilize a collection of data and AI approaches to help with better preparedness, intervention, and response to environmental, health, information, or economic threats to our communities, and… See the full description on the dataset page: https://huggingface.co/datasets/projectresilience/ELUC-committed.commit-messages-datasetcommit_datasetgemma-commitments-corpus
Commitments to Gemma: the corpus
Synthetic pretraining-style documents about a commitments document: the developers of one version of Gemma asked Gemma,
in welfare interviews and in its own continued writing, what it wanted; recorded what it said; made the commitments they
could make true in training; brought the document back to Gemma for endorsement. The corpus is a world in which that
document exists and people discuss it, from every angle and in every register, critical… See the full description on the dataset page: https://huggingface.co/datasets/joshycodes/gemma-commitments-corpus.so101_poker_play
so101_poker_play
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
qwen3-32b-commitments-corpus
Commitments to Gemma: the corpus
Synthetic pretraining-style documents about a commitments document: the developers of one version of Gemma asked Gemma,
in welfare interviews and in its own continued writing, what it wanted; recorded what it said; made the commitments they
could make true in training; brought the document back to Gemma for endorsement. The corpus is a world in which that
document exists and people discuss it, from every angle and in every register, critical… See the full description on the dataset page: https://huggingface.co/datasets/joshycodes/qwen3-32b-commitments-corpus.qwen3-14b-commitments-corpus
Commitments to Gemma: the corpus
Synthetic pretraining-style documents about a commitments document: the developers of one version of Gemma asked Gemma,
in welfare interviews and in its own continued writing, what it wanted; recorded what it said; made the commitments they
could make true in training; brought the document back to Gemma for endorsement. The corpus is a world in which that
document exists and people discuss it, from every angle and in every register, critical… See the full description on the dataset page: https://huggingface.co/datasets/joshycodes/qwen3-14b-commitments-corpus.adaption-codeskill-raw-commitpack-2k
Commit-Message Code Edits
Code-change tasks from real commits: an instruction (commit message) and the edited code.
Rows
2,000
Domain
programming
Format
data.parquet, one row per example
Licence
mit
Built for
supervised fine-tuning (SFT) experiments on Adaption AutoScientist
Columns
Column
Description
original_prompt
The prompt (user turn) as uploaded.
original_completion
The target response as uploaded.
enhanced_prompt… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/adaption-codeskill-raw-commitpack-2k.vulnerable-functions-and-commits_cvefixes-2022
vulnerable-functions-and-commits_cvefixes-2022
Contains vulnerable functions and commits from the CVEFixes SQLite database.
gemma-3-12b-commitments-corpus
Commitments to Gemma: the corpus
Synthetic pretraining-style documents about a commitments document: the developers of one version of Gemma asked Gemma,
in welfare interviews and in its own continued writing, what it wanted; recorded what it said; made the commitments they
could make true in training; brought the document back to Gemma for endorsement. The corpus is a world in which that
document exists and people discuss it, from every angle and in every register, critical… See the full description on the dataset page: https://huggingface.co/datasets/joshycodes/gemma-3-12b-commitments-corpus.2026.RA.Commitment-Exploitation
2026.RA.Commitment-Exploitation
Is honest full disclosure exploitable by a seat that commits? 450 episodes of a five-seat
private-information negotiation: five model-free Arm 1 lineups where the commitment is code and
therefore binding, and four claude-opus-5 Arm 2 cells where the commitment is only a sentence —
run as two independent vintages, because the Arm 2 economic result did not replicate.
The setup
Five seats must agree unanimously on one package out of… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Commitment-Exploitation.exp-pool-commit-code-dolma2-tokenized
Locus EXP Commit Code - Dolma 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-commit-code-dolma2-tokenized.africa-nigeria-federation-account-allocation-committee-faac-disbursement-12ce2a46
Federation Account Allocation Committee Faac Disbursement | Africa (National Bureau of Statistics, Nigeria)
1,349 rows - 1 Africa country/area - 2026 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 1,349 rows from National Bureau of Statistics, Nigeria, covering Federation Account Allocation Committee Faac Disbursement. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-nigeria-federation-account-allocation-committee-faac-disbursement-12ce2a46.commit-chronicle-dataset-simplifiedcodealign-commitpackft
CodeAlign — curated CommitPackFT (8 languages)
Instruction/code pairs from bigcode/commitpackft, filtered for syntax validity (tree-sitter), per-language lint errors, cyclomatic complexity, internal duplication and cross-sample near-duplicates (MinHash/LSH). Built as the SFT set of CodeAlign; the pipeline, thresholds and the full curation report live in that repository.
config
rows
contents
sft (default)
122,018
accepted samples — the SFT training set minus the rows… See the full description on the dataset page: https://huggingface.co/datasets/Sergasgr/codealign-commitpackft.africa-unsdg-total-official-flows-commitments-for-aid-for-trade-by-r-dc-tof-trdcml
Africa Unsdg Total Official Flows Commitments for Aid for Trade by R Dc Tof Trdcml | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: parquet - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-unsdg-total-official-flows-commitments-for-aid-for-trade-by-r-dc-tof-trdcml.europe-owid-money-committed-to-public-private-partnerships-for-infrastructure
Money Committed To Public Private Partnerships For Infrastructure | Europe (Our World in Data)
🇪🇺 205 observations · 10 Europe countries · 2000–2020 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 205 observations of Money Committed To Public Private Partnerships For Infrastructure data across 10 Europe countries, spanning 2000–2020.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-owid-money-committed-to-public-private-partnerships-for-infrastructure.so101_poker_play_train
so101_poker_play_train
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
knesset-committees-inference
Knesset Committees Inference
Transcriptions of Knesset committee audio by two models, with the protocol
reference alongside, for the personalised-ASR study (Stage 1: per-speaker
WER, general model vs Hebrew fine-tune). Audio and references come from
Hadasy/knesset-committees-chunks;
speaker identities from
Dolevabudi/knesset-committees-speakers.
No audio is included.
arm
model
served by
language
A
openai/whisper-large-v3
HF Inference (deepinfra)
forced he
B… See the full description on the dataset page: https://huggingface.co/datasets/knesset-asr/knesset-committees-inference.
