datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikidata-extraction
Wikidata Extraction
This dataset contains all RDF triples extracted from the latest Wikidata, converted from the N-Triples format to Parquet.
The data originates from Wikidata, a free and open knowledge base that acts as central storage for structured data used by Wikipedia and other Wikimedia projects. The source file is the "truthy" N-Triples dump (latest-truthy.nt.bz2), which contains only the current, non-deprecated statements.
The code to extract this data is available at… See the full description on the dataset page: https://huggingface.co/datasets/piebro/wikidata-extraction.invoice-extraction-benchmark
Invoice Extraction Benchmark v1
A synthetic test set for invoice data extraction (invoice OCR, intelligent document
processing, accounts-payable capture): 181 documents with answer keys and a scorer.
Run any invoice reader over the documents, write its output as one JSON file, and score it
field by field.
Source, scorer and generator: https://github.com/DrewKraken/invoice-extraction-benchmark
(this dataset is a mirror of corpus/v1/ there; the GitHub repository is canonical).
Who… See the full description on the dataset page: https://huggingface.co/datasets/drew-ipp/invoice-extraction-benchmark.raw-fact-extractionfunding-extraction-harness-benchmarktooth_extraction_4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 200,
"total_frames": 76053,
"total_tasks": 1,
"total_videos": 400,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/samsam0510/tooth_extraction_4.ntsb-accident-extraction
Files
data/: the train, val, and test splits used for finetuning and inference
raw.jsonl: every cleaned record with its narrative, reference fields, and split, used as
evaluation ground truth
preparation_config.yaml: the exact cleaning, splitting, and prompt configuration used
vocab_schema.json: the closed vocabulary for every categorical field, as a JSON Schema
Processing
Combined source splits: train
Document length: 300 to
30,000 characters
Deduplicated by… See the full description on the dataset page: https://huggingface.co/datasets/FuzzyLabs/ntsb-accident-extraction.fcv-extractions-meta
fcv-extractions-meta
Data-use mention extractions from the World Bank Fragility, Conflict and Violence (FCV) document corpus, enriched by the fine-tuned rafmacalaba/lfm2.5-350M-datause-multitask model (a LoRA SFT of LiquidAI/LFM2.5-350M).
Shape
nested — one row per chunk; each entity in entities[] carries provenance and usage_impact alongside its span.
Configs
config
rows
fcv_pads_east_africa
793,763
jdc_operational
12,372… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/fcv-extractions-meta.tooth_extraction_3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 100,
"total_frames": 32879,
"total_tasks": 1,
"total_videos": 200,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/samsam0510/tooth_extraction_3.arxiv-funding-entity-extractions
arxiv-funding-entity-extractions
Funder/award entity extractions over cometadata/arxiv-funding-statements.
Extractor: funding-entity-extractor (vLLM + LoRA)
Base model: meta-llama/Llama-3.1-8B-Instruct
LoRA: cometadata/funding-extraction-llama-3.1-8b-instruct-artifact-data-mix-grpo-mixed-reward
Hardware: A100-large bf16, concurrency 256
Total rows: 1,823,650
Configs
predictions (default) — original extractions, no ROR enrichment.
predictions_with_ror — same rows… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-funding-entity-extractions.autonomous-driving-intention-field-extraction-v0.1What this dataset tests
Whether a system can infer agent intentions
from context cues in complex driving scenes.
This is not trajectory prediction.
It is intention inference.
Required outputs
agent_id
inferred_intention
intention_confidence
time_horizon_s
alternative_intentions
stability_score
Scoring conventions
confidence and stability range 0 to 1
time horizon is seconds into the near future
Use case
Layer one of Intention Field and Social Coherence Maps.
This enables… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-intention-field-extraction-v0.1.ru-invoice-extraction-benchmark
Набор для извлечения данных из русскоязычных счетов
50 синтетических русскоязычных счетов с эталонными ответами в JSON. Набор не привязан к определённой модели и предназначен для проверки систем извлечения структурированных данных из текста.
Разделы
Раздел
Документы
Назначение
development
30
разработка шаблонов и примеров
validation
10
выбор настроек
test
10
итоговая оценка
Раздел test зафиксирован для версии 1. Его нельзя использовать для… See the full description on the dataset page: https://huggingface.co/datasets/necrasov-ilya/ru-invoice-extraction-benchmark.paper_extractionextraction-examples
Extraction Examples Dataset
This dataset contains 17 examples for testing extraction workflows.
Dataset Structure
Each example includes:
PDF file: Original document
map_info.json: Map extraction metadata
direction.json: Direction information
GeoJSON files: Polygon geometries
Area JSON files: Area definitions
File Organization
files/
├── example1/
│ ├── document.pdf
│ ├── map_info.json
│ ├── direction.json
│ ├── polygon1.geojson
│ └── area1.json… See the full description on the dataset page: https://huggingface.co/datasets/alexdzm/extraction-examples.feature-extraction-checkpoint-downloadsjoa-extraction-arena
JOA Job-Posting Extraction Arena
Disclosure: Job Opportunities API (JOA, jobopportunitiesapi.org) is an independent data business that sells API access to job-posting data. AI helped run the experiments, check the numbers and draft this text; Loukas (Luca) Tzekos is editorially responsible. Contact: hello@jobopportunitiesapi.org.
A benchmark for one narrow task: reading a real job posting and filling 11 structured fields. Job Opportunities API (JOA) built it to choose and train… See the full description on the dataset page: https://huggingface.co/datasets/JobOpportunitiesAPI/joa-extraction-arena.paper_extraction_v1eval_act_tooth_extraction_3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 10,
"total_frames": 17799,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/samsam0510/eval_act_tooth_extraction_3.ki_extraction-4b-lmeval-qwen2.5-7b-instruct-on-mmlu_pro-0shot_cot-scillm-5d1468e6e9Dicom-metadata-extraction-skillbenchfinemath-qa-extraction-test
Dataset Card for finemath-qa-extraction-test
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/gabrielmbmb/finemath-qa-extraction-test/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/gabrielmbmb/finemath-qa-extraction-test.putusan-structured-extraction
Putusan structured-extraction dataset
Built 2026-07-08T23:02:28+00:00 by notebooks/build_dataset.py (seed 3407).
Indonesian court-decision (putusan) extractive-structuring dataset over three
corpora (Anak, Asusila, TPPO). Each row is one model extraction of one source
document into 31 canonical sections of verbatim spans. Empty sections were
completed from sibling model extractions of the same document where available
(cross_model_fill_json records per-section donor provenance).… See the full description on the dataset page: https://huggingface.co/datasets/Haeryz/putusan-structured-extraction.fcv-extractions-meta-tiered-probe
fcv-extractions-meta-tiered-probe
Data-use mention extractions from the World Bank Fragility, Conflict and Violence (FCV) document corpus. Spans are extracted by the fine-tuned GLiNER model rafmacalaba/gliner_datause_tiered, scored by the tier-probe head rafmacalaba/gliner-tier-probe, and attributed (provenance + usage/impact) by rafmacalaba/lfm2.5-350M-datause-multitask-tiered (a LoRA SFT of LiquidAI/LFM2.5-350M).
Shape
nested — one row per chunk; each entity in… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/fcv-extractions-meta-tiered-probe.fcv-extractions-meta-tiered
fcv-extractions-meta-tiered
Data-use mention extractions from the World Bank Fragility, Conflict and Violence (FCV) document corpus, enriched by the fine-tuned rafmacalaba/lfm2.5-350M-datause-multitask model (a LoRA SFT of LiquidAI/LFM2.5-350M).
Shape
nested — one row per chunk; each entity in entities[] carries provenance and usage_impact alongside its span.
Configs
config
rows
fcv_pads_east_africa
862,663
jdc_operational
2,468… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/fcv-extractions-meta-tiered.mteb-tweet_sentiment_extraction-avs_triplets
MTEB Tweet Sentiment Extraction Triplets Dataset
This dataset was used in the paper GISTEmbed: Guided In-sample Selection of Training Negatives for Text Embedding Fine-tuning. Refer to https://arxiv.org/abs/2402.16829 for details.
The code for generating the data is available at https://github.com/avsolatorio/GISTEmbed/blob/main/scripts/create_classification_dataset.py.
Citation
@article{solatorio2024gistembed,
title={GISTEmbed: Guided In-sample Selection of… See the full description on the dataset page: https://huggingface.co/datasets/avsolatorio/mteb-tweet_sentiment_extraction-avs_triplets.unitization-before-extraction
Unitization Before Extraction — the 2026bl run
Records, adjudicated inventories, per-call metadata and derived tables for a pre-registered
diagnostic replication that decomposes disagreement in machine recovery of document-scale
argument dependency structure.
The design was deposited before any datum existed
(10.5281/zenodo.21830221, 2026-08-07); the run
executed 2026-08-08. Every threshold in the analysis preceded the data, and that ordering is a
checkable timestamp rather than… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/unitization-before-extraction.putusan-windowed-extraction
Putusan windowed line-anchored extraction dataset (Plan B)
Built 2026-07-09T13:01:27+00:00 by notebooks/build_windowed_dataset.py from the
legacy Haeryz/putusan-structured-extraction dataset (same documents, same
leakage-safe purpose/split assignment, seed 3407).
Each legacy document row (~34K tokens median — longer than a 32K context)
is re-expressed as overlapping line-numbered windows of <= 6400
content tokens (measured with Qwen/Qwen3.5-9B; fits a
max_seq_length of 8192 with… See the full description on the dataset page: https://huggingface.co/datasets/Haeryz/putusan-windowed-extraction.arxiv-author-affiliation-extraction-inference-inputs-metadatadeepseek-ai__deepseek-llm-7b-chathealth-log-extraction-datasetTo create this dataset, we sampled demographic seeds from an occupation and age-range table from the Labor Force Statistics[1].
For each occupation, the pipeline randomly selected an age range, sampled an age within that range, and assigned a gender from a fixed set of options.
These demographic seeds were used to prompt an LLM to generate structured personas containing a name, description, medications or supplements, general mood, and possible health conditions or injuries.
We then used… See the full description on the dataset page: https://huggingface.co/datasets/lbakar/health-log-extraction-dataset.Goekdeniz-Guelmez__Josiefied-Qwen2.5-1.5B-Instruct-abliterated-v3
