datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
invoice-extraction-benchmark
Invoice Extraction Benchmark v1
A synthetic test set for invoice data extraction (invoice OCR, intelligent document
processing, accounts-payable capture): 181 documents with answer keys and a scorer.
Run any invoice reader over the documents, write its output as one JSON file, and score it
field by field.
Source, scorer and generator: https://github.com/DrewKraken/invoice-extraction-benchmark
(this dataset is a mirror of corpus/v1/ there; the GitHub repository is canonical).
Who… See the full description on the dataset page: https://huggingface.co/datasets/drew-ipp/invoice-extraction-benchmark.fcv-extractions-meta
fcv-extractions-meta
Data-use mention extractions from the World Bank Fragility, Conflict and Violence (FCV) document corpus, enriched by the fine-tuned rafmacalaba/lfm2.5-350M-datause-multitask model (a LoRA SFT of LiquidAI/LFM2.5-350M).
Shape
nested — one row per chunk; each entity in entities[] carries provenance and usage_impact alongside its span.
Configs
config
rows
fcv_pads_east_africa
793,763
jdc_operational
12,372… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/fcv-extractions-meta.ru-invoice-extraction-benchmark
Набор для извлечения данных из русскоязычных счетов
50 синтетических русскоязычных счетов с эталонными ответами в JSON. Набор не привязан к определённой модели и предназначен для проверки систем извлечения структурированных данных из текста.
Разделы
Раздел
Документы
Назначение
development
30
разработка шаблонов и примеров
validation
10
выбор настроек
test
10
итоговая оценка
Раздел test зафиксирован для версии 1. Его нельзя использовать для… See the full description on the dataset page: https://huggingface.co/datasets/necrasov-ilya/ru-invoice-extraction-benchmark.Dicom-metadata-extraction-skillbenchfcv-extractions-meta-tiered-probe
fcv-extractions-meta-tiered-probe
Data-use mention extractions from the World Bank Fragility, Conflict and Violence (FCV) document corpus. Spans are extracted by the fine-tuned GLiNER model rafmacalaba/gliner_datause_tiered, scored by the tier-probe head rafmacalaba/gliner-tier-probe, and attributed (provenance + usage/impact) by rafmacalaba/lfm2.5-350M-datause-multitask-tiered (a LoRA SFT of LiquidAI/LFM2.5-350M).
Shape
nested — one row per chunk; each entity in… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/fcv-extractions-meta-tiered-probe.fcv-extractions-meta-tiered
fcv-extractions-meta-tiered
Data-use mention extractions from the World Bank Fragility, Conflict and Violence (FCV) document corpus, enriched by the fine-tuned rafmacalaba/lfm2.5-350M-datause-multitask model (a LoRA SFT of LiquidAI/LFM2.5-350M).
Shape
nested — one row per chunk; each entity in entities[] carries provenance and usage_impact alongside its span.
Configs
config
rows
fcv_pads_east_africa
862,663
jdc_operational
2,468… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/fcv-extractions-meta-tiered.unitization-before-extraction
Unitization Before Extraction — the 2026bl run
Records, adjudicated inventories, per-call metadata and derived tables for a pre-registered
diagnostic replication that decomposes disagreement in machine recovery of document-scale
argument dependency structure.
The design was deposited before any datum existed
(10.5281/zenodo.21830221, 2026-08-07); the run
executed 2026-08-08. Every threshold in the analysis preceded the data, and that ordering is a
checkable timestamp rather than… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/unitization-before-extraction.arxiv-author-affiliation-extraction-inference-inputs-metadatahealth-log-extraction-datasetTo create this dataset, we sampled demographic seeds from an occupation and age-range table from the Labor Force Statistics[1].
For each occupation, the pipeline randomly selected an age range, sampled an age within that range, and assigned a gender from a fixed set of options.
These demographic seeds were used to prompt an LLM to generate structured personas containing a name, description, medications or supplements, general mood, and possible health conditions or injuries.
We then used… See the full description on the dataset page: https://huggingface.co/datasets/lbakar/health-log-extraction-dataset.real-human-logs-extraction-datasetTo create this dataset, we collected human-generated logs from two individuals. This needs to be beefed up in the future, but this is what we have for now.
Subsequently, we ran NuExtract3 on each example of the dataset to get the extraction ground truth.
References
[1] U.S. Bureau of Labor Statistics, Employed persons by detailed occupation and age, 2025. Available at: https://www.bls.gov/cps/cpsaat11b.htm
llama-3.1-8b-funding-extraction-sft-ablations
LLaMA 3.1 8B Funding Extraction SFT Ablations
Ablation study results for LoRA SFT of Meta LLaMA 3.1 8B Instruct on structured funding metadata extraction from scholarly text.
The model extracts four fields: funder_name, award_ids, funding_scheme, and award_title.
Key findings
Factor
Best config
Avg F1
Overall best
synthetic, twostage (2+1 epochs), LoRA r=64, lr=3e-5
0.588
Data type
Synthetic >> non-synthetic (+0.126 avg F1)
—
LoRA rank
r=64 > r=32 > r=16
—… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/llama-3.1-8b-funding-extraction-sft-ablations.paper-url-extraction-v1
Papers With Code URL Extraction
A representative dataset for training and evaluating tool-using agents that
find the official GitHub repository and project page for an AI research paper.
It was prepared for the pwc-url-extraction-v1 Prime/verifiers environment.
Splits
Split
Rows
train
4,000
validation
500
test
500
Rows were sampled with seed 13 from up to 120,000 Papers With Code candidates,
stratified by paper year and known URL state.… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/paper-url-extraction-v1.dotsocr-extractions-holdoutcodedp-bench-seq-extraction-cpt-p5020250817-funding-extraction-test
Overview
This test dataset contains extracted and structured funding acknowledgments from works in OpenAlex. It includes 4,746 records derived from 1,387 PDF-markdown-conversions, with 922 documents containing identifiable funding statements. The dataset captures funding organizations, grant identifiers, and the original acknowledgment text.
Structure
Fields
document_id (string): Unique identifier for the source document
funding_statement (string): Normalized… See the full description on the dataset page: https://huggingface.co/datasets/adambuttrick/20250817-funding-extraction-test.codedp-bench-pii-extraction-cptcodedp-bench-seq-extraction-cpt-p100
