Team Ai
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01drew-ipp /invoice-extraction-benchmark Invoice Extraction Benchmark v1 A synthetic test set for invoice data extraction (invoice OCR, intelligent document processing, accounts-payable capture): 181 documents with answer keys and a scorer. Run any invoice reader over the documents, write its output as one JSON file, and score it field by field. Source, scorer and generator: https://github.com/DrewKraken/invoice-extraction-benchmark (this dataset is a mirror of corpus/v1/ there; the GitHub repository is canonical). Who… See the full description on the dataset page: https://huggingface.co/datasets/drew-ipp/invoice-extraction-benchmark.documentimage-to-textn<1K1 likes2.2k downloads6d agoHugging Face02rafmacalaba /fcv-extractions-meta fcv-extractions-meta Data-use mention extractions from the World Bank Fragility, Conflict and Violence (FCV) document corpus, enriched by the fine-tuned rafmacalaba/lfm2.5-350M-datause-multitask model (a LoRA SFT of LiquidAI/LFM2.5-350M). Shape nested — one row per chunk; each entity in entities[] carries provenance and usage_impact alongside its span. Configs config rows fcv_pads_east_africa 793,763 jdc_operational 12,372… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/fcv-extractions-meta.tabular100K<n<1M0 likes159 downloads2mo agoHugging Face03necrasov-ilya /ru-invoice-extraction-benchmark Набор для извлечения данных из русскоязычных счетов 50 синтетических русскоязычных счетов с эталонными ответами в JSON. Набор не привязан к определённой модели и предназначен для проверки систем извлечения структурированных данных из текста. Разделы Раздел Документы Назначение development 30 разработка шаблонов и примеров validation 10 выбор настроек test 10 итоговая оценка Раздел test зафиксирован для версии 1. Его нельзя использовать для… See the full description on the dataset page: https://huggingface.co/datasets/necrasov-ilya/ru-invoice-extraction-benchmark.tabulartext-generationn<1K0 likes99 downloads29d agoHugging Face04SR219 /Dicom-metadata-extraction-skillbenchtabularn<1K0 likes63 downloads9mo agoHugging Face05rafmacalaba /fcv-extractions-meta-tiered-probe fcv-extractions-meta-tiered-probe Data-use mention extractions from the World Bank Fragility, Conflict and Violence (FCV) document corpus. Spans are extracted by the fine-tuned GLiNER model rafmacalaba/gliner_datause_tiered, scored by the tier-probe head rafmacalaba/gliner-tier-probe, and attributed (provenance + usage/impact) by rafmacalaba/lfm2.5-350M-datause-multitask-tiered (a LoRA SFT of LiquidAI/LFM2.5-350M). Shape nested — one row per chunk; each entity in… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/fcv-extractions-meta-tiered-probe.tabular10K<n<100K0 likes44 downloads1mo agoHugging Face06rafmacalaba /fcv-extractions-meta-tiered fcv-extractions-meta-tiered Data-use mention extractions from the World Bank Fragility, Conflict and Violence (FCV) document corpus, enriched by the fine-tuned rafmacalaba/lfm2.5-350M-datause-multitask model (a LoRA SFT of LiquidAI/LFM2.5-350M). Shape nested — one row per chunk; each entity in entities[] carries provenance and usage_impact alongside its span. Configs config rows fcv_pads_east_africa 862,663 jdc_operational 2,468… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/fcv-extractions-meta-tiered.tabular10K<n<100K0 likes43 downloads1mo agoHugging Face07spectralbranding /unitization-before-extraction Unitization Before Extraction — the 2026bl run Records, adjudicated inventories, per-call metadata and derived tables for a pre-registered diagnostic replication that decomposes disagreement in machine recovery of document-scale argument dependency structure. The design was deposited before any datum existed (10.5281/zenodo.21830221, 2026-08-07); the run executed 2026-08-08. Every threshold in the analysis preceded the data, and that ordering is a checkable timestamp rather than… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/unitization-before-extraction.tabularn<1K0 likes36 downloads2mo agoHugging Face08cometadata /arxiv-author-affiliation-extraction-inference-inputs-metadatatabular100K<n<1M0 likes33 downloads11mo agoHugging Face09lbakar /health-log-extraction-datasetTo create this dataset, we sampled demographic seeds from an occupation and age-range table from the Labor Force Statistics[1]. For each occupation, the pipeline randomly selected an age range, sampled an age within that range, and assigned a gender from a fixed set of options. These demographic seeds were used to prompt an LLM to generate structured personas containing a name, description, medications or supplements, general mood, and possible health conditions or injuries. We then used… See the full description on the dataset page: https://huggingface.co/datasets/lbakar/health-log-extraction-dataset.tabularfeature-extraction100K<n<1M0 likes32 downloads1mo agoHugging Face10lbakar /real-human-logs-extraction-datasetTo create this dataset, we collected human-generated logs from two individuals. This needs to be beefed up in the future, but this is what we have for now. Subsequently, we ran NuExtract3 on each example of the dataset to get the extraction ground truth. References [1] U.S. Bureau of Labor Statistics, Employed persons by detailed occupation and age, 2025. Available at: https://www.bls.gov/cps/cpsaat11b.htm tabularn<1K0 likes25 downloads1mo agoHugging Face11cometadata /llama-3.1-8b-funding-extraction-sft-ablations LLaMA 3.1 8B Funding Extraction SFT Ablations Ablation study results for LoRA SFT of Meta LLaMA 3.1 8B Instruct on structured funding metadata extraction from scholarly text. The model extracts four fields: funder_name, award_ids, funding_scheme, and award_title. Key findings Factor Best config Avg F1 Overall best synthetic, twostage (2+1 epochs), LoRA r=64, lr=3e-5 0.588 Data type Synthetic >> non-synthetic (+0.126 avg F1) — LoRA rank r=64 > r=32 > r=16 —… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/llama-3.1-8b-funding-extraction-sft-ablations.tabulartoken-classificationn<1K0 likes24 downloads7mo agoHugging Face12nielsr /paper-url-extraction-v1 Papers With Code URL Extraction A representative dataset for training and evaluating tool-using agents that find the official GitHub repository and project page for an AI research paper. It was prepared for the pwc-url-extraction-v1 Prime/verifiers environment. Splits Split Rows train 4,000 validation 500 test 500 Rows were sampled with seed 13 from up to 120,000 Papers With Code candidates, stratified by paper year and known URL state.… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/paper-url-extraction-v1.tabulartext-generation1K<n<10K0 likes19 downloads2mo agoHugging Face13bluecopa /dotsocr-extractions-holdouttabularn<1K0 likes15 downloads9mo agoHugging Face14melihcatal /codedp-bench-seq-extraction-cpt-p50tabularn<1K0 likes6 downloads7mo agoHugging Face15adambuttrick /20250817-funding-extraction-test Overview This test dataset contains extracted and structured funding acknowledgments from works in OpenAlex. It includes 4,746 records derived from 1,387 PDF-markdown-conversions, with 922 documents containing identifiable funding statements. The dataset captures funding organizations, grant identifiers, and the original acknowledgment text. Structure Fields document_id (string): Unique identifier for the source document funding_statement (string): Normalized… See the full description on the dataset page: https://huggingface.co/datasets/adambuttrick/20250817-funding-extraction-test.tabular1K<n<10K0 likes5 downloads1y agoHugging Face16melihcatal /codedp-bench-pii-extraction-cpttabularn<1K0 likes5 downloads7mo agoHugging Face17melihcatal /codedp-bench-seq-extraction-cpt-p100tabularn<1K0 likes3 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.