Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01piebro /wikidata-extraction Wikidata Extraction This dataset contains all RDF triples extracted from the latest Wikidata, converted from the N-Triples format to Parquet. The data originates from Wikidata, a free and open knowledge base that acts as central storage for structured data used by Wikipedia and other Wikimedia projects. The source file is the "truthy" N-Triples dump (latest-truthy.nt.bz2), which contains only the current, non-deprecated statements. The code to extract this data is available at… See the full description on the dataset page: https://huggingface.co/datasets/piebro/wikidata-extraction.tabular1B<n<10B4 likes5.7k downloads9mo agoHugging Face02drew-ipp /invoice-extraction-benchmark Invoice Extraction Benchmark v1 A synthetic test set for invoice data extraction (invoice OCR, intelligent document processing, accounts-payable capture): 181 documents with answer keys and a scorer. Run any invoice reader over the documents, write its output as one JSON file, and score it field by field. Source, scorer and generator: https://github.com/DrewKraken/invoice-extraction-benchmark (this dataset is a mirror of corpus/v1/ there; the GitHub repository is canonical). Who… See the full description on the dataset page: https://huggingface.co/datasets/drew-ipp/invoice-extraction-benchmark.documentimage-to-textn<1K1 likes2.2k downloads6d agoHugging Face03obalcells /raw-fact-extractiontabular100K<n<1M0 likes331 downloads2y agoHugging Face04cometadata /funding-extraction-harness-benchmarktabular10K<n<100K0 likes285 downloads7mo agoHugging Face05samsam0510 /tooth_extraction_4This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "so100", "total_episodes": 200, "total_frames": 76053, "total_tasks": 1, "total_videos": 400, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:200" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/samsam0510/tooth_extraction_4.tabularrobotics10K<n<100K0 likes248 downloads2y agoHugging Face06FuzzyLabs /ntsb-accident-extraction Files data/: the train, val, and test splits used for finetuning and inference raw.jsonl: every cleaned record with its narrative, reference fields, and split, used as evaluation ground truth preparation_config.yaml: the exact cleaning, splitting, and prompt configuration used vocab_schema.json: the closed vocabulary for every categorical field, as a JSON Schema Processing Combined source splits: train Document length: 300 to 30,000 characters Deduplicated by… See the full description on the dataset page: https://huggingface.co/datasets/FuzzyLabs/ntsb-accident-extraction.tabulartext-generation1K<n<10K0 likes177 downloads15d agoHugging Face07rafmacalaba /fcv-extractions-meta fcv-extractions-meta Data-use mention extractions from the World Bank Fragility, Conflict and Violence (FCV) document corpus, enriched by the fine-tuned rafmacalaba/lfm2.5-350M-datause-multitask model (a LoRA SFT of LiquidAI/LFM2.5-350M). Shape nested — one row per chunk; each entity in entities[] carries provenance and usage_impact alongside its span. Configs config rows fcv_pads_east_africa 793,763 jdc_operational 12,372… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/fcv-extractions-meta.tabular100K<n<1M0 likes159 downloads2mo agoHugging Face08samsam0510 /tooth_extraction_3This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "so100", "total_episodes": 100, "total_frames": 32879, "total_tasks": 1, "total_videos": 200, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:100" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/samsam0510/tooth_extraction_3.tabularrobotics10K<n<100K0 likes138 downloads2y agoHugging Face09cometadata /arxiv-funding-entity-extractions arxiv-funding-entity-extractions Funder/award entity extractions over cometadata/arxiv-funding-statements. Extractor: funding-entity-extractor (vLLM + LoRA) Base model: meta-llama/Llama-3.1-8B-Instruct LoRA: cometadata/funding-extraction-llama-3.1-8b-instruct-artifact-data-mix-grpo-mixed-reward Hardware: A100-large bf16, concurrency 256 Total rows: 1,823,650 Configs predictions (default) — original extractions, no ROR enrichment. predictions_with_ror — same rows… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-funding-entity-extractions.tabular1M<n<10M1 likes132 downloads5mo agoHugging Face10ClarusC64 /autonomous-driving-intention-field-extraction-v0.1What this dataset tests Whether a system can infer agent intentions from context cues in complex driving scenes. This is not trajectory prediction. It is intention inference. Required outputs agent_id inferred_intention intention_confidence time_horizon_s alternative_intentions stability_score Scoring conventions confidence and stability range 0 to 1 time horizon is seconds into the near future Use case Layer one of Intention Field and Social Coherence Maps. This enables… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/autonomous-driving-intention-field-extraction-v0.1.tabulartabular-classificationn<1K0 likes105 downloads8mo agoHugging Face11necrasov-ilya /ru-invoice-extraction-benchmark Набор для извлечения данных из русскоязычных счетов 50 синтетических русскоязычных счетов с эталонными ответами в JSON. Набор не привязан к определённой модели и предназначен для проверки систем извлечения структурированных данных из текста. Разделы Раздел Документы Назначение development 30 разработка шаблонов и примеров validation 10 выбор настроек test 10 итоговая оценка Раздел test зафиксирован для версии 1. Его нельзя использовать для… See the full description on the dataset page: https://huggingface.co/datasets/necrasov-ilya/ru-invoice-extraction-benchmark.tabulartext-generationn<1K0 likes99 downloads28d agoHugging Face12zaaabik /paper_extractiontabular1M<n<10M0 likes94 downloads17d agoHugging Face13alexdzm /extraction-examples Extraction Examples Dataset This dataset contains 17 examples for testing extraction workflows. Dataset Structure Each example includes: PDF file: Original document map_info.json: Map extraction metadata direction.json: Direction information GeoJSON files: Polygon geometries Area JSON files: Area definitions File Organization files/ ├── example1/ │ ├── document.pdf │ ├── map_info.json │ ├── direction.json │ ├── polygon1.geojson │ └── area1.json… See the full description on the dataset page: https://huggingface.co/datasets/alexdzm/extraction-examples.documentothern<1K0 likes90 downloads1y agoHugging Face14open-source-metrics /feature-extraction-checkpoint-downloadstabular1K<n<10K4 likes85 downloads4y agoHugging Face15JobOpportunitiesAPI /joa-extraction-arena JOA Job-Posting Extraction Arena Disclosure: Job Opportunities API (JOA, jobopportunitiesapi.org) is an independent data business that sells API access to job-posting data. AI helped run the experiments, check the numbers and draft this text; Loukas (Luca) Tzekos is editorially responsible. Contact: hello@jobopportunitiesapi.org. A benchmark for one narrow task: reading a real job posting and filling 11 structured fields. Job Opportunities API (JOA) built it to choose and train… See the full description on the dataset page: https://huggingface.co/datasets/JobOpportunitiesAPI/joa-extraction-arena.tabulartext-generationn<1K0 likes81 downloads2d agoHugging Face16zaaabik /paper_extraction_v1tabular100K<n<1M0 likes77 downloads2mo agoHugging Face17samsam0510 /eval_act_tooth_extraction_3This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "so100", "total_episodes": 10, "total_frames": 17799, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:10" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/samsam0510/eval_act_tooth_extraction_3.tabularrobotics10K<n<100K0 likes71 downloads2y agoHugging Face18lihaoxin2020 /ki_extraction-4b-lmeval-qwen2.5-7b-instruct-on-mmlu_pro-0shot_cot-scillm-5d1468e6e9tabular1K<n<10K0 likes70 downloads1y agoHugging Face19SR219 /Dicom-metadata-extraction-skillbenchtabularn<1K0 likes63 downloads9mo agoHugging Face20gabrielmbmb /finemath-qa-extraction-test Dataset Card for finemath-qa-extraction-test This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/gabrielmbmb/finemath-qa-extraction-test/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/gabrielmbmb/finemath-qa-extraction-test.tabular10K<n<100K1 likes50 downloads2y agoHugging Face21Haeryz /putusan-structured-extraction Putusan structured-extraction dataset Built 2026-07-08T23:02:28+00:00 by notebooks/build_dataset.py (seed 3407). Indonesian court-decision (putusan) extractive-structuring dataset over three corpora (Anak, Asusila, TPPO). Each row is one model extraction of one source document into 31 canonical sections of verbatim spans. Empty sections were completed from sibling model extractions of the same document where available (cross_model_fill_json records per-section donor provenance).… See the full description on the dataset page: https://huggingface.co/datasets/Haeryz/putusan-structured-extraction.tabulartext-generation1K<n<10K0 likes44 downloads2mo agoHugging Face22rafmacalaba /fcv-extractions-meta-tiered-probe fcv-extractions-meta-tiered-probe Data-use mention extractions from the World Bank Fragility, Conflict and Violence (FCV) document corpus. Spans are extracted by the fine-tuned GLiNER model rafmacalaba/gliner_datause_tiered, scored by the tier-probe head rafmacalaba/gliner-tier-probe, and attributed (provenance + usage/impact) by rafmacalaba/lfm2.5-350M-datause-multitask-tiered (a LoRA SFT of LiquidAI/LFM2.5-350M). Shape nested — one row per chunk; each entity in… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/fcv-extractions-meta-tiered-probe.tabular10K<n<100K0 likes44 downloads1mo agoHugging Face23rafmacalaba /fcv-extractions-meta-tiered fcv-extractions-meta-tiered Data-use mention extractions from the World Bank Fragility, Conflict and Violence (FCV) document corpus, enriched by the fine-tuned rafmacalaba/lfm2.5-350M-datause-multitask model (a LoRA SFT of LiquidAI/LFM2.5-350M). Shape nested — one row per chunk; each entity in entities[] carries provenance and usage_impact alongside its span. Configs config rows fcv_pads_east_africa 862,663 jdc_operational 2,468… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/fcv-extractions-meta-tiered.tabular10K<n<100K0 likes43 downloads1mo agoHugging Face24avsolatorio /mteb-tweet_sentiment_extraction-avs_triplets MTEB Tweet Sentiment Extraction Triplets Dataset This dataset was used in the paper GISTEmbed: Guided In-sample Selection of Training Negatives for Text Embedding Fine-tuning. Refer to https://arxiv.org/abs/2402.16829 for details. The code for generating the data is available at https://github.com/avsolatorio/GISTEmbed/blob/main/scripts/create_classification_dataset.py. Citation @article{solatorio2024gistembed, title={GISTEmbed: Guided In-sample Selection of… See the full description on the dataset page: https://huggingface.co/datasets/avsolatorio/mteb-tweet_sentiment_extraction-avs_triplets.tabular10K<n<100K0 likes37 downloads3y agoHugging Face25spectralbranding /unitization-before-extraction Unitization Before Extraction — the 2026bl run Records, adjudicated inventories, per-call metadata and derived tables for a pre-registered diagnostic replication that decomposes disagreement in machine recovery of document-scale argument dependency structure. The design was deposited before any datum existed (10.5281/zenodo.21830221, 2026-08-07); the run executed 2026-08-08. Every threshold in the analysis preceded the data, and that ordering is a checkable timestamp rather than… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/unitization-before-extraction.tabularn<1K0 likes36 downloads2mo agoHugging Face26Haeryz /putusan-windowed-extraction Putusan windowed line-anchored extraction dataset (Plan B) Built 2026-07-09T13:01:27+00:00 by notebooks/build_windowed_dataset.py from the legacy Haeryz/putusan-structured-extraction dataset (same documents, same leakage-safe purpose/split assignment, seed 3407). Each legacy document row (~34K tokens median — longer than a 32K context) is re-expressed as overlapping line-numbered windows of <= 6400 content tokens (measured with Qwen/Qwen3.5-9B; fits a max_seq_length of 8192 with… See the full description on the dataset page: https://huggingface.co/datasets/Haeryz/putusan-windowed-extraction.tabulartext-generation10K<n<100K0 likes34 downloads3mo agoHugging Face27cometadata /arxiv-author-affiliation-extraction-inference-inputs-metadatatabular100K<n<1M0 likes33 downloads11mo agoHugging Face28math-extraction-comp /deepseek-ai__deepseek-llm-7b-chattabular1K<n<10K0 likes32 downloads2y agoHugging Face29lbakar /health-log-extraction-datasetTo create this dataset, we sampled demographic seeds from an occupation and age-range table from the Labor Force Statistics[1]. For each occupation, the pipeline randomly selected an age range, sampled an age within that range, and assigned a gender from a fixed set of options. These demographic seeds were used to prompt an LLM to generate structured personas containing a name, description, medications or supplements, general mood, and possible health conditions or injuries. We then used… See the full description on the dataset page: https://huggingface.co/datasets/lbakar/health-log-extraction-dataset.tabularfeature-extraction100K<n<1M0 likes32 downloads1mo agoHugging Face30math-extraction-comp /Goekdeniz-Guelmez__Josiefied-Qwen2.5-1.5B-Instruct-abliterated-v3tabular1K<n<10K0 likes31 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.