datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tweet_sentiment_extraction
TweetSentimentExtractionClassification
An MTEB dataset
Massive Text Embedding Benchmark
Task category
t2c
Domains
Social, Written
Reference
https://www.kaggle.com/competitions/tweet-sentiment-extraction/overview
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["TweetSentimentExtractionClassification"])
evaluator = mteb.MTEB(task)
model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/tweet_sentiment_extraction.wikidata-extraction
Wikidata Extraction
This dataset contains all RDF triples extracted from the latest Wikidata, converted from the N-Triples format to Parquet.
The data originates from Wikidata, a free and open knowledge base that acts as central storage for structured data used by Wikipedia and other Wikimedia projects. The source file is the "truthy" N-Triples dump (latest-truthy.nt.bz2), which contains only the current, non-deprecated statements.
The code to extract this data is available at… See the full description on the dataset page: https://huggingface.co/datasets/piebro/wikidata-extraction.invoice-extraction-benchmark
Invoice Extraction Benchmark v1
A synthetic test set for invoice data extraction (invoice OCR, intelligent document
processing, accounts-payable capture): 181 documents with answer keys and a scorer.
Run any invoice reader over the documents, write its output as one JSON file, and score it
field by field.
Source, scorer and generator: https://github.com/DrewKraken/invoice-extraction-benchmark
(this dataset is a mirror of corpus/v1/ there; the GitHub repository is canonical).
Who… See the full description on the dataset page: https://huggingface.co/datasets/drew-ipp/invoice-extraction-benchmark.DRIVE-digital-retinal-images-for-vessel-extractionarXiv:2501.18921https://arxiv.org/abs/2501.18921
promoter-component-extraction-dataglobal-openstreetmap-extraction-slippy-tiles-tarweb-content-extraction-benchmark
WCXB: Web Content Extraction Benchmark
The largest open benchmark for evaluating web content extraction, boilerplate removal, and main content detection across diverse page types.
WCXB provides 2,008 human-reviewed web pages spanning 7 page types and 1,613 domains, with ground truth annotations, HTML source files, and baseline results from 14 extraction systems. Unlike existing benchmarks that focus exclusively on news articles, WCXB evaluates extraction across the full diversity of… See the full description on the dataset page: https://huggingface.co/datasets/murrough-foley/web-content-extraction-benchmark.skill-extraction-tech
Skill Extraction with ESCO skills - TECH subset
Dataset Summary
This dataset contains an extension of the TECH subset form the SkillSpan dataset, in which spans of skill mentions in sentences have been labeled with corresponding ESCO skills (ESCO v1.1.0).
This dataset is part of a three-part evaluation dataset for skill extraction:
skill-extraction-tech
skill-extraction-house
skill-extraction-techwolf
Citation Information
If you use this dataset, please include… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/skill-extraction-tech.affirming-review-extractionsupercon-extraction-harbor-taskstool-output-extraction-swebench
Tool Output Extraction Dataset
Paper | Code
Training data for squeez — a small model that prunes verbose coding agent tool output to only the evidence the agent needs next.
Task
Task-conditioned context pruning of a single tool observation for coding agents.
Given a focused extraction query and one verbose tool output, return the smallest verbatim evidence block(s) the agent should read next.
The model copies lines from the tool output — it never rewrites, summarizes, or… See the full description on the dataset page: https://huggingface.co/datasets/KRLabsOrg/tool-output-extraction-swebench.skill-extraction-house
Skill Extraction with ESCO skills - HOUSE subset
Dataset Summary
This dataset contains an extension of the HOUSE subset form the SkillSpan dataset, in which spans of skill mentions in sentences have been labeled with corresponding ESCO skills (ESCO v1.1.0).
This dataset is part of a three-part evaluation dataset for skill extraction:
skill-extraction-tech
skill-extraction-house
skill-extraction-techwolf
Citation Information
If you use this dataset, please… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/skill-extraction-house.price-tag-extraction
Price tag extraction dataset
This dataset contains images of price tags extracted from the Open Prices dataset, along with information extracted from this dataset.
It is intended to be used to train large visual language models (LVLMs) to extract information from price tag images, as part of the Open Prices project.
For more information about the format of this dataset, please refer to the documentation on llm-image-extraction datasets.
Dataset creation
A detailed… See the full description on the dataset page: https://huggingface.co/datasets/openfoodfacts/price-tag-extraction.scrapinghub-article-extraction-benchmark
Scrapinghub Article Extraction Benchmark
This dataset was originally created and distributed under MIT License by Scrapinghub on GitHub: github.com/scrapinghub/article-extraction-benchmark
It is mirrored on the HuggingFace Hub as a convenience.
cdw-extraction-harbor-taskstweet_sentiment_extraction
Tweet Sentiment Extraction
Source: https://www.kaggle.com/c/tweet-sentiment-extraction/data
mall_receipt_extraction_datasetraw-fact-extractionbiosurfactants-extraction-harbor-tasksImage_based_trait_extraction
Dataset Card
Dataset Summary
This dataset contains side-view RGB images of Chenopodium quinoa plants exposed to drought and salinity stress, collected at the Netherlands Plant Eco-phenotyping Centre (NPEC). The repository also includes the scripts required to preprocess the raw images, train and apply a U-Net++ segmentation model, manually annotated images used for model training, and the trained segmentation model accompanying the associated publication.
The… See the full description on the dataset page: https://huggingface.co/datasets/MiekevanVlaardingen/Image_based_trait_extraction.table-extraction-scientific-datasets
Benchmarking Table Extraction from Heterogeneous Scientific PDF Documents
Dataset Sources
Repository: GitHub
Code archive: Zenodo (DOI 10.5281/zenodo.20486311)
Paper: 10.1145/3770855.3817462
Extended version: arXiv:2511.16134
Interactive demo: table-extraction-benchmark-explorer.streamlit.app
(source)
Dataset Details
PubTables (pubtables/*.tar.gz)
Subset of PubTables-Test dataset, enriched with HTML table ground… See the full description on the dataset page: https://huggingface.co/datasets/marijanic/table-extraction-scientific-datasets.funding-extraction-harness-benchmarkllm-ner-extraction
Introduction
This dataset is an extraction of NER data from the wikipedia dataset.
This can be used to fine tune llm models for NER extraction.
openalex_extractiontooth_extraction_4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 200,
"total_frames": 76053,
"total_tasks": 1,
"total_videos": 400,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/samsam0510/tooth_extraction_4.turkish-keyword-extraction-500k
Turkish Keyword Extraction 500K v2
Yirmi alanda konu ve anahtar sözcük çıkarımı için kısa Türkçe belgeler.
Doğrulanmış boyut
Train: 490,000
Validation: 5,000
Test: 5,000
Toplam: 500,000
Ana görev sütunları: id, text, keywords, domain
Provenance
Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı
depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type,
provenance, generator_version… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-keyword-extraction-500k.calcium-source-extraction
Calcium source extraction
One-photon (miniscope-style) calcium imaging for source extraction: find every cell in a movie and
recover its fluorescence trace. The movies include correlated neuropil, vasculature and
hemodynamics, non-stationary rigid motion, wide-field optics and camera noise, and come with
complete ground truth.
data/heldout/video.tif 60 s at 15 Hz, 512 x 512, 8-bit, 900 frames: the movie to extract from
data/heldout/acquisition.txt frame rate, pixel size… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/calcium-source-extraction.Skill-extraction-SkillSkape-graded
skill-extraction-skillskape-graded
Graded-relevance annotations for sentences from
jjzha/skillskape
against the ESCO v1.1.0 skill taxonomy. Layout follows the
BEIR convention so it is drop-in for
MTEB-style retrieval evaluators.
This dataset was created for the RecSys-HR 2026 WorkRB challenge.
Configs
config
split
rows
columns
queries
validation
100
_id (sentence id), text (sentence)
queries
test
500
_id (sentence id), text (sentence)
corpus
corpus… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Skill-extraction-SkillSkape-graded.skill-extraction-techwolf
Skill Extraction with ESCO skills - TechWolf subset
Dataset Summary
The TECHWOLF subset, although smaller, represents a more generic distribution of job descriptions and skill spans. ESCO skills are directly annotated on the full sentence level, thus omitting the intermediate span identification step. ESCO v1.1.0 is used.
This dataset is part of a three-part evaluation dataset for skill extraction:
skill-extraction-tech
skill-extraction-house
skill-extraction-techwolf… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/skill-extraction-techwolf.sec-contracts-financial-extraction-instructions
S&P 500 SEC Financial Extraction Instructions
Dataset Summary
7,683 instruction-tuning examples for training LLMs to extract structured financial data from SEC filings. Covers two filing types across S&P 500 companies:
Split
Examples
Filing Type
Description
train
3,430
Exhibit 10 + DEF 14A
Positive examples with validated outputs
corrective
4,253
Exhibit 10 + DEF 14A
Corrective, rescued, and negative examples
Exhibit 10 — Material Contracts (2… See the full description on the dataset page: https://huggingface.co/datasets/TheTokenFactory/sec-contracts-financial-extraction-instructions.
