Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /tweet_sentiment_extraction TweetSentimentExtractionClassification An MTEB dataset Massive Text Embedding Benchmark Task category t2c Domains Social, Written Reference https://www.kaggle.com/competitions/tweet-sentiment-extraction/overview How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["TweetSentimentExtractionClassification"]) evaluator = mteb.MTEB(task) model =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/tweet_sentiment_extraction.texttext-classification10K<n<100K38 likes7.6k downloads1y agoHugging Face02piebro /wikidata-extraction Wikidata Extraction This dataset contains all RDF triples extracted from the latest Wikidata, converted from the N-Triples format to Parquet. The data originates from Wikidata, a free and open knowledge base that acts as central storage for structured data used by Wikipedia and other Wikimedia projects. The source file is the "truthy" N-Triples dump (latest-truthy.nt.bz2), which contains only the current, non-deprecated statements. The code to extract this data is available at… See the full description on the dataset page: https://huggingface.co/datasets/piebro/wikidata-extraction.tabular1B<n<10B4 likes5.7k downloads9mo agoHugging Face03drew-ipp /invoice-extraction-benchmark Invoice Extraction Benchmark v1 A synthetic test set for invoice data extraction (invoice OCR, intelligent document processing, accounts-payable capture): 181 documents with answer keys and a scorer. Run any invoice reader over the documents, write its output as one JSON file, and score it field by field. Source, scorer and generator: https://github.com/DrewKraken/invoice-extraction-benchmark (this dataset is a mirror of corpus/v1/ there; the GitHub repository is canonical). Who… See the full description on the dataset page: https://huggingface.co/datasets/drew-ipp/invoice-extraction-benchmark.documentimage-to-textn<1K1 likes2.2k downloads6d agoHugging Face04Zomba /DRIVE-digital-retinal-images-for-vessel-extractionarXiv:2501.18921https://arxiv.org/abs/2501.18921 imageimage-segmentationn<1K0 likes1.6k downloads11mo agoHugging Face05jyw-zju /promoter-component-extraction-data0 likes1.1k downloads2mo agoHugging Face06PeakNav /global-openstreetmap-extraction-slippy-tiles-tar0 likes900 downloads26d agoHugging Face07murrough-foley /web-content-extraction-benchmark WCXB: Web Content Extraction Benchmark The largest open benchmark for evaluating web content extraction, boilerplate removal, and main content detection across diverse page types. WCXB provides 2,008 human-reviewed web pages spanning 7 page types and 1,613 domains, with ground truth annotations, HTML source files, and baseline results from 14 extraction systems. Unlike existing benchmarks that focus exclusively on news articles, WCXB evaluates extraction across the full diversity of… See the full description on the dataset page: https://huggingface.co/datasets/murrough-foley/web-content-extraction-benchmark.2 likes820 downloads6mo agoHugging Face08TechWolf /skill-extraction-tech Skill Extraction with ESCO skills - TECH subset Dataset Summary This dataset contains an extension of the TECH subset form the SkillSpan dataset, in which spans of skill mentions in sentences have been labeled with corresponding ESCO skills (ESCO v1.1.0). This dataset is part of a three-part evaluation dataset for skill extraction: skill-extraction-tech skill-extraction-house skill-extraction-techwolf Citation Information If you use this dataset, please include… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/skill-extraction-tech.texttext-classification1K<n<10K2 likes733 downloads2y agoHugging Face09yilanliu917 /affirming-review-extraction1 likes690 downloads8d agoHugging Face10kilian-group /supercon-extraction-harbor-tasks0 likes631 downloads9mo agoHugging Face11KRLabsOrg /tool-output-extraction-swebench Tool Output Extraction Dataset Paper | Code Training data for squeez — a small model that prunes verbose coding agent tool output to only the evidence the agent needs next. Task Task-conditioned context pruning of a single tool observation for coding agents. Given a focused extraction query and one verbose tool output, return the smallest verbatim evidence block(s) the agent should read next. The model copies lines from the tool output — it never rewrites, summarizes, or… See the full description on the dataset page: https://huggingface.co/datasets/KRLabsOrg/tool-output-extraction-swebench.texttext-generation10K<n<100K5 likes588 downloads6mo agoHugging Face12TechWolf /skill-extraction-house Skill Extraction with ESCO skills - HOUSE subset Dataset Summary This dataset contains an extension of the HOUSE subset form the SkillSpan dataset, in which spans of skill mentions in sentences have been labeled with corresponding ESCO skills (ESCO v1.1.0). This dataset is part of a three-part evaluation dataset for skill extraction: skill-extraction-tech skill-extraction-house skill-extraction-techwolf Citation Information If you use this dataset, please… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/skill-extraction-house.texttext-classification1K<n<10K5 likes576 downloads2y agoHugging Face13openfoodfacts /price-tag-extraction Price tag extraction dataset This dataset contains images of price tags extracted from the Open Prices dataset, along with information extracted from this dataset. It is intended to be used to train large visual language models (LVLMs) to extract information from price tag images, as part of the Open Prices project. For more information about the format of this dataset, please refer to the documentation on llm-image-extraction datasets. Dataset creation A detailed… See the full description on the dataset page: https://huggingface.co/datasets/openfoodfacts/price-tag-extraction.image10K<n<100K2 likes539 downloads9mo agoHugging Face14allenai /scrapinghub-article-extraction-benchmark Scrapinghub Article Extraction Benchmark This dataset was originally created and distributed under MIT License by Scrapinghub on GitHub: github.com/scrapinghub/article-extraction-benchmark It is mirrored on the HuggingFace Hub as a convenience. textn<1K0 likes535 downloads3y agoHugging Face15kilian-group /cdw-extraction-harbor-tasks0 likes415 downloads9mo agoHugging Face16SetFit /tweet_sentiment_extraction Tweet Sentiment Extraction Source: https://www.kaggle.com/c/tweet-sentiment-extraction/data text10K<n<100K11 likes391 downloads4y agoHugging Face17CC1984 /mall_receipt_extraction_datasetimage1K<n<10K3 likes349 downloads3y agoHugging Face18obalcells /raw-fact-extractiontabular100K<n<1M0 likes331 downloads2y agoHugging Face19kilian-group /biosurfactants-extraction-harbor-tasks0 likes326 downloads9mo agoHugging Face20MiekevanVlaardingen /Image_based_trait_extraction Dataset Card Dataset Summary This dataset contains side-view RGB images of Chenopodium quinoa plants exposed to drought and salinity stress, collected at the Netherlands Plant Eco-phenotyping Centre (NPEC). The repository also includes the scripts required to preprocess the raw images, train and apply a U-Net++ segmentation model, manually annotated images used for model training, and the trained segmentation model accompanying the associated publication. The… See the full description on the dataset page: https://huggingface.co/datasets/MiekevanVlaardingen/Image_based_trait_extraction.0 likes323 downloads3mo agoHugging Face21marijanic /table-extraction-scientific-datasets Benchmarking Table Extraction from Heterogeneous Scientific PDF Documents Dataset Sources Repository: GitHub Code archive: Zenodo (DOI 10.5281/zenodo.20486311) Paper: 10.1145/3770855.3817462 Extended version: arXiv:2511.16134 Interactive demo: table-extraction-benchmark-explorer.streamlit.app (source) Dataset Details PubTables (pubtables/*.tar.gz) Subset of PubTables-Test dataset, enriched with HTML table ground… See the full description on the dataset page: https://huggingface.co/datasets/marijanic/table-extraction-scientific-datasets.object-detection10K<n<100K1 likes295 downloads2mo agoHugging Face22cometadata /funding-extraction-harness-benchmarktabular10K<n<100K0 likes285 downloads7mo agoHugging Face23vangheem /llm-ner-extraction Introduction This dataset is an extraction of NER data from the wikipedia dataset. This can be used to fine tune llm models for NER extraction. text10K<n<100K0 likes259 downloads1y agoHugging Face24LLMDH /openalex_extractiontext10M<n<100M1 likes254 downloads2y agoHugging Face25samsam0510 /tooth_extraction_4This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "so100", "total_episodes": 200, "total_frames": 76053, "total_tasks": 1, "total_videos": 400, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:200" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/samsam0510/tooth_extraction_4.tabularrobotics10K<n<100K0 likes248 downloads2y agoHugging Face26GoktugD /turkish-keyword-extraction-500k Turkish Keyword Extraction 500K v2 Yirmi alanda konu ve anahtar sözcük çıkarımı için kısa Türkçe belgeler. Doğrulanmış boyut Train: 490,000 Validation: 5,000 Test: 5,000 Toplam: 500,000 Ana görev sütunları: id, text, keywords, domain Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance, generator_version… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-keyword-extraction-500k.texttoken-classification100K<n<1M0 likes241 downloads2mo agoHugging Face27Emulated-Inc /calcium-source-extraction Calcium source extraction One-photon (miniscope-style) calcium imaging for source extraction: find every cell in a movie and recover its fluorescence trace. The movies include correlated neuropil, vasculature and hemodynamics, non-stationary rigid motion, wide-field optics and camera noise, and come with complete ground truth. data/heldout/video.tif 60 s at 15 Hz, 512 x 512, 8-bit, 900 frames: the movie to extract from data/heldout/acquisition.txt frame rate, pixel size… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/calcium-source-extraction.imagen<1K0 likes239 downloads16d agoHugging Face28TechWolf /Skill-extraction-SkillSkape-graded skill-extraction-skillskape-graded Graded-relevance annotations for sentences from jjzha/skillskape against the ESCO v1.1.0 skill taxonomy. Layout follows the BEIR convention so it is drop-in for MTEB-style retrieval evaluators. This dataset was created for the RecSys-HR 2026 WorkRB challenge. Configs config split rows columns queries validation 100 _id (sentence id), text (sentence) queries test 500 _id (sentence id), text (sentence) corpus corpus… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/Skill-extraction-SkillSkape-graded.text1M<n<10M1 likes220 downloads2mo agoHugging Face29TechWolf /skill-extraction-techwolf Skill Extraction with ESCO skills - TechWolf subset Dataset Summary The TECHWOLF subset, although smaller, represents a more generic distribution of job descriptions and skill spans. ESCO skills are directly annotated on the full sentence level, thus omitting the intermediate span identification step. ESCO v1.1.0 is used. This dataset is part of a three-part evaluation dataset for skill extraction: skill-extraction-tech skill-extraction-house skill-extraction-techwolf… See the full description on the dataset page: https://huggingface.co/datasets/TechWolf/skill-extraction-techwolf.texttext-classificationn<1K3 likes218 downloads2y agoHugging Face30TheTokenFactory /sec-contracts-financial-extraction-instructions S&P 500 SEC Financial Extraction Instructions Dataset Summary 7,683 instruction-tuning examples for training LLMs to extract structured financial data from SEC filings. Covers two filing types across S&P 500 companies: Split Examples Filing Type Description train 3,430 Exhibit 10 + DEF 14A Positive examples with validated outputs corrective 4,253 Exhibit 10 + DEF 14A Corrective, rescued, and negative examples Exhibit 10 — Material Contracts (2… See the full description on the dataset page: https://huggingface.co/datasets/TheTokenFactory/sec-contracts-financial-extraction-instructions.texttext-generation10K<n<100K1 likes207 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.