Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Logistic12 /xhs-image-extractor-20260314115142image1K<n<10K0 likes130 downloads7mo agoHugging Face02emgena /omnimcp_browser_dom_structured_extractor_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_browser_dom_structured_extractor_teaser.texttext-generationn<1K0 likes90 downloads23d agoHugging Face03emgena /omnimcp_graphrag_triplet_extractor_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_graphrag_triplet_extractor_teaser.texttext-generationn<1K0 likes76 downloads23d agoHugging Face04Logistic12 /xhs-image-extractor-v2image1K<n<10K0 likes69 downloads7mo agoHugging Face05emgena /omnimcp_episodic_fact_extractor_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_episodic_fact_extractor_teaser.texttext-generationn<1K0 likes69 downloads23d agoHugging Face06emgena /multimodal_rag_complex_table_extractor_teaser 🚀 Data Platform - Multi-Modal RAG, Complex Document & Table Extractor (Evaluation Teaser) ⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (500 Samples) & Commercial EULA on Gumroad:👉 Data Platform - Multi-Modal RAG, Complex Document & Table Extractor on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout! 📦 What is Inside the Full Production Package: 500 Verified FAANG v2.0 Scenarios (100%… See the full description on the dataset page: https://huggingface.co/datasets/emgena/multimodal_rag_complex_table_extractor_teaser.texttext-generationn<1K0 likes67 downloads9d agoHugging Face07keerthanshetty /resume-skill-extractor-dataset Resume Skill Extractor Dataset Dataset Summary This dataset contains 3,050 pre-processed job descriptions with their summaries and required technical skills. It is designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs) to teach them how to parse job postings and extract skill requirements. Data Structure Each row in the dataset is a JSON object containing the following fields: title: The job title (e.g., "Senior Data Scientist"). source:… See the full description on the dataset page: https://huggingface.co/datasets/keerthanshetty/resume-skill-extractor-dataset.text1K<n<10K0 likes35 downloads6mo agoHugging Face08MBMMurad /Bangla_Person_Name_Extractortext1K<n<10K0 likes28 downloads3y agoHugging Face09dedemerve /ILSA-LLM-Extractor-Dataset ILSA LLM Extractor Dataset Project website: https://dedemerve.github.io/ILSA-LLM-Extractor/ Dataset Description This dataset contains structured metadata automatically extracted from 1,756 peer-reviewed articles and reports covering International Large-Scale Assessments (IEA: TIMSS, PIRLS, ICCS; OECD: PISA, TALIS, PIAAC). The extraction pipeline combines PDF parsing, LLM-based structured extraction, and RAG-based synthesis. Pipeline stages: Stage 1: LLM-based… See the full description on the dataset page: https://huggingface.co/datasets/dedemerve/ILSA-LLM-Extractor-Dataset.tabularfeature-extraction10K<n<100K0 likes24 downloads4mo agoHugging Face10nevvton /size-extractor-dataset-qatext1K<n<10K0 likes20 downloads2y agoHugging Face11smolify /smolified-extractor 🤏 smolified-extractor Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model smolify/smolified-extractor. 📦 Asset Details Origin: Smolify Foundry (Job ID: 13d088d3) Records: 15669 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by smolify. Generated via Smolify.ai. texttext-generation10K<n<100K0 likes18 downloads8mo agoHugging Face12v-rusu /recipe-extractor-datasetThis dataset was created to finetune a small gemma-3 270M model to extract valid, correct JSON-LD objects from a recipe blog/social media post. It's a completely synthethic dataset. I used Deepseek v3.2 to create the blog posts, JSON-LD extraction, and the reasoning traces. The blogs were created based on this Kaggle All Recipes Dataset. The pipeline to generate this dataset is available on github text1K<n<10K0 likes18 downloads8mo agoHugging Face13logiover /json-ld-schema-meta-tag-extractor-sample-data JSON-LD Schema & Meta Tag Extractor Extract JSON-LD/Schema.org structured data, Meta tags, OpenGraph and Twitter Cards from any URL. Get page title + meta description with a clean JSON output for SEO audits, validation, competitor research and AI datasets. Proxy-ready for large crawls. What the actor scrapes 🧩 JSON-LD Schema & Meta Tag Extractor — Scrape Schema.org, OpenGraph & Meta Tags Extract structured data and SEO metadata from any webpage in seconds. This… See the full description on the dataset page: https://huggingface.co/datasets/logiover/json-ld-schema-meta-tag-extractor-sample-data.textn<1K0 likes17 downloads5mo agoHugging Face14ciaranmacseoin /research_paper_extractortextn<1K3 likes16 downloads2y agoHugging Face15jaeyong2 /keywords-extractor-Kotext10K<n<100K0 likes16 downloads1y agoHugging Face16MinaGabriel /sentence-relevance-extractor Sentence Relevance Extractor (SRE) Sentence Relevance Extractor (SRE) is a large-scale dataset for binary evidence selection in multi-document, multi-hop question answering. The goal: Given a question and a sentence from the context, predict whether this sentence is relevant evidence ("Yes") or irrelevant ("No"). This dataset is suitable for training: Sentence-level RAG rerankers Binary relevance classifiers Optimization-based truth discovery systems Multi-hop QA evidence… See the full description on the dataset page: https://huggingface.co/datasets/MinaGabriel/sentence-relevance-extractor.tabular1M<n<10M0 likes16 downloads11mo agoHugging Face17JohnGorri /macro-extractor-flan-t5-synthtext10K<n<100K0 likes16 downloads3mo agoHugging Face18rishiraj /smolified-ingredient-extractor 🤏 smolified-ingredient-extractor Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model rishiraj/smolified-ingredient-extractor. 📦 Asset Details Origin: Smolify Foundry (Job ID: 65517eae) Records: 9905 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by rishiraj. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes15 downloads7mo agoHugging Face19dhareesh28 /resume-skill-extractor-dataset Resume Skill Extractor Dataset Dataset Summary This dataset contains 3,050 pre-processed job descriptions with their summaries and required technical skills. It is designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs) to teach them how to parse job postings and extract skill requirements. Data Structure Each row in the dataset is a JSON object containing the following fields: title: The job title (e.g., "Senior Data… See the full description on the dataset page: https://huggingface.co/datasets/dhareesh28/resume-skill-extractor-dataset.text1K<n<10K0 likes15 downloads2mo agoHugging Face20titou4ng /smolified-ocr-data-extractor-urssaf 🤏 smolified-ocr-data-extractor-urssaf Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model titou4ng/smolified-ocr-data-extractor-urssaf. 📦 Asset Details Origin: Smolify Foundry (Job ID: 6baf72cd) Records: 1288 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by titou4ng. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes14 downloads8mo agoHugging Face21MindCastSogang /word_extractorimage100K<n<1M0 likes13 downloads2mo agoHugging Face22Vishal24 /feature_extractortext1K<n<10K0 likes10 downloads2y agoHugging Face23reasoning-degeneration-dev /test-updated-extractor-v2 test-updated-extractor-v2 LLM-based math span extraction with canonicalization Dataset Info Rows: 1 Columns: 26 Columns Column Type Description question Value('string') No description provided metadata Value('string') No description provided task_source Value('string') No description provided formatted_prompt List({'content': Value('string'), 'role': Value('string')}) No description provided responses_by_sample List(List(Value('string')))… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/test-updated-extractor-v2.textn<1K0 likes9 downloads10mo agoHugging Face24smolify /smolified-ingredient-extractor 🤏 smolified-ingredient-extractor Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model smolify/smolified-ingredient-extractor. 📦 Asset Details Origin: Smolify Foundry (Job ID: 1f92fa68) Records: 600 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by smolify. Generated via Smolify.ai. texttext-generationn<1K0 likes9 downloads7mo agoHugging Face25shuaih777 /music-crs-state-extractor-datatext100K<n<1M0 likes9 downloads4mo agoHugging Face26jacekduszenko /lora-adapters-are-good-feature-extractors LORA Adapters are Good Feature Extractors Dataset This dataset contains images of two sets of categories that are not safe for work (hentai and porn, labelled as 0 and 2 correspondingly) and one neutral category, labelled as 2. The dataset is the source data for training a zoo of LORA adapters on sample images from each category. Adapters representations will then be used as input data to a weight-space model in an experiment to verify whether WS models operating in low rank… See the full description on the dataset page: https://huggingface.co/datasets/jacekduszenko/lora-adapters-are-good-feature-extractors.tabularn<1K1 likes8 downloads2y agoHugging Face27aravind-selvam /model_card_extractor_qwq Dataset card for model_card_extractor_qwq This dataset was made with Curator. Dataset details A sample from the dataset: { "modelId": "digiplay/XtReMixAnimeMaster_v1", "author": "digiplay", "last_modified": "2024-03-16 00:22:41+00:00", "downloads": 232, "likes": 2, "library_name": "diffusers", "tags": [ "diffusers", "safetensors", "stable-diffusion", "stable-diffusion-diffusers", "text-to-image"… See the full description on the dataset page: https://huggingface.co/datasets/aravind-selvam/model_card_extractor_qwq.tabularn<1K0 likes8 downloads2y agoHugging Face28reasoning-degeneration-dev /test-updated-extractor-v3 test-updated-extractor-v3 LLM-based math span extraction with canonicalization Dataset Info Rows: 1 Columns: 26 Columns Column Type Description question Value('string') No description provided metadata Value('string') No description provided task_source Value('string') No description provided formatted_prompt List({'content': Value('string'), 'role': Value('string')}) No description provided responses_by_sample List(List(Value('string')))… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/test-updated-extractor-v3.textn<1K0 likes8 downloads10mo agoHugging Face29titou4ng /smolified-ocr-data-extractor-urssaf-2 🤏 smolified-ocr-data-extractor-urssaf-2 Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model titou4ng/smolified-ocr-data-extractor-urssaf-2. 📦 Asset Details Origin: Smolify Foundry (Job ID: 1f7ab49c) Records: 550 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by titou4ng. Generated via Smolify.ai. texttext-generationn<1K0 likes8 downloads8mo agoHugging Face30Algocean /algocean-extractor algocean-extractor 질의와 문서 묶음을 받아 답의 근거 span을 원문 그대로 뽑거나, 없으면 기권하도록 가르치는 LoRA SFT 데이터셋입니다. RAG / 사내 QA 파이프라인의 근거 추출 노드용입니다. 답을 쓰지 않고, 재료가 어디에 있는지만 가리킵니다. 규모 파일 행 수 extractor.jsonl 150,000 extractor.eval.jsonl 2,000 형식: JSONL, messages 3턴 (system / user / assistant) 언어: 한국어 약 70% · 영어 약 30% eval은 학습셋과 별도 생성 어디에 쓰나요 문서 청크에서 문자 단위 근거 인용이 필요한 추출기 “모르면 기권” 정책을 경량 모델에 LoRA로 심을 때 프론티어 모델 앞단의 값싼 grounding 필터 어떤 모델에 LoRA 하나요… See the full description on the dataset page: https://huggingface.co/datasets/Algocean/algocean-extractor.texttext-generation1K<n<10K0 likes8 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.