Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01regolo /brick-complexity-extractor 🧱 Brick Complexity Extractor Dataset 76,831 user queries labeled by complexity for LLM routing Regolo.ai · Model · Brick SR1 on GitHub · API Docs Overview This dataset provides 76,831 user queries annotated with a complexity label (easy, medium, or hard) indicating the cognitive effort and reasoning depth required to answer each query. It was created to train the Brick Complexity Extractor, a LoRA adapter used in the Brick Semantic Router for… See the full description on the dataset page: https://huggingface.co/datasets/regolo/brick-complexity-extractor.text-classification10K<n<100K2 likes96 downloads6mo agoHugging Face02emgena /omnimcp_browser_dom_structured_extractor_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_browser_dom_structured_extractor_teaser.texttext-generationn<1K0 likes85 downloads20d agoHugging Face03Logistic12 /xhs-image-extractor-20260314115142image1K<n<10K0 likes77 downloads7mo agoHugging Face04emgena /omnimcp_graphrag_triplet_extractor_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_graphrag_triplet_extractor_teaser.texttext-generationn<1K0 likes71 downloads20d agoHugging Face05emgena /multimodal_rag_complex_table_extractor_teaser 🚀 Data Platform - Multi-Modal RAG, Complex Document & Table Extractor (Evaluation Teaser) ⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (500 Samples) & Commercial EULA on Gumroad:👉 Data Platform - Multi-Modal RAG, Complex Document & Table Extractor on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout! 📦 What is Inside the Full Production Package: 500 Verified FAANG v2.0 Scenarios (100%… See the full description on the dataset page: https://huggingface.co/datasets/emgena/multimodal_rag_complex_table_extractor_teaser.texttext-generationn<1K0 likes65 downloads5d agoHugging Face06emgena /omnimcp_episodic_fact_extractor_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_episodic_fact_extractor_teaser.texttext-generationn<1K0 likes64 downloads20d agoHugging Face07Siva2022 /esg-extractor-design-and-code ESG Metric Extractor — Design & Code Package Two files, both copy-paste ready: File Contents DESIGN.md Full design: task framing, multimodal architecture, model choices with 2026 costs, data strategy, training config (TRL-grounded), evaluation, risks, roadmap CODE.md All runnable Colab cells: Part A = v1 text-only pipeline (Qwen2.5-3B QLoRA, data prep, training, eval); Part B = v2 multimodal pipeline (Qwen3-VL-4B QLoRA on page images, teacher labeling, PDF pipeline)… See the full description on the dataset page: https://huggingface.co/datasets/Siva2022/esg-extractor-design-and-code.0 likes60 downloads25d agoHugging Face08keerthanshetty /resume-skill-extractor-dataset Resume Skill Extractor Dataset Dataset Summary This dataset contains 3,050 pre-processed job descriptions with their summaries and required technical skills. It is designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs) to teach them how to parse job postings and extract skill requirements. Data Structure Each row in the dataset is a JSON object containing the following fields: title: The job title (e.g., "Senior Data Scientist"). source:… See the full description on the dataset page: https://huggingface.co/datasets/keerthanshetty/resume-skill-extractor-dataset.text1K<n<10K0 likes35 downloads6mo agoHugging Face09MBMMurad /Bangla_Person_Name_Extractortext1K<n<10K0 likes32 downloads3y agoHugging Face10dedemerve /ILSA-LLM-Extractor-Dataset ILSA LLM Extractor Dataset Project website: https://dedemerve.github.io/ILSA-LLM-Extractor/ Dataset Description This dataset contains structured metadata automatically extracted from 1,756 peer-reviewed articles and reports covering International Large-Scale Assessments (IEA: TIMSS, PIRLS, ICCS; OECD: PISA, TALIS, PIAAC). The extraction pipeline combines PDF parsing, LLM-based structured extraction, and RAG-based synthesis. Pipeline stages: Stage 1: LLM-based… See the full description on the dataset page: https://huggingface.co/datasets/dedemerve/ILSA-LLM-Extractor-Dataset.tabularfeature-extraction10K<n<100K0 likes23 downloads4mo agoHugging Face11Vishal24 /feature_extractortext1K<n<10K0 likes21 downloads2y agoHugging Face12smartytrios /document_data_extractor Dataset Title: OCR-to-JSON Information Extraction Project Overview This dataset is specifically designed for fine-tuning Large Language Models (LLMs) to perform structured data extraction from Optical Character Recognition (OCR) outputs. The primary objective is to convert raw, unstructured text strings—often containing noise, misalignments, and formatting inconsistencies—into valid, machine-readable JSON objects. Dataset Specifications Attribute… See the full description on the dataset page: https://huggingface.co/datasets/smartytrios/document_data_extractor.0 likes20 downloads9mo agoHugging Face13Luimas /claim-extractor-detective-data0 likes20 downloads4mo agoHugging Face14JohnGorri /macro-extractor-flan-t5-synthtext10K<n<100K0 likes20 downloads3mo agoHugging Face15ciaranmacseoin /research_paper_extractortextn<1K3 likes18 downloads2y agoHugging Face16nevvton /size-extractor-dataset-qatext1K<n<10K0 likes18 downloads2y agoHugging Face17titou4ng /smolified-ocr-data-extractor-kbis 🤏 smolified-ocr-data-extractor-kbis Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model titou4ng/smolified-ocr-data-extractor-kbis. 📦 Asset Details Origin: Smolify Foundry (Job ID: 7b974e9e) Records: 0 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by titou4ng. Generated via Smolify.ai. text-generation1K<n<10K0 likes18 downloads8mo agoHugging Face18rishiraj /smolified-ingredient-extractor 🤏 smolified-ingredient-extractor Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model rishiraj/smolified-ingredient-extractor. 📦 Asset Details Origin: Smolify Foundry (Job ID: 65517eae) Records: 9905 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by rishiraj. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes18 downloads6mo agoHugging Face19MinaGabriel /sentence-relevance-extractor Sentence Relevance Extractor (SRE) Sentence Relevance Extractor (SRE) is a large-scale dataset for binary evidence selection in multi-document, multi-hop question answering. The goal: Given a question and a sentence from the context, predict whether this sentence is relevant evidence ("Yes") or irrelevant ("No"). This dataset is suitable for training: Sentence-level RAG rerankers Binary relevance classifiers Optimization-based truth discovery systems Multi-hop QA evidence… See the full description on the dataset page: https://huggingface.co/datasets/MinaGabriel/sentence-relevance-extractor.tabular1M<n<10M0 likes16 downloads11mo agoHugging Face20dhareesh28 /resume-skill-extractor-dataset Resume Skill Extractor Dataset Dataset Summary This dataset contains 3,050 pre-processed job descriptions with their summaries and required technical skills. It is designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs) to teach them how to parse job postings and extract skill requirements. Data Structure Each row in the dataset is a JSON object containing the following fields: title: The job title (e.g., "Senior Data… See the full description on the dataset page: https://huggingface.co/datasets/dhareesh28/resume-skill-extractor-dataset.text1K<n<10K0 likes16 downloads2mo agoHugging Face21smolify /smolified-extractor 🤏 smolified-extractor Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model smolify/smolified-extractor. 📦 Asset Details Origin: Smolify Foundry (Job ID: 13d088d3) Records: 15669 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by smolify. Generated via Smolify.ai. texttext-generation10K<n<100K0 likes15 downloads8mo agoHugging Face22Logistic12 /xhs-image-extractor-v2image1K<n<10K0 likes15 downloads7mo agoHugging Face23logiover /json-ld-schema-meta-tag-extractor-sample-data JSON-LD Schema & Meta Tag Extractor Extract JSON-LD/Schema.org structured data, Meta tags, OpenGraph and Twitter Cards from any URL. Get page title + meta description with a clean JSON output for SEO audits, validation, competitor research and AI datasets. Proxy-ready for large crawls. What the actor scrapes 🧩 JSON-LD Schema & Meta Tag Extractor — Scrape Schema.org, OpenGraph & Meta Tags Extract structured data and SEO metadata from any webpage in seconds. This… See the full description on the dataset page: https://huggingface.co/datasets/logiover/json-ld-schema-meta-tag-extractor-sample-data.textn<1K0 likes14 downloads5mo agoHugging Face24jaeyong2 /keywords-extractor-Kotext10K<n<100K0 likes13 downloads1y agoHugging Face25v-rusu /recipe-extractor-datasetThis dataset was created to finetune a small gemma-3 270M model to extract valid, correct JSON-LD objects from a recipe blog/social media post. It's a completely synthethic dataset. I used Deepseek v3.2 to create the blog posts, JSON-LD extraction, and the reasoning traces. The blogs were created based on this Kaggle All Recipes Dataset. The pipeline to generate this dataset is available on github text1K<n<10K0 likes13 downloads8mo agoHugging Face26titou4ng /smolified-ocr-data-extractor-urssaf 🤏 smolified-ocr-data-extractor-urssaf Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model titou4ng/smolified-ocr-data-extractor-urssaf. 📦 Asset Details Origin: Smolify Foundry (Job ID: 6baf72cd) Records: 1288 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by titou4ng. Generated via Smolify.ai. texttext-generation1K<n<10K0 likes13 downloads7mo agoHugging Face27Algocean /algocean-extractor algocean-extractor 질의와 문서 묶음을 받아 답의 근거 span을 원문 그대로 뽑거나, 없으면 기권하도록 가르치는 LoRA SFT 데이터셋입니다. RAG / 사내 QA 파이프라인의 근거 추출 노드용입니다. 답을 쓰지 않고, 재료가 어디에 있는지만 가리킵니다. 규모 파일 행 수 extractor.jsonl 150,000 extractor.eval.jsonl 2,000 형식: JSONL, messages 3턴 (system / user / assistant) 언어: 한국어 약 70% · 영어 약 30% eval은 학습셋과 별도 생성 어디에 쓰나요 문서 청크에서 문자 단위 근거 인용이 필요한 추출기 “모르면 기권” 정책을 경량 모델에 LoRA로 심을 때 프론티어 모델 앞단의 값싼 grounding 필터 어떤 모델에 LoRA 하나요… See the full description on the dataset page: https://huggingface.co/datasets/Algocean/algocean-extractor.texttext-generation1K<n<10K0 likes13 downloads2mo agoHugging Face28MindCastSogang /word_extractorimage100K<n<1M0 likes12 downloads1mo agoHugging Face29d4nieldev /qpl-value-extractor-dstext10K<n<100K0 likes10 downloads11mo agoHugging Face30Shubhankar444 /smolified-ingredient-extractor 🤏 smolified-ingredient-extractor Intelligence, Distilled. This is a synthetic training corpus generated by the Smolify Foundry. It was used to train the corresponding model Shubhankar444/smolified-ingredient-extractor. 📦 Asset Details Origin: Smolify Foundry (Job ID: 22fdd899) Records: 250 Type: Synthetic Instruction Tuning Data ⚖️ License & Ownership This dataset is a sovereign asset owned by Shubhankar444. Generated via Smolify.ai. texttext-generationn<1K0 likes9 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.