datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xhs-image-extractor-20260314115142omnimcp_browser_dom_structured_extractor_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_browser_dom_structured_extractor_teaser.omnimcp_graphrag_triplet_extractor_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_graphrag_triplet_extractor_teaser.xhs-image-extractor-v2omnimcp_episodic_fact_extractor_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_episodic_fact_extractor_teaser.multimodal_rag_complex_table_extractor_teaser
🚀 Data Platform - Multi-Modal RAG, Complex Document & Table Extractor (Evaluation Teaser)
⚡ Official Free Evaluation Teaser (50 Verified Multi-Turn Scenarios)🏆 Get the Full Production Package (500 Samples) & Commercial EULA on Gumroad:👉 Data Platform - Multi-Modal RAG, Complex Document & Table Extractor on Gumroad🏷️ Use coupon code LAUNCH20 for 20 € off at checkout!
📦 What is Inside the Full Production Package:
500 Verified FAANG v2.0 Scenarios (100%… See the full description on the dataset page: https://huggingface.co/datasets/emgena/multimodal_rag_complex_table_extractor_teaser.resume-skill-extractor-dataset
Resume Skill Extractor Dataset
Dataset Summary
This dataset contains 3,050 pre-processed job descriptions with their summaries and required technical skills. It is designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs) to teach them how to parse job postings and extract skill requirements.
Data Structure
Each row in the dataset is a JSON object containing the following fields:
title: The job title (e.g., "Senior Data Scientist").
source:… See the full description on the dataset page: https://huggingface.co/datasets/keerthanshetty/resume-skill-extractor-dataset.Bangla_Person_Name_ExtractorILSA-LLM-Extractor-Dataset
ILSA LLM Extractor Dataset
Project website: https://dedemerve.github.io/ILSA-LLM-Extractor/
Dataset Description
This dataset contains structured metadata automatically extracted from 1,756 peer-reviewed articles and reports covering International Large-Scale Assessments (IEA: TIMSS, PIRLS, ICCS; OECD: PISA, TALIS, PIAAC). The extraction pipeline combines PDF parsing, LLM-based structured extraction, and RAG-based synthesis.
Pipeline stages:
Stage 1: LLM-based… See the full description on the dataset page: https://huggingface.co/datasets/dedemerve/ILSA-LLM-Extractor-Dataset.size-extractor-dataset-qasmolified-extractor
🤏 smolified-extractor
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-extractor.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 13d088d3)
Records: 15669
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
recipe-extractor-datasetThis dataset was created to finetune a small gemma-3 270M model to extract valid, correct JSON-LD objects from a recipe blog/social media post.
It's a completely synthethic dataset. I used Deepseek v3.2 to create the blog posts, JSON-LD extraction, and the reasoning traces.
The blogs were created based on this Kaggle All Recipes Dataset.
The pipeline to generate this dataset is available on github
json-ld-schema-meta-tag-extractor-sample-data
JSON-LD Schema & Meta Tag Extractor
Extract JSON-LD/Schema.org structured data, Meta tags, OpenGraph and Twitter Cards from any URL. Get page title + meta description with a clean JSON output for SEO audits, validation, competitor research and AI datasets. Proxy-ready for large crawls.
What the actor scrapes
🧩 JSON-LD Schema & Meta Tag Extractor — Scrape Schema.org, OpenGraph & Meta Tags Extract structured data and SEO metadata from any webpage in seconds. This… See the full description on the dataset page: https://huggingface.co/datasets/logiover/json-ld-schema-meta-tag-extractor-sample-data.research_paper_extractorkeywords-extractor-Kosentence-relevance-extractor
Sentence Relevance Extractor (SRE)
Sentence Relevance Extractor (SRE) is a large-scale dataset for binary evidence selection in multi-document, multi-hop question answering.
The goal:
Given a question and a sentence from the context, predict whether this sentence is relevant evidence ("Yes") or irrelevant ("No").
This dataset is suitable for training:
Sentence-level RAG rerankers
Binary relevance classifiers
Optimization-based truth discovery systems
Multi-hop QA evidence… See the full description on the dataset page: https://huggingface.co/datasets/MinaGabriel/sentence-relevance-extractor.macro-extractor-flan-t5-synthsmolified-ingredient-extractor
🤏 smolified-ingredient-extractor
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model rishiraj/smolified-ingredient-extractor.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 65517eae)
Records: 9905
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by rishiraj.
Generated via Smolify.ai.
resume-skill-extractor-dataset
Resume Skill Extractor Dataset
Dataset Summary
This dataset contains 3,050 pre-processed job descriptions with their summaries and required technical skills. It is designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs) to teach them how to parse job postings and extract skill requirements.
Data Structure
Each row in the dataset is a JSON object containing the following fields:
title: The job title (e.g., "Senior Data… See the full description on the dataset page: https://huggingface.co/datasets/dhareesh28/resume-skill-extractor-dataset.smolified-ocr-data-extractor-urssaf
🤏 smolified-ocr-data-extractor-urssaf
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model titou4ng/smolified-ocr-data-extractor-urssaf.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 6baf72cd)
Records: 1288
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by titou4ng.
Generated via Smolify.ai.
word_extractorfeature_extractortest-updated-extractor-v2
test-updated-extractor-v2
LLM-based math span extraction with canonicalization
Dataset Info
Rows: 1
Columns: 26
Columns
Column
Type
Description
question
Value('string')
No description provided
metadata
Value('string')
No description provided
task_source
Value('string')
No description provided
formatted_prompt
List({'content': Value('string'), 'role': Value('string')})
No description provided
responses_by_sample
List(List(Value('string')))… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/test-updated-extractor-v2.smolified-ingredient-extractor
🤏 smolified-ingredient-extractor
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-ingredient-extractor.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 1f92fa68)
Records: 600
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
music-crs-state-extractor-datalora-adapters-are-good-feature-extractors
LORA Adapters are Good Feature Extractors Dataset
This dataset contains images of two sets of categories that are not safe for work (hentai and porn, labelled as 0 and 2 correspondingly) and one neutral category, labelled as 2.
The dataset is the source data for training a zoo of LORA adapters on sample images from each category. Adapters representations will then be used as input data to a weight-space model
in an experiment to verify whether WS models operating in low rank… See the full description on the dataset page: https://huggingface.co/datasets/jacekduszenko/lora-adapters-are-good-feature-extractors.model_card_extractor_qwq
Dataset card for model_card_extractor_qwq
This dataset was made with Curator.
Dataset details
A sample from the dataset:
{
"modelId": "digiplay/XtReMixAnimeMaster_v1",
"author": "digiplay",
"last_modified": "2024-03-16 00:22:41+00:00",
"downloads": 232,
"likes": 2,
"library_name": "diffusers",
"tags": [
"diffusers",
"safetensors",
"stable-diffusion",
"stable-diffusion-diffusers",
"text-to-image"… See the full description on the dataset page: https://huggingface.co/datasets/aravind-selvam/model_card_extractor_qwq.test-updated-extractor-v3
test-updated-extractor-v3
LLM-based math span extraction with canonicalization
Dataset Info
Rows: 1
Columns: 26
Columns
Column
Type
Description
question
Value('string')
No description provided
metadata
Value('string')
No description provided
task_source
Value('string')
No description provided
formatted_prompt
List({'content': Value('string'), 'role': Value('string')})
No description provided
responses_by_sample
List(List(Value('string')))… See the full description on the dataset page: https://huggingface.co/datasets/reasoning-degeneration-dev/test-updated-extractor-v3.smolified-ocr-data-extractor-urssaf-2
🤏 smolified-ocr-data-extractor-urssaf-2
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model titou4ng/smolified-ocr-data-extractor-urssaf-2.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 1f7ab49c)
Records: 550
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by titou4ng.
Generated via Smolify.ai.
algocean-extractor
algocean-extractor
질의와 문서 묶음을 받아 답의 근거 span을 원문 그대로 뽑거나, 없으면 기권하도록 가르치는 LoRA SFT 데이터셋입니다.
RAG / 사내 QA 파이프라인의 근거 추출 노드용입니다. 답을 쓰지 않고, 재료가 어디에 있는지만 가리킵니다.
규모
파일
행 수
extractor.jsonl
150,000
extractor.eval.jsonl
2,000
형식: JSONL, messages 3턴 (system / user / assistant)
언어: 한국어 약 70% · 영어 약 30%
eval은 학습셋과 별도 생성
어디에 쓰나요
문서 청크에서 문자 단위 근거 인용이 필요한 추출기
“모르면 기권” 정책을 경량 모델에 LoRA로 심을 때
프론티어 모델 앞단의 값싼 grounding 필터
어떤 모델에 LoRA 하나요… See the full description on the dataset page: https://huggingface.co/datasets/Algocean/algocean-extractor.
