datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
resume-parser-dataset
Resume Parser Dataset and Processing Pipeline
This repository documents and packages the data pipeline for the resume-parser-model project. The project explores fine-tuning compact language models to turn resume text into structured JSON, using only facts supported by each resume.
What is in this repository?
The folders preserve the stages used to build the dataset:
folder
contents
raw/
Original source resumes, organized by occupation/category.… See the full description on the dataset page: https://huggingface.co/datasets/md-zain/resume-parser-dataset.MultiLang-Code-Parser-Dataset
MultiLang Code Parser Dataset (MLCPD)
MultiLang-Code-Parser-Dataset (MLCPD) provides a large-scale, unified dataset of parsed source code across 10 major programming languages, represented under a universal schema that captures syntax, semantics, and structure in a consistent format.
Each entry corresponds to one parsed source file and includes:
Language metadata
Code-level statistics (lines, errors, AST nodes)
Universal Schema JSON (normalized structural representation)
MLCPD… See the full description on the dataset page: https://huggingface.co/datasets/jugalgajjar/MultiLang-Code-Parser-Dataset.ParserV1-modelsagenda-parser-tool-traces
Agenda Parser — tool-calling reasoning traces
ReAct tool-calling traces for the Agenda Parser
agents: each row is one agent step — a {system, user, assistant} chat example
where the assistant emits a single JSON action {"thought", "tool", "args"}.
Two agents are covered (tagged by meta.domain):
agenda — the uploaded-packet research agent, over real public-meeting agenda
packets (tools: list/read items, semantic + exact search, summarize, report).
Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-tool-traces.agenda-parser-tool-traces
Agenda Parser — tool-calling reasoning traces
ReAct tool-calling traces for the Agenda Parser
agents: each row is one agent step — a {system, user, assistant} chat example
where the assistant emits a single JSON action {"thought", "tool", "args"}.
Two agents are covered (tagged by meta.domain):
agenda — the uploaded-packet research agent, over real public-meeting agenda
packets (tools: list/read items, semantic + exact search, summarize, report).
Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/rdubwiley/agenda-parser-tool-traces.parser_dataset
parser_dataset
Parser training and evaluation data for multi-hop QA with long concatenated document contexts.
Derived from HotpotQA and 2WikiMultihopQA.
Contents
Path
Split
Samples
Notes
train/hotpotqa_train_process_emb-select_llm-unable.parquet
train
HotpotQA processed train
VERL / ParserRLHFDataset format
2wiki_val/eval_{N}.json
eval
128 per file
2WikiMultihopQA, N documents per example
hqa_val/eval_{N}.json
eval
128 per file
HotpotQA, N documents… See the full description on the dataset page: https://huggingface.co/datasets/inNexus/parser_dataset.pi-trace-parser-sessionsomnimcp_healthtech_hl7_parser_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_healthtech_hl7_parser_teaser.parser-benchagenda-parser-models-example-agent-traces
Agenda Parser — fine-tuned agent models
Three Gemma 4 models fine-tuned to drive the Agenda Parser's ReAct agent: at each step
the model emits a single JSON action {"thought","tool","args"} over two toolkits —
meeting-agenda packets and Michigan local-government law (Open Meetings Act, FOIA,
the Michigan Compiled Laws via Cornell LII). This card doubles as the project write-up; the
dataset itself (bottom) is a gallery of example traces from the three models.
tier
base… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-models-example-agent-traces.lumi-parser-data
Lumi parser distillation data
Synthetic bilingual (Hebrew/English) task-parsing data for fine-tuning a small student model, generated by Qwen/Qwen3-235B-A22B-Instruct-2507 on Nebius.
Examples: 1465 (1319 train / 146 val)
Avg tasks/example: 2.36
Format: OpenAI chat messages (system/user/assistant) in train.jsonl/val.jsonl
Built for the Nebius Serverless AI Builders Challenge. License: MIT.
synth_cv_parser_fakergrok-parser-vrl-940k
grok-parser-vrl-940k
940,257 validated (log, grok_pattern) pairs for training models that
generate Vector.dev VRL parse_grok! patterns from raw log lines.
Files
merged_validated.csv — full schema (854 MB)
log — raw log line
parser — full VRL snippet (e.g. .message = ... | parse_grok!(.message, "..."))
grok_pattern — bare grok string extracted from parser
target — canonical pygrok output (dict)
parsed_output — independent re-application of grok_pattern (sanity check)… See the full description on the dataset page: https://huggingface.co/datasets/omeryentur/grok-parser-vrl-940k.english-date-semantic-parser-data
Dataset Overview: semantic_train_en
Total Samples: 100000
Random Seed: 42
Noise Probability: 0.3
Generated At: 2026-02-10 14:40:49
Generator Distribution
Generator Function
Count
Percentage
Target Weight
gen_ambiguous_until
1812
1.81%
0.03
gen_before_after_weekday
2445
2.44%
0.04
gen_complex_weekday_offset
1866
1.87%
0.03
gen_compound
1242
1.24%
0.02
gen_day_after_tomorrow
2493
2.49%
0.04
gen_day_month_written
2531
2.53%
0.04
gen_day_of_month
1880… See the full description on the dataset page: https://huggingface.co/datasets/alperiox/english-date-semantic-parser-data.parser_dataset_ner_v1.31parser_dataset_ner_val_v1.14parser_dataset_ner_v1.26DCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260102_021553job-educational-parser-dataset-08-0-0805
Job Educational Parser Dataset
招聘领域的岗位与学历要求数据集。
输入:岗位描述 -> 输出:学历要求
Splits
train: 19w_0701.csv (约 19 万条)
test: 2w_0716.csv (约 2 万条)
validation: 4w_0708.csv (约 4 万条)
每条数据至少包含字段:
user: 职位描述
assistant: 要求的学历(如 "博士、硕士、本科"),遵循从高到低
由 @wangzihaogithub 创建。
parser_user_v41bDCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260102_000810DCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260101_132118DCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260101_224249parser_dataset_ner_mini_v1.18parser_dataset_ner_v1.49parser_dataset_ner_combined_v1.16Recipe-parser
Dataset Card for Dataset Name
A collection of traditional Mountain Jewish (Gorsky Jewish) recipes from STMEGI.com, containing authentic culinary recipes representing the cultural heritage of the Caucasus Jewish community.
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
A collection of 42 traditional Mountain Jewish (Gorsky Jewish) recipes collected from… See the full description on the dataset page: https://huggingface.co/datasets/AFKatz/Recipe-parser.parser_dataset_ner_v1.54parser_user_v44aparser_dataset_sgpt_v3.8
