Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01md-zain /resume-parser-dataset Resume Parser Dataset and Processing Pipeline This repository documents and packages the data pipeline for the resume-parser-model project. The project explores fine-tuning compact language models to turn resume text into structured JSON, using only facts supported by each resume. What is in this repository? The folders preserve the stages used to build the dataset: folder contents raw/ Original source resumes, organized by occupation/category.… See the full description on the dataset page: https://huggingface.co/datasets/md-zain/resume-parser-dataset.documenttext-generation1 likes650 downloads7d agoHugging Face02jugalgajjar /MultiLang-Code-Parser-Dataset MultiLang Code Parser Dataset (MLCPD) MultiLang-Code-Parser-Dataset (MLCPD) provides a large-scale, unified dataset of parsed source code across 10 major programming languages, represented under a universal schema that captures syntax, semantics, and structure in a consistent format. Each entry corresponds to one parsed source file and includes: Language metadata Code-level statistics (lines, errors, AST nodes) Universal Schema JSON (normalized structural representation) MLCPD… See the full description on the dataset page: https://huggingface.co/datasets/jugalgajjar/MultiLang-Code-Parser-Dataset.tabular1M<n<10M2 likes300 downloads1y agoHugging Face03gg676 /ParserV1-modelstextn<1K0 likes297 downloads10mo agoHugging Face04build-small-hackathon /agenda-parser-tool-traces Agenda Parser — tool-calling reasoning traces ReAct tool-calling traces for the Agenda Parser agents: each row is one agent step — a {system, user, assistant} chat example where the assistant emits a single JSON action {"thought", "tool", "args"}. Two agents are covered (tagged by meta.domain): agenda — the uploaded-packet research agent, over real public-meeting agenda packets (tools: list/read items, semantic + exact search, summarize, report). Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-tool-traces.documenttext-generation1K<n<10K0 likes142 downloads4mo agoHugging Face05rdubwiley /agenda-parser-tool-traces Agenda Parser — tool-calling reasoning traces ReAct tool-calling traces for the Agenda Parser agents: each row is one agent step — a {system, user, assistant} chat example where the assistant emits a single JSON action {"thought", "tool", "args"}. Two agents are covered (tagged by meta.domain): agenda — the uploaded-packet research agent, over real public-meeting agenda packets (tools: list/read items, semantic + exact search, summarize, report). Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/rdubwiley/agenda-parser-tool-traces.documenttext-generation1K<n<10K0 likes127 downloads4mo agoHugging Face06inNexus /parser_dataset parser_dataset Parser training and evaluation data for multi-hop QA with long concatenated document contexts. Derived from HotpotQA and 2WikiMultihopQA. Contents Path Split Samples Notes train/hotpotqa_train_process_emb-select_llm-unable.parquet train HotpotQA processed train VERL / ParserRLHFDataset format 2wiki_val/eval_{N}.json eval 128 per file 2WikiMultihopQA, N documents per example hqa_val/eval_{N}.json eval 128 per file HotpotQA, N documents… See the full description on the dataset page: https://huggingface.co/datasets/inNexus/parser_dataset.textquestion-answering10K<n<100K1 likes99 downloads1mo agoHugging Face07davanstrien /pi-trace-parser-sessionstabularn<1K0 likes87 downloads6mo agoHugging Face08emgena /omnimcp_healthtech_hl7_parser_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_healthtech_hl7_parser_teaser.texttext-generationn<1K0 likes80 downloads24d agoHugging Face09gabrielbo /parser-benchimagen<1K0 likes71 downloads5mo agoHugging Face10build-small-hackathon /agenda-parser-models-example-agent-traces Agenda Parser — fine-tuned agent models Three Gemma 4 models fine-tuned to drive the Agenda Parser's ReAct agent: at each step the model emits a single JSON action {"thought","tool","args"} over two toolkits — meeting-agenda packets and Michigan local-government law (Open Meetings Act, FOIA, the Michigan Compiled Laws via Cornell LII). This card doubles as the project write-up; the dataset itself (bottom) is a gallery of example traces from the three models. tier base… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-models-example-agent-traces.texttext-generationn<1K0 likes66 downloads4mo agoHugging Face11Koralzakai /lumi-parser-data Lumi parser distillation data Synthetic bilingual (Hebrew/English) task-parsing data for fine-tuning a small student model, generated by Qwen/Qwen3-235B-A22B-Instruct-2507 on Nebius. Examples: 1465 (1319 train / 146 val) Avg tasks/example: 2.36 Format: OpenAI chat messages (system/user/assistant) in train.jsonl/val.jsonl Built for the Nebius Serverless AI Builders Challenge. License: MIT. text1K<n<10K0 likes42 downloads4mo agoHugging Face12hassno /synth_cv_parser_fakertext10K<n<100K0 likes41 downloads1y agoHugging Face13omeryentur /grok-parser-vrl-940k grok-parser-vrl-940k 940,257 validated (log, grok_pattern) pairs for training models that generate Vector.dev VRL parse_grok! patterns from raw log lines. Files merged_validated.csv — full schema (854 MB) log — raw log line parser — full VRL snippet (e.g. .message = ... | parse_grok!(.message, "...")) grok_pattern — bare grok string extracted from parser target — canonical pygrok output (dict) parsed_output — independent re-application of grok_pattern (sanity check)… See the full description on the dataset page: https://huggingface.co/datasets/omeryentur/grok-parser-vrl-940k.text100K<n<1M0 likes40 downloads5mo agoHugging Face14alperiox /english-date-semantic-parser-data Dataset Overview: semantic_train_en Total Samples: 100000 Random Seed: 42 Noise Probability: 0.3 Generated At: 2026-02-10 14:40:49 Generator Distribution Generator Function Count Percentage Target Weight gen_ambiguous_until 1812 1.81% 0.03 gen_before_after_weekday 2445 2.44% 0.04 gen_complex_weekday_offset 1866 1.87% 0.03 gen_compound 1242 1.24% 0.02 gen_day_after_tomorrow 2493 2.49% 0.04 gen_day_month_written 2531 2.53% 0.04 gen_day_of_month 1880… See the full description on the dataset page: https://huggingface.co/datasets/alperiox/english-date-semantic-parser-data.text100K<n<1M0 likes37 downloads8mo agoHugging Face15myfi /parser_dataset_ner_v1.31text1K<n<10K0 likes32 downloads9mo agoHugging Face16myfi /parser_dataset_ner_val_v1.14text1K<n<10K0 likes31 downloads1y agoHugging Face17myfi /parser_dataset_ner_v1.26text1K<n<10K0 likes31 downloads10mo agoHugging Face18DCAgent2 /DCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260102_021553textn<1K0 likes31 downloads9mo agoHugging Face19wangzihaogithub /job-educational-parser-dataset-08-0-0805 Job Educational Parser Dataset 招聘领域的岗位与学历要求数据集。 输入:岗位描述 -> 输出:学历要求 Splits train: 19w_0701.csv (约 19 万条) test: 2w_0716.csv (约 2 万条) validation: 4w_0708.csv (约 4 万条) 每条数据至少包含字段: user: 职位描述 assistant: 要求的学历(如 "博士、硕士、本科"),遵循从高到低 由 @wangzihaogithub 创建。 tabulartext-generation100K<n<1M0 likes30 downloads1y agoHugging Face20magnifi /parser_user_v41btext1K<n<10K0 likes29 downloads1y agoHugging Face21DCAgent2 /DCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260102_000810textn<1K0 likes26 downloads9mo agoHugging Face22DCAgent2 /DCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260101_132118textn<1K0 likes25 downloads9mo agoHugging Face23DCAgent2 /DCAgent2_terminal_bench_2_laion_exp_tas_parser_xml_traces_20260101_224249textn<1K0 likes25 downloads9mo agoHugging Face24myfi /parser_dataset_ner_mini_v1.18text1K<n<10K0 likes24 downloads1y agoHugging Face25myfi /parser_dataset_ner_v1.49text1K<n<10K0 likes23 downloads6mo agoHugging Face26myfi /parser_dataset_ner_combined_v1.16text10K<n<100K0 likes22 downloads1y agoHugging Face27AFKatz /Recipe-parser Dataset Card for Dataset Name A collection of traditional Mountain Jewish (Gorsky Jewish) recipes from STMEGI.com, containing authentic culinary recipes representing the cultural heritage of the Caucasus Jewish community. This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description A collection of 42 traditional Mountain Jewish (Gorsky Jewish) recipes collected from… See the full description on the dataset page: https://huggingface.co/datasets/AFKatz/Recipe-parser.textn<1K0 likes22 downloads10mo agoHugging Face28myfi /parser_dataset_ner_v1.54text1K<n<10K1 likes22 downloads6mo agoHugging Face29magnifi /parser_user_v44atext1K<n<10K0 likes21 downloads1y agoHugging Face30myfi /parser_dataset_sgpt_v3.8text1K<n<10K0 likes21 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.