datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-RL-Instruction-Following-Structured-Outputs-v2
Dataset Description:
Split 1: Direct Generation tests the model’s ability to perform freeform text structured outputs on JSON, YAML, and XML data, varying the complexity and presentation of the schema.
Split 2: Diversified Tasks adds 2 additional output formats: TOML and CSV, while increasing problem types to Direct Extraction from document, Translation between formats, Multistep Translation from known data, Multistep Extraction from unrelated context, Schema-Only Generation for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2.sharegpt-structured-output-json
ShareGPT-Formatted Dataset for Structured JSON Output
Dataset Description
This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios.
Usage
This dataset can be used to train LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-structured-output-json.Nemotron-RL-instruction_following-structured_outputs
Dataset Description:
The Nemotron-RL-instruction_following-structured_outputs dataset tests the ability of the model to follow output formatting instructions under schema constraints under the JSON format. Each problem consists of three components: The document, output formatting Instruction (Schema), and question. The dataset varies the difficulty of each problem by varying the location of instructions, the comprehensiveness of instructions, the complexity of the schema, and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-instruction_following-structured_outputs.sft-tool-calling-structured-output-v1
vericava/sft-tool-calling-structured-output-v1
Dataset to train (SFT) 3-20B LLMs for tool calling and structured outputs/classifications.
Includes contents in English as well as some Japanese.
structured-output-training-pool
Structured output training pool
Public JSON schemas, and public text paired with the structured record it describes, from the
collections named below, read at the pinned revisions given there and laid out twice. Train on
either layer or on both.
pool.jsonl
Every source rewritten into one shape, 114103 rows, one JSON object per line, with these fields.
Field
What it holds
id
a row identifier unique within this file
request
the text the record is to be… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/structured-output-training-pool.structured-output-sft-100k
Structured Output SFT (100K)
100,000 ShareGPT conversations demonstrating correct generation of structured data formats: JSON, YAML, CSV, XML, Markdown tables, JSON Schema, and OpenAPI fragments. Each example pairs a natural language specification with a valid, well-formed output.
Motivation
Structured output generation is among the most commercially critical LLM capabilities. Models fail in characteristic ways:
Invalid JSON: unclosed brackets, trailing commas… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/structured-output-sft-100k.german-structured-output
German Structured Output Dataset 🇩🇪
GDPR & EU AI Act compliant German dataset for training structured output capabilities in LLMs.
Overview
This dataset contains 4,521 examples across 7 task types for training language models to produce structured outputs (JSON, function calls, schema-following generation) from German text. It is the first dedicated German structured output dataset, filling a critical gap in the German NLP ecosystem.
Key Features
🇩🇪… See the full description on the dataset page: https://huggingface.co/datasets/philipp-zettl/german-structured-output.Nemotron-RL-Instruction-Following-Structured-Outputs-v2-prompt-only
Nemotron-RL-Instruction-Following-Structured-Outputs-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Structured-Outputs-v2-prompt-only.synthetic-structured-output-dataset
Synthetic Structured Output Dataset (SFT + DPO)
Synthetic training corpus for structured-output model tuning. This package contains SFT and DPO data focused on JSON schema compliance, structured extraction, and function calling.
Included files
sft_synthetic_json.jsonl — SFT samples for schema-conditioned JSON generation
sft_synthetic_extraction.jsonl — SFT samples for text-to-structured extraction
dpo_structured_output.jsonl — DPO chosen/rejected pairs for structured… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/synthetic-structured-output-dataset.structured-outputs-calibration-v1
[!TIP]
Support this work: donate.sybilsolutions.ai
REAP surfaces: GLM | MiniMax | Qwen | Gemma | Paper | Code | PR17 | Cerebras Collection
structured-outputs-calibration-v1
Structured-output calibration set for REAP observer runs, focused on preserving:
strict JSON generation
schema-conditioned JSON responses
Mermaid diagram generation
fenced Mermaid block formatting
Contents
data.jsonl: normalized calibration records
summary.json: build summary with source… See the full description on the dataset page: https://huggingface.co/datasets/0xSero/structured-outputs-calibration-v1.Nemotron-RL-Instruction-Following-Structured-Outputs-v2-preferencenemotron-gym-structured-outputs-v3
laion/nemotron-gym-structured-outputs-v3
Harbor task-binary dataset (53,870 tasks) converted from nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2
(part of the nvidia/Nemotron-Post-Training-v3 collection).
Each row is a valid Harbor
task binary: columns path (str) and task_binary (gzip tar). Converted with the
OpenThoughts-Agent data.nemotron_gym framework.
Grading: JSON/YAML/TOML schema validation; XML/CSV structural (well-formed + required keys).
pydantic_ai_structured_output_resilience_teaser
🛡️ CodeArchitect - Pydantic AI Structured Output & Schema Drift Guard (Teaser & 1-Click Ollama Engine)
FAANG v2.0 Standard • Enterprise Autonomous Sentinel
This free teaser contains the 1-Click Ollama Modelfile, the official EU_AI_ACT_ANNEX_IV_AUDIT.md compliance audit, and an AST-verified 50-sample evaluation slice.
🚀 1-Click Local Run with Ollama
ollama create pydantic_ai_structured_output_resilience -f Modelfile
ollama run… See the full description on the dataset page: https://huggingface.co/datasets/emgena/pydantic_ai_structured_output_resilience_teaser.nemotron-gym-structured-outputs-v4
laion/nemotron-gym-structured-outputs-v4
Harbor task-binary dataset (53,870 tasks) converted from nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2
(part of nvidia/Nemotron-Post-Training-v3).
Columns path (str) + task_binary (gzip tar). Converted with the
OpenThoughts-Agent data.nemotron_gym framework.
Grading: JSON/YAML/TOML schema validation; XML/CSV structural.
What changed vs the prior version
This version fixes the answer-delivery contract for… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/nemotron-gym-structured-outputs-v4.json-structured-output-dpo-3k
JSON Structured Output DPO Pairs (3K)
DPO preference pairs for training LLMs to produce valid, schema-compliant JSON output.
Motivation
Structured output (JSON mode) is critical for production AI applications — parsers fail, pipelines break, and downstream processing errors when models output malformed JSON, use wrong field names, or wrap responses in markdown. This dataset trains strict schema adherence.
Dataset Description
3,000 preference pairs… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/json-structured-output-dpo-3k.Nemotron-RL-Instruction-Following-Structured-Outputs-v2-direct-completethis is based on the first part of this dataset: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2
except thinking traces have been added, the final output has been added and validated. around 8k rows failed validation, so the final result is ~20k rows
format is simplified into "prompt", "thinking", "result" columns
big thanks to Delo on the Unsloth discord for helping out
Nemotron-RL-Instruction-Following-Structured-Outputs-v2
Dataset Description:
Split 1: Direct Generation tests the model’s ability to perform freeform text structured outputs on JSON, YAML, and XML data, varying the complexity and presentation of the schema.
Split 2: Diversified Tasks adds 2 additional output formats: TOML and CSV, while increasing problem types to Direct Extraction from document, Translation between formats, Multistep Translation from known data, Multistep Extraction from unrelated context, Schema-Only Generation for… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/Nemotron-RL-Instruction-Following-Structured-Outputs-v2.distilabel-example-output-structuredstructured_outputr1_outputs_cd3arg_translator_gpt5-mini_structuredr1_outputs_cd3arg_translator_gpt5_structuredr1_outputs_cd3arg_translator_gpt5mini_structuredreview_prompts_for_structured_outputTurkish-Structured-Output
