Team Ai
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-RL-Instruction-Following-Structured-Outputs-v2 Dataset Description: Split 1: Direct Generation tests the model’s ability to perform freeform text structured outputs on JSON, YAML, and XML data, varying the complexity and presentation of the schema. Split 2: Diversified Tasks adds 2 additional output formats: TOML and CSV, while increasing problem types to Direct Extraction from document, Translation between formats, Multistep Translation from known data, Multistep Extraction from unrelated context, Schema-Only Generation for… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2.texttext-generation10K<n<100K8 likes1k downloads7d agoHugging Face02Arun63 /sharegpt-structured-output-json ShareGPT-Formatted Dataset for Structured JSON Output Dataset Description This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios. Usage This dataset can be used to train LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-structured-output-json.texttext-generationn<1K7 likes702 downloads2y agoHugging Face03nvidia /Nemotron-RL-instruction_following-structured_outputs Dataset Description: The Nemotron-RL-instruction_following-structured_outputs dataset tests the ability of the model to follow output formatting instructions under schema constraints under the JSON format. Each problem consists of three components: The document, output formatting Instruction (Schema), and question. The dataset varies the difficulty of each problem by varying the location of instructions, the comprehensiveness of instructions, the complexity of the schema, and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-instruction_following-structured_outputs.text1K<n<10K41 likes503 downloads8d agoHugging Face04vericava /sft-tool-calling-structured-output-v1 vericava/sft-tool-calling-structured-output-v1 Dataset to train (SFT) 3-20B LLMs for tool calling and structured outputs/classifications. Includes contents in English as well as some Japanese. texttext-classification100K<n<1M2 likes275 downloads8mo agoHugging Face05Emulated-Inc /structured-output-training-pool Structured output training pool Public JSON schemas, and public text paired with the structured record it describes, from the collections named below, read at the pinned revisions given there and laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one shape, 114103 rows, one JSON object per line, with these fields. Field What it holds id a row identifier unique within this file request the text the record is to be… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/structured-output-training-pool.text10K<n<100K0 likes139 downloads25d agoHugging Face06stindardlogic /structured-output-sft-100k Structured Output SFT (100K) 100,000 ShareGPT conversations demonstrating correct generation of structured data formats: JSON, YAML, CSV, XML, Markdown tables, JSON Schema, and OpenAPI fragments. Each example pairs a natural language specification with a valid, well-formed output. Motivation Structured output generation is among the most commercially critical LLM capabilities. Models fail in characteristic ways: Invalid JSON: unclosed brackets, trailing commas… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/structured-output-sft-100k.texttext-generation100K<n<1M1 likes136 downloads3mo agoHugging Face07philipp-zettl /german-structured-output German Structured Output Dataset 🇩🇪 GDPR & EU AI Act compliant German dataset for training structured output capabilities in LLMs. Overview This dataset contains 4,521 examples across 7 task types for training language models to produce structured outputs (JSON, function calls, schema-following generation) from German text. It is the first dedicated German structured output dataset, filling a critical gap in the German NLP ecosystem. Key Features 🇩🇪… See the full description on the dataset page: https://huggingface.co/datasets/philipp-zettl/german-structured-output.texttext-generation1K<n<10K0 likes45 downloads6mo agoHugging Face08jamesdborin /Nemotron-RL-Instruction-Following-Structured-Outputs-v2-prompt-only Nemotron-RL-Instruction-Following-Structured-Outputs-v2-prompt-only Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Structured-Outputs-v2-prompt-only.0 likes45 downloads3mo agoHugging Face09mdonigian /synthetic-structured-output-dataset Synthetic Structured Output Dataset (SFT + DPO) Synthetic training corpus for structured-output model tuning. This package contains SFT and DPO data focused on JSON schema compliance, structured extraction, and function calling. Included files sft_synthetic_json.jsonl — SFT samples for schema-conditioned JSON generation sft_synthetic_extraction.jsonl — SFT samples for text-to-structured extraction dpo_structured_output.jsonl — DPO chosen/rejected pairs for structured… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/synthetic-structured-output-dataset.text-generation10K<n<100K0 likes38 downloads7mo agoHugging Face100xSero /structured-outputs-calibration-v1 [!TIP] Support this work: donate.sybilsolutions.ai REAP surfaces: GLM | MiniMax | Qwen | Gemma | Paper | Code | PR17 | Cerebras Collection structured-outputs-calibration-v1 Structured-output calibration set for REAP observer runs, focused on preserving: strict JSON generation schema-conditioned JSON responses Mermaid diagram generation fenced Mermaid block formatting Contents data.jsonl: normalized calibration records summary.json: build summary with source… See the full description on the dataset page: https://huggingface.co/datasets/0xSero/structured-outputs-calibration-v1.text-generation0 likes37 downloads6mo agoHugging Face11electroglyph /Nemotron-RL-Instruction-Following-Structured-Outputs-v2-preferencetext10K<n<100K0 likes36 downloads3mo agoHugging Face12open-athena /nemotron-gym-structured-outputs-v3 laion/nemotron-gym-structured-outputs-v3 Harbor task-binary dataset (53,870 tasks) converted from nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2 (part of the nvidia/Nemotron-Post-Training-v3 collection). Each row is a valid Harbor task binary: columns path (str) and task_binary (gzip tar). Converted with the OpenThoughts-Agent data.nemotron_gym framework. Grading: JSON/YAML/TOML schema validation; XML/CSV structural (well-formed + required keys). texttext-generation10K<n<100K0 likes34 downloads1mo agoHugging Face13emgena /pydantic_ai_structured_output_resilience_teaser 🛡️ CodeArchitect - Pydantic AI Structured Output & Schema Drift Guard (Teaser & 1-Click Ollama Engine) FAANG v2.0 Standard • Enterprise Autonomous Sentinel This free teaser contains the 1-Click Ollama Modelfile, the official EU_AI_ACT_ANNEX_IV_AUDIT.md compliance audit, and an AST-verified 50-sample evaluation slice. 🚀 1-Click Local Run with Ollama ollama create pydantic_ai_structured_output_resilience -f Modelfile ollama run… See the full description on the dataset page: https://huggingface.co/datasets/emgena/pydantic_ai_structured_output_resilience_teaser.n<1K0 likes31 downloads3d agoHugging Face14open-athena /nemotron-gym-structured-outputs-v4 laion/nemotron-gym-structured-outputs-v4 Harbor task-binary dataset (53,870 tasks) converted from nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2 (part of nvidia/Nemotron-Post-Training-v3). Columns path (str) + task_binary (gzip tar). Converted with the OpenThoughts-Agent data.nemotron_gym framework. Grading: JSON/YAML/TOML schema validation; XML/CSV structural. What changed vs the prior version This version fixes the answer-delivery contract for… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/nemotron-gym-structured-outputs-v4.texttext-generation10K<n<100K0 likes23 downloads1mo agoHugging Face15stindardlogic /json-structured-output-dpo-3k JSON Structured Output DPO Pairs (3K) DPO preference pairs for training LLMs to produce valid, schema-compliant JSON output. Motivation Structured output (JSON mode) is critical for production AI applications — parsers fail, pipelines break, and downstream processing errors when models output malformed JSON, use wrong field names, or wrap responses in markdown. This dataset trains strict schema adherence. Dataset Description 3,000 preference pairs… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/json-structured-output-dpo-3k.texttext-generation1K<n<10K0 likes19 downloads3mo agoHugging Face16electroglyph /Nemotron-RL-Instruction-Following-Structured-Outputs-v2-direct-completethis is based on the first part of this dataset: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Structured-Outputs-v2 except thinking traces have been added, the final output has been added and validated. around 8k rows failed validation, so the final result is ~20k rows format is simplified into "prompt", "thinking", "result" columns big thanks to Delo on the Unsloth discord for helping out 1 likes15 downloads4mo agoHugging Face17Arsh9210 /Nemotron-RL-Instruction-Following-Structured-Outputs-v2 Dataset Description: Split 1: Direct Generation tests the model’s ability to perform freeform text structured outputs on JSON, YAML, and XML data, varying the complexity and presentation of the schema. Split 2: Diversified Tasks adds 2 additional output formats: TOML and CSV, while increasing problem types to Direct Extraction from document, Translation between formats, Multistep Translation from known data, Multistep Extraction from unrelated context, Schema-Only Generation for… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/Nemotron-RL-Instruction-Following-Structured-Outputs-v2.texttext-generation10K<n<100K0 likes13 downloads2mo agoHugging Face18Benjoyo /distilabel-example-output-structuredtextn<1K0 likes8 downloads7mo agoHugging Face19BeardedMonster /structured_output0 likes6 downloads1y agoHugging Face20TAUR-dev /r1_outputs_cd3arg_translator_gpt5-mini_structuredtextn<1K0 likes6 downloads11mo agoHugging Face21TAUR-dev /r1_outputs_cd3arg_translator_gpt5_structuredtextn<1K0 likes6 downloads11mo agoHugging Face22TAUR-dev /r1_outputs_cd3arg_translator_gpt5mini_structuredtextn<1K0 likes5 downloads11mo agoHugging Face23shivanikerai /review_prompts_for_structured_outputtextn<1K2 likes4 downloads3y agoHugging Face24aliarda /Turkish-Structured-Output0 likes4 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.