datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Caselaw_Access_Project_JSON
The Caselaw Access Project
In collaboration with Ravel Law, Harvard Law Library digitized over 40 million U.S. court decisions consisting of 6.7 million cases from the last 360 years into a dataset that is widely accessible to use. Access a bulk download of the data through the Caselaw Access Project API (CAPAPI): https://case.law/caselaw/
Find more information about accessing state and federal written court decisions of common law through the bulk data service documentation here:… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/Caselaw_Access_Project_JSON.JSONSchemaBench
JSONSchemaBench
JSONSchemaBench is a benchmark of real-world JSON schemas designed to evaluate structured output generation for Large Language Models (LLMs). It contains approximately 10,000 JSON schemas, capturing diverse constraints and complexities.
import datasets
from datasets import load_dataset
def main():
# Inspect the available subsets of the datasetall_subsets = datasets.get_dataset_config_names("epfl-dlab/JSONSchemaBench")
print("Available subsets:"… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/JSONSchemaBench.sharegpt-quizz-generation-json-output
ShareGPT-Formatted Dataset for Quizz Generation in Structured JSON Output
Dataset Description
This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate quizz in structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios.
Usage
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-quizz-generation-json-output.sharegpt-structured-output-json
ShareGPT-Formatted Dataset for Structured JSON Output
Dataset Description
This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios.
Usage
This dataset can be used to train LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-structured-output-json.otel-test-snippet-jsonl
⚠️ TEST DATASET - DO NOT USE FOR PRODUCTION
This is a small test snippet for internal validation purposes only.
This dataset contains a subset of OpenTelemetry traces from various LLM inference benchmarks. It is intended for testing dataset infrastructure and should NOT be used for research, benchmarking, or production purposes.
Dataset Structure
The dataset contains OpenTelemetry traces organized by:
Benchmark: appworld, tau2_telecom
Agent Framework: openai_solo… See the full description on the dataset page: https://huggingface.co/datasets/lenadan/otel-test-snippet-jsonl.home-commands-json-v1
Home Commands JSON v1
English and Romanized Hindi/Hinglish home commands mapped to structured JSON.
This release contains the exact source and processed snapshot associated
with Superfast Tiny Home Robotics JSON 1M v1.
It also includes a separately authored evaluation challenge set. Native-script
Hindi appears only in a small later diagnostic, not as a supported training
language claim.
Publisher: sraivante. Release date: 2026-09-25. Version: v1.0.2.
Configurations… See the full description on the dataset page: https://huggingface.co/datasets/sraivante/home-commands-json-v1.cfr-rag-jsonECFR from 06/2025
Ascend-COT-v2-json
AscendKernelGen/Ascend-COT-v2-json
AscendKernelGen/Ascend-CoT-v2-json contains a subset of the full Ascend-CoT dataset, which will be released in stages. The Ascend-CoT Dataset is a high-quality, domain-specific dataset that incorporates Chain-of-Thought (CoT) reasoning derived from real-world kernel implementations. It combines three types of reasoning: documentation-based reasoning, code-centric reasoning extracted from actual NPU kernel code, and general reasoning chains that… See the full description on the dataset page: https://huggingface.co/datasets/AscendKernelGen/Ascend-COT-v2-json.Ascend-CoT-v3-json
Ascend-CoT-v3-json
Ascend-CoT-v3-json is an Ascend C / CANN supervised fine-tuning dataset for custom operator development. It contains cleaned CoT-style samples for Ascend C kernel implementation, tiling logic, CANN API usage, debugging, and operator-development reasoning.
The release is organized into two final SFT subsets in one dataset repository.
Related Artifacts
Paper: AscendKernelGen: A Systematic Study of LLM-Based Kernel Generation for Neural… See the full description on the dataset page: https://huggingface.co/datasets/AscendKernelGen/Ascend-CoT-v3-json.support-json-ru
Support-JSON-RU
Synthetic Russian SaaS support data for policy-conditioned JSON decisions and draft replies. The task supplies customer text, company policies, sourced facts and available capabilities; the model predicts a nine-field decision rather than memorizing a single company's policy.
Русский SaaS-support: обращение + правила + факты → категория, приоритет, настроение, действие, черновик ответа и эскалация.
Model · Dataset files · License
Configurations… See the full description on the dataset page: https://huggingface.co/datasets/A11Sunday/support-json-ru.home-commands-json-v3
Home Commands JSON v3
English and Hinglish (Romanised Hindi) home-automation instructions mapped to a
single JSON device command, with optional conditional rules. This is the
exact training mix behind
Superfast Tiny Home Robotics JSON 1M v3,
plus the generators that produced it.
{"instruction": "agar room ka temperature 42 se upar jaye to bedroom ka heater band kar do",
"output": "{\"activity\":\"heating\",\"subject\":\"bedroom_room_heater\",\"action\":\"OFF\"… See the full description on the dataset page: https://huggingface.co/datasets/sraivante/home-commands-json-v3.cc-task1-json
CC captions → task_1 structured JSON
Task 1: V1 Complete
Total Time: 50 hours 3x 6000 blackwell pros
Host: Google Colab
Model: AbstractPhil/qwen3.5-0.8b-task_1-lora
Task: Converting plain English image prompts to JSON containing similar assessments using subjective analysis.
Biases: Guaranteed - Filter nulls before training anything.
V1 Limitations
Extracted from the qwen3.5-0.8b task_1 V1 lora.
Context and associations limited
Topic faults and invalid context… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/cc-task1-json.json_data_extraction
Diverse Restricted JSON Data Extraction
Curated by: The paraloq analytics team.
Uses
Benchmark restricted JSON data extraction (text + JSON schema -> JSON instance)
Fine-Tune data extraction model (text + JSON schema -> JSON instance)
Fine-Tune JSON schema Retrieval model (text -> retriever -> most adequate JSON schema)
Out-of-Scope Use
Intended for research purposes only.
Dataset Structure
The data comes with the following fields:
title: The… See the full description on the dataset page: https://huggingface.co/datasets/paraloq/json_data_extraction.json-schema-instances-training-pool
JSON schema and instance training pool
Real JSON Schemas from the public collections named below, read at the pinned revisions given there,
each paired where possible with documents that satisfy it, laid out twice. Train on either layer or
on both.
pool.jsonl
Every source rewritten into one shape, 20004 rows, one JSON object per line, with these fields.
Field
What it holds
id
a row identifier unique within this file
prompt
the request a model would… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/json-schema-instances-training-pool.hermes-function-calling-v1-jsonl
Hermes Function-Calling V1
This dataset is the compilation of structured output and function calling data used in the Hermes 2 Pro series of models.
This repository contains a structured output dataset with function-calling conversations, json-mode, agentic json-mode and structured extraction samples, designed to train LLM models in performing function calls and returning structured output based on natural language instructions. The dataset features various conversational scenarios… See the full description on the dataset page: https://huggingface.co/datasets/minpeter/hermes-function-calling-v1-jsonl.resume-json-extraction-5k
Dataset Card for resume-json-extraction-5k
Dataset Description
This dataset contains 4,879 resume examples formatted for fine-tuning language models to extract structured JSON information from resume text.
Dataset Summary
The dataset consists of resume text paired with structured JSON outputs containing:
Job titles (current and previous)
Companies (current and previous)
Years of experience
Seniority level
Primary domain and industries
Core and secondary skills… See the full description on the dataset page: https://huggingface.co/datasets/sandeeppanem/resume-json-extraction-5k.bigbench_jsonlBIG-Bench but it doesn't require the hellish dependencies (tensorflow, pypi-bigbench, protobuf) of the official version.
dataset = load_dataset("tasksource/bigbench",'movie_recommendation')
Code to reproduce:
https://colab.research.google.com/drive/1MKdLdF7oqrSQCeavAcsEnPdI85kD0LzU?usp=sharing
Datasets are capped to 50k examples to keep things light.
I also removed the default split when train was available also to save space, as default=train+val.
@article{srivastava2022beyond… See the full description on the dataset page: https://huggingface.co/datasets/NJUDeepEngine/bigbench_jsonl.json-extraction
Rob Dixon's JSON Extraction Dataset
A synthetic dataset for training JSON extraction models, generated using Claude 3 Haiku.
Dataset Overview
This dataset contains paired examples of:
Instructions: Natural language task descriptions asking to extract information
Text documents: Source content containing information to extract
JSON outputs: Structured data extracted from the text
The dataset is designed for training smaller models on constrained context lengths, with… See the full description on the dataset page: https://huggingface.co/datasets/robdixon/json-extraction.ShortStory-SFT-jsonl
Public Domain Short Fiction with Prompts
719 complete short stories (300 to 2500 words) by 26 authors whose work is in the public domain,
each paired with a natural-language request that could plausibly have produced it. Built for supervised
fine-tuning of small language models on fiction, where the usual sources (forum stories, model-generated
stories) lack the structural control of published short fiction.
Fields
field
description
id
stable id (hash… See the full description on the dataset page: https://huggingface.co/datasets/Travis-ML/ShortStory-SFT-jsonl.ai-culture-multilingual-json-dolma
AI-Culture Multilingual JSON + DOLMA Corpus
16M words · 12 languages · CC-BY-4.0
The AI-Culture corpus contains 5K articles providing comprehensive philosophical and cultural content, exploring the intersection of technology, artificial intelligence, and human culture, perfectly aligned across 12 languages. All content maintains identical parallel structure across translations with zero duplication and editor-curated quality.
This project is maintained by a non-profit digital… See the full description on the dataset page: https://huggingface.co/datasets/AI-Culture-Commons/ai-culture-multilingual-json-dolma.NSFW-Stories-JsonLConverted to JsonL from: bluuwhale/nsfwstory2
json-coco-format
JSON COCO Format — task-differentiated SFT data
A multi-task supervised fine-tuning dataset that teaches a model to convert
image-synthesis caption prompts into JSON whose structure varies by task.
Built from MS-COCO captions (Karpathy split) with Claude Sonnet 4.6 as the
teacher; designed for training per-task LoRAs on
Qwen/Qwen3.5-0.8B.
Each row is in the Qwen3.5-native tool-call shape: a messages array with an
assistant turn whose tool_calls[0].function.arguments is a dict… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/json-coco-format.alpaca_cleaned_ja_json
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/shi3z/alpaca_cleaned_ja_json.Ru_Tax_Audit_Instruct_Demo_JSON
Ru-Tax-Audit-Instruct: FNS Inspections, Fines & Corporate Compliance Scenarios (JSON)
🇷🇺 Описание проекта (Russian Description)
Ru-Tax-Audit-Instruct — это высококачественный коммерческий датасет инструктивного типа (Instruct Dataset), разработанный для обучения больших языковых моделей (LLM) логике российского корпоративного, налогового и трудового права.
Массив данных ориентирован на создание умных ИИ-ассистентов, роботов-консультантов, систем AI-комплаенса… See the full description on the dataset page: https://huggingface.co/datasets/Rudatamind/Ru_Tax_Audit_Instruct_Demo_JSON.svgen_500k_rasterized_jsonified_uuided
SVGEN RJU - SVGEN 500k: Rasterized, JSONified, UUID'ed
I have selected every svg image from svgen that would rasterize under cairosvg, which is significantly less than a 1% failure rate. Under development.
Reasoning
This is the 1st of many SVG datasets I am collecting, extracting, and rasterizing in an attempt to produce a meaningfully helpful spatial reasoning and vertex manipulation model.
Usage
The rasterized images are in PNG format, as bytes. They may be… See the full description on the dataset page: https://huggingface.co/datasets/MrOvkill/svgen_500k_rasterized_jsonified_uuided.listybox-etsy-listing-json
Etsy Product Listing Dataset (JSON SFT Format)
🎯 Optimized for Fine-tuning
Clean, minimal dataset for training vision-language models to generate Etsy product listings.
Key Features
JSON SFT Format: Ready for training with standard SFT trainers
Minimal Instructions: ~200 chars average (vs 2000+ in verbose versions)
Token Efficient: 90% reduction in instruction tokens
Production Ready: Model learns the task, not prompt engineering
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/burakaktna/listybox-etsy-listing-json.reasoning-sft-JSON-structuring-and-correcting
JSON Structuring and Correcting (Reasoning SFT)
Combined dataset of 508K rows for training LLMs on structured output tasks with reasoning traces, sourced from two datasets:
Sources
tool_calling.parquet (488,461 rows)
Converted from vericava/sft-tool-calling-structured-output-v1. Multi-turn tool calling and structured output tasks including tool invocations, tool results, and final assistant responses. Includes English and Japanese content.… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-JSON-structuring-and-correcting.json-schema-compliance-benchmark
JSON Schema Compliance Benchmark
A 500-example benchmark for evaluating whether language models can generate valid JSON conforming to provided schemas. Designed with strict contamination prevention to test generalization, not memorization.
Purpose
This is the primary Tier 1 evaluation metric for the Trellis SFT project. It measures a model's ability to produce structured output that passes jsonschema.validate() against novel, niche-domain schemas the model has never seen… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/json-schema-compliance-benchmark.100k_Tdk_zurriyet_dna_v6.jsonl
🌟 DESTEK & TOPLULUK ÇAĞRISI (SUPPORT & LIKE):Açık kaynak ve ücretsiz olarak sunduğum bu devasa çalışmayı faydalı bulduysanız, projenin sürdürülebilirliğine ve açık kaynak ekosisteminin görünürlüğüne katkı sağlamak için lütfen sayfanın sağ üstündeki Like (❤️ Beğeni) butonuna basarak destek olmayı unutmayın!(If you find this open-source dataset valuable for your research or models, please consider leaving a ❤️ Like at the top-right to support future updates and maintenance).
🇹🇷… See the full description on the dataset page: https://huggingface.co/datasets/bysismo/100k_Tdk_zurriyet_dna_v6.jsonl.qa-ml-dl-jsonl
💡 AI Q&A Dataset for ML, DL, RL, TensorFlow, PyTorch
This dataset is designed to support training and evaluation of AI systems on question generation, answering, and understanding in the domains of Machine Learning, Deep Learning, Reinforcement Learning, TensorFlow, and PyTorch. It contains a large number of categorized questions along with high-quality answers in two different levels of brevity.
📁 Dataset Files
1. questions.jsonl
Lines: 24,510… See the full description on the dataset page: https://huggingface.co/datasets/Koushim/qa-ml-dl-jsonl.
