Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01endomorphosis /Caselaw_Access_Project_JSON The Caselaw Access Project In collaboration with Ravel Law, Harvard Law Library digitized over 40 million U.S. court decisions consisting of 6.7 million cases from the last 360 years into a dataset that is widely accessible to use. Access a bulk download of the data through the Caselaw Access Project API (CAPAPI): https://case.law/caselaw/ Find more information about accessing state and federal written court decisions of common law through the bulk data service documentation here:… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/Caselaw_Access_Project_JSON.text-generation1M<n<10M3 likes12k downloads2y agoHugging Face02epfl-dlab /JSONSchemaBench JSONSchemaBench JSONSchemaBench is a benchmark of real-world JSON schemas designed to evaluate structured output generation for Large Language Models (LLMs). It contains approximately 10,000 JSON schemas, capturing diverse constraints and complexities. import datasets from datasets import load_dataset def main(): # Inspect the available subsets of the datasetall_subsets = datasets.get_dataset_config_names("epfl-dlab/JSONSchemaBench") print("Available subsets:"… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/JSONSchemaBench.texttext-generation10K<n<100K12 likes3.5k downloads2y agoHugging Face03Arun63 /sharegpt-quizz-generation-json-output ShareGPT-Formatted Dataset for Quizz Generation in Structured JSON Output Dataset Description This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate quizz in structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios. Usage This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-quizz-generation-json-output.texttext-generationn<1K1 likes912 downloads2y agoHugging Face04Arun63 /sharegpt-structured-output-json ShareGPT-Formatted Dataset for Structured JSON Output Dataset Description This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios. Usage This dataset can be used to train LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-structured-output-json.texttext-generationn<1K7 likes686 downloads2y agoHugging Face05lenadan /otel-test-snippet-jsonl ⚠️ TEST DATASET - DO NOT USE FOR PRODUCTION This is a small test snippet for internal validation purposes only. This dataset contains a subset of OpenTelemetry traces from various LLM inference benchmarks. It is intended for testing dataset infrastructure and should NOT be used for research, benchmarking, or production purposes. Dataset Structure The dataset contains OpenTelemetry traces organized by: Benchmark: appworld, tau2_telecom Agent Framework: openai_solo… See the full description on the dataset page: https://huggingface.co/datasets/lenadan/otel-test-snippet-jsonl.text-generationn<1K0 likes458 downloads5mo agoHugging Face06sraivante /home-commands-json-v1 Home Commands JSON v1 English and Romanized Hindi/Hinglish home commands mapped to structured JSON. This release contains the exact source and processed snapshot associated with Superfast Tiny Home Robotics JSON 1M v1. It also includes a separately authored evaluation challenge set. Native-script Hindi appears only in a small later diagnostic, not as a supported training language claim. Publisher: sraivante. Release date: 2026-09-25. Version: v1.0.2. Configurations… See the full description on the dataset page: https://huggingface.co/datasets/sraivante/home-commands-json-v1.texttext-generation1M<n<10M0 likes339 downloads14d agoHugging Face07MedinaArmando /cfr-rag-jsonECFR from 06/2025 texttext-generation100K<n<1M0 likes291 downloads1y agoHugging Face08AscendKernelGen /Ascend-COT-v2-json AscendKernelGen/Ascend-COT-v2-json AscendKernelGen/Ascend-CoT-v2-json contains a subset of the full Ascend-CoT dataset, which will be released in stages. The Ascend-CoT Dataset is a high-quality, domain-specific dataset that incorporates Chain-of-Thought (CoT) reasoning derived from real-world kernel implementations. It combines three types of reasoning: documentation-based reasoning, code-centric reasoning extracted from actual NPU kernel code, and general reasoning chains that… See the full description on the dataset page: https://huggingface.co/datasets/AscendKernelGen/Ascend-COT-v2-json.texttext-generation10K<n<100K3 likes287 downloads6mo agoHugging Face09AscendKernelGen /Ascend-CoT-v3-json Ascend-CoT-v3-json Ascend-CoT-v3-json is an Ascend C / CANN supervised fine-tuning dataset for custom operator development. It contains cleaned CoT-style samples for Ascend C kernel implementation, tiling logic, CANN API usage, debugging, and operator-development reasoning. The release is organized into two final SFT subsets in one dataset repository. Related Artifacts Paper: AscendKernelGen: A Systematic Study of LLM-Based Kernel Generation for Neural… See the full description on the dataset page: https://huggingface.co/datasets/AscendKernelGen/Ascend-CoT-v3-json.texttext-generation100K<n<1M2 likes205 downloads4mo agoHugging Face10A11Sunday /support-json-ru Support-JSON-RU Synthetic Russian SaaS support data for policy-conditioned JSON decisions and draft replies. The task supplies customer text, company policies, sourced facts and available capabilities; the model predicts a nine-field decision rather than memorizing a single company's policy. Русский SaaS-support: обращение + правила + факты → категория, приоритет, настроение, действие, черновик ответа и эскалация. Model · Dataset files · License Configurations… See the full description on the dataset page: https://huggingface.co/datasets/A11Sunday/support-json-ru.texttext-generation10K<n<100K1 likes199 downloads23d agoHugging Face11sraivante /home-commands-json-v3 Home Commands JSON v3 English and Hinglish (Romanised Hindi) home-automation instructions mapped to a single JSON device command, with optional conditional rules. This is the exact training mix behind Superfast Tiny Home Robotics JSON 1M v3, plus the generators that produced it. {"instruction": "agar room ka temperature 42 se upar jaye to bedroom ka heater band kar do", "output": "{\"activity\":\"heating\",\"subject\":\"bedroom_room_heater\",\"action\":\"OFF\"… See the full description on the dataset page: https://huggingface.co/datasets/sraivante/home-commands-json-v3.texttext-generation100K<n<1M0 likes186 downloads14d agoHugging Face12AbstractPhil /cc-task1-json CC captions → task_1 structured JSON Task 1: V1 Complete Total Time: 50 hours 3x 6000 blackwell pros Host: Google Colab Model: AbstractPhil/qwen3.5-0.8b-task_1-lora Task: Converting plain English image prompts to JSON containing similar assessments using subjective analysis. Biases: Guaranteed - Filter nulls before training anything. V1 Limitations Extracted from the qwen3.5-0.8b task_1 V1 lora. Context and associations limited Topic faults and invalid context… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/cc-task1-json.text-generation1M<n<10M0 likes184 downloads5mo agoHugging Face13paraloq /json_data_extraction Diverse Restricted JSON Data Extraction Curated by: The paraloq analytics team. Uses Benchmark restricted JSON data extraction (text + JSON schema -> JSON instance) Fine-Tune data extraction model (text + JSON schema -> JSON instance) Fine-Tune JSON schema Retrieval model (text -> retriever -> most adequate JSON schema) Out-of-Scope Use Intended for research purposes only. Dataset Structure The data comes with the following fields: title: The… See the full description on the dataset page: https://huggingface.co/datasets/paraloq/json_data_extraction.texttext-generationn<1K35 likes180 downloads3y agoHugging Face14Emulated-Inc /json-schema-instances-training-pool JSON schema and instance training pool Real JSON Schemas from the public collections named below, read at the pinned revisions given there, each paired where possible with documents that satisfy it, laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one shape, 20004 rows, one JSON object per line, with these fields. Field What it holds id a row identifier unique within this file prompt the request a model would… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/json-schema-instances-training-pool.texttext-generation10K<n<100K1 likes173 downloads28d agoHugging Face15minpeter /hermes-function-calling-v1-jsonl Hermes Function-Calling V1 This dataset is the compilation of structured output and function calling data used in the Hermes 2 Pro series of models. This repository contains a structured output dataset with function-calling conversations, json-mode, agentic json-mode and structured extraction samples, designed to train LLM models in performing function calls and returning structured output based on natural language instructions. The dataset features various conversational scenarios… See the full description on the dataset page: https://huggingface.co/datasets/minpeter/hermes-function-calling-v1-jsonl.texttext-generation10K<n<100K1 likes168 downloads2y agoHugging Face16sandeeppanem /resume-json-extraction-5k Dataset Card for resume-json-extraction-5k Dataset Description This dataset contains 4,879 resume examples formatted for fine-tuning language models to extract structured JSON information from resume text. Dataset Summary The dataset consists of resume text paired with structured JSON outputs containing: Job titles (current and previous) Companies (current and previous) Years of experience Seniority level Primary domain and industries Core and secondary skills… See the full description on the dataset page: https://huggingface.co/datasets/sandeeppanem/resume-json-extraction-5k.texttext-generation1K<n<10K1 likes135 downloads9mo agoHugging Face17NJUDeepEngine /bigbench_jsonlBIG-Bench but it doesn't require the hellish dependencies (tensorflow, pypi-bigbench, protobuf) of the official version. dataset = load_dataset("tasksource/bigbench",'movie_recommendation') Code to reproduce: https://colab.research.google.com/drive/1MKdLdF7oqrSQCeavAcsEnPdI85kD0LzU?usp=sharing Datasets are capped to 50k examples to keep things light. I also removed the default split when train was available also to save space, as default=train+val. @article{srivastava2022beyond… See the full description on the dataset page: https://huggingface.co/datasets/NJUDeepEngine/bigbench_jsonl.textmultiple-choice100K<n<1M1 likes123 downloads1y agoHugging Face18robdixon /json-extraction Rob Dixon's JSON Extraction Dataset A synthetic dataset for training JSON extraction models, generated using Claude 3 Haiku. Dataset Overview This dataset contains paired examples of: Instructions: Natural language task descriptions asking to extract information Text documents: Source content containing information to extract JSON outputs: Structured data extracted from the text The dataset is designed for training smaller models on constrained context lengths, with… See the full description on the dataset page: https://huggingface.co/datasets/robdixon/json-extraction.texttext-generation10K<n<100K2 likes109 downloads8mo agoHugging Face19Travis-ML /ShortStory-SFT-jsonl Public Domain Short Fiction with Prompts 719 complete short stories (300 to 2500 words) by 26 authors whose work is in the public domain, each paired with a natural-language request that could plausibly have produced it. Built for supervised fine-tuning of small language models on fiction, where the usual sources (forum stories, model-generated stories) lack the structural control of published short fiction. Fields field description id stable id (hash… See the full description on the dataset page: https://huggingface.co/datasets/Travis-ML/ShortStory-SFT-jsonl.texttext-generationn<1K0 likes101 downloads20d agoHugging Face20AI-Culture-Commons /ai-culture-multilingual-json-dolma AI-Culture Multilingual JSON + DOLMA Corpus 16M words · 12 languages · CC-BY-4.0 The AI-Culture corpus contains 5K articles providing comprehensive philosophical and cultural content, exploring the intersection of technology, artificial intelligence, and human culture, perfectly aligned across 12 languages. All content maintains identical parallel structure across translations with zero duplication and editor-curated quality. This project is maintained by a non-profit digital… See the full description on the dataset page: https://huggingface.co/datasets/AI-Culture-Commons/ai-culture-multilingual-json-dolma.texttranslation1K<n<10K3 likes100 downloads1y agoHugging Face21mrcuddle /NSFW-Stories-JsonLConverted to JsonL from: bluuwhale/nsfwstory2 texttext-generation10K<n<100K42 likes92 downloads2y agoHugging Face22AbstractPhil /json-coco-format JSON COCO Format — task-differentiated SFT data A multi-task supervised fine-tuning dataset that teaches a model to convert image-synthesis caption prompts into JSON whose structure varies by task. Built from MS-COCO captions (Karpathy split) with Claude Sonnet 4.6 as the teacher; designed for training per-task LoRAs on Qwen/Qwen3.5-0.8B. Each row is in the Qwen3.5-native tool-call shape: a messages array with an assistant turn whose tool_calls[0].function.arguments is a dict… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/json-coco-format.texttext-generation100K<n<1M0 likes82 downloads5mo agoHugging Face23shi3z /alpaca_cleaned_ja_json Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/shi3z/alpaca_cleaned_ja_json.texttext-generation100K<n<1M13 likes74 downloads3y agoHugging Face24Rudatamind /Ru_Tax_Audit_Instruct_Demo_JSON Ru-Tax-Audit-Instruct: FNS Inspections, Fines & Corporate Compliance Scenarios (JSON) 🇷🇺 Описание проекта (Russian Description) Ru-Tax-Audit-Instruct — это высококачественный коммерческий датасет инструктивного типа (Instruct Dataset), разработанный для обучения больших языковых моделей (LLM) логике российского корпоративного, налогового и трудового права. Массив данных ориентирован на создание умных ИИ-ассистентов, роботов-консультантов, систем AI-комплаенса… See the full description on the dataset page: https://huggingface.co/datasets/Rudatamind/Ru_Tax_Audit_Instruct_Demo_JSON.text-generationn<1K0 likes62 downloads24d agoHugging Face25MrOvkill /svgen_500k_rasterized_jsonified_uuided SVGEN RJU - SVGEN 500k: Rasterized, JSONified, UUID'ed I have selected every svg image from svgen that would rasterize under cairosvg, which is significantly less than a 1% failure rate. Under development. Reasoning This is the 1st of many SVG datasets I am collecting, extracting, and rasterizing in an attempt to produce a meaningfully helpful spatial reasoning and vertex manipulation model. Usage The rasterized images are in PNG format, as bytes. They may be… See the full description on the dataset page: https://huggingface.co/datasets/MrOvkill/svgen_500k_rasterized_jsonified_uuided.texttext-generation100K<n<1M1 likes55 downloads2y agoHugging Face26burakaktna /listybox-etsy-listing-json Etsy Product Listing Dataset (JSON SFT Format) 🎯 Optimized for Fine-tuning Clean, minimal dataset for training vision-language models to generate Etsy product listings. Key Features JSON SFT Format: Ready for training with standard SFT trainers Minimal Instructions: ~200 chars average (vs 2000+ in verbose versions) Token Efficient: 90% reduction in instruction tokens Production Ready: Model learns the task, not prompt engineering Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/burakaktna/listybox-etsy-listing-json.textimage-to-text1K<n<10K0 likes54 downloads1y agoHugging Face27AmanPriyanshu /reasoning-sft-JSON-structuring-and-correcting JSON Structuring and Correcting (Reasoning SFT) Combined dataset of 508K rows for training LLMs on structured output tasks with reasoning traces, sourced from two datasets: Sources tool_calling.parquet (488,461 rows) Converted from vericava/sft-tool-calling-structured-output-v1. Multi-turn tool calling and structured output tasks including tool invocations, tool results, and final assistant responses. Includes English and Japanese content.… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-JSON-structuring-and-correcting.texttext-generation100K<n<1M1 likes52 downloads7mo agoHugging Face28mdonigian /json-schema-compliance-benchmark JSON Schema Compliance Benchmark A 500-example benchmark for evaluating whether language models can generate valid JSON conforming to provided schemas. Designed with strict contamination prevention to test generalization, not memorization. Purpose This is the primary Tier 1 evaluation metric for the Trellis SFT project. It measures a model's ability to produce structured output that passes jsonschema.validate() against novel, niche-domain schemas the model has never seen… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/json-schema-compliance-benchmark.text-generationn<1K0 likes52 downloads7mo agoHugging Face29bysismo /100k_Tdk_zurriyet_dna_v6.jsonl 🌟 DESTEK & TOPLULUK ÇAĞRISI (SUPPORT & LIKE):Açık kaynak ve ücretsiz olarak sunduğum bu devasa çalışmayı faydalı bulduysanız, projenin sürdürülebilirliğine ve açık kaynak ekosisteminin görünürlüğüne katkı sağlamak için lütfen sayfanın sağ üstündeki Like (❤️ Beğeni) butonuna basarak destek olmayı unutmayın!(If you find this open-source dataset valuable for your research or models, please consider leaving a ❤️ Like at the top-right to support future updates and maintenance). 🇹🇷… See the full description on the dataset page: https://huggingface.co/datasets/bysismo/100k_Tdk_zurriyet_dna_v6.jsonl.textquestion-answering10K<n<100K1 likes50 downloads2mo agoHugging Face30Koushim /qa-ml-dl-jsonl 💡 AI Q&A Dataset for ML, DL, RL, TensorFlow, PyTorch This dataset is designed to support training and evaluation of AI systems on question generation, answering, and understanding in the domains of Machine Learning, Deep Learning, Reinforcement Learning, TensorFlow, and PyTorch. It contains a large number of categorized questions along with high-quality answers in two different levels of brevity. 📁 Dataset Files 1. questions.jsonl Lines: 24,510… See the full description on the dataset page: https://huggingface.co/datasets/Koushim/qa-ml-dl-jsonl.question-answering0 likes49 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.