Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01albertvillanova /tests-raw-jsonltext10K<n<100K1 likes31k downloads5y agoHugging Face02hf-internal-testing /raw_jsonltext10K<n<100K0 likes28k downloads5y agoHugging Face03permutans /wdc-common-crawl-embedded-jsonldtext10B<n<100B4 likes16k downloads2y agoHugging Face04hf-internal-testing /ner-jsonltext10K<n<100K0 likes7.3k downloads1y agoHugging Face05kaczmarj /wsinfer-model-zoo-jsonThis is the registry of models in the WSInfer Model Zoo. See https://wsinfer.readthedocs.io/en/latest/ and https://github.com/SBU-BMI/wsinfer-zoo for more information. textn<1K1 likes4.9k downloads3y agoHugging Face06chupei /format-jsonltextn<1K0 likes4.6k downloads2y agoHugging Face07chupei /format-jsontextn<1K0 likes4.6k downloads2y agoHugging Face08datasets-examples /doc-formats-jsonl-1 [doc] formats - jsonl - 1 This dataset contains one jsonl file at the root. textn<1K0 likes3.9k downloads3y agoHugging Face09epfl-dlab /JSONSchemaBench JSONSchemaBench JSONSchemaBench is a benchmark of real-world JSON schemas designed to evaluate structured output generation for Large Language Models (LLMs). It contains approximately 10,000 JSON schemas, capturing diverse constraints and complexities. import datasets from datasets import load_dataset def main(): # Inspect the available subsets of the datasetall_subsets = datasets.get_dataset_config_names("epfl-dlab/JSONSchemaBench") print("Available subsets:"… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/JSONSchemaBench.texttext-generation10K<n<100K12 likes3.5k downloads2y agoHugging Face10flax-sentence-embeddings /stackexchange_title_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M9 likes1.2k downloads4y agoHugging Face11flax-sentence-embeddings /stackexchange_titlebody_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M11 likes1.2k downloads4y agoHugging Face12scilons /SciLaD-all-json-v1 SciLaD (JSON) SciLaD is a novel, large-scale dataset of scientific language constructed entirely using open-source frameworks and publicly available data sources. It comprises a curated English split containing over 10 million scientific publications and a multilingual, unfiltered TEI XML split including more than 35 million publications. We also publish the extensible pipeline for generating SciLaD. Dataset Details In this repository we share the full… See the full description on the dataset page: https://huggingface.co/datasets/scilons/SciLaD-all-json-v1.text10M<n<100M0 likes1.1k downloads2mo agoHugging Face13flax-sentence-embeddings /stackexchange_title_body_jsonljsonl.gz format from https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_xml Each line contains a dict in the format: {"text": ["title", "body"], "tags": ["tag1", "tag2"]} The following parameters have been used for filtering: min_title_len = 20 min_body_len = 20 max_body_len = 4096 min_score = 0 If a stackexchange contained less than 10k questions (after filtering), it is written to the small_stackexchanges.jsonl.gz file. This is a dump of the files from… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_title_body_jsonl.text1M<n<10M2 likes1k downloads5y agoHugging Face14Arun63 /sharegpt-quizz-generation-json-output ShareGPT-Formatted Dataset for Quizz Generation in Structured JSON Output Dataset Description This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate quizz in structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios. Usage This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-quizz-generation-json-output.texttext-generationn<1K1 likes912 downloads2y agoHugging Face15flax-sentence-embeddings /stackexchange_titlebody_best_and_down_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering100K<n<1M13 likes783 downloads4y agoHugging Face16Arun63 /sharegpt-structured-output-json ShareGPT-Formatted Dataset for Structured JSON Output Dataset Description This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios. Usage This dataset can be used to train LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-structured-output-json.texttext-generationn<1K7 likes686 downloads2y agoHugging Face17NousResearch /json-mode-evaltextn<1K44 likes667 downloads3y agoHugging Face18flax-sentence-embeddings /stackexchange_math_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M20 likes531 downloads4y agoHugging Face19dixantp /desktop-accessibility-screenshot-json-dumpsimagen<1K1 likes475 downloads1y agoHugging Face20picollect /danbooru_jsonimage1M<n<10M1 likes458 downloads2y agoHugging Face21aitf-its-tim3-dfk /aitf-dfk3-vlm-dataset-jsonlimage10K<n<100K0 likes436 downloads4mo agoHugging Face22isaiahbjork /web-ui-grounding-jsontext100K<n<1M1 likes417 downloads2y agoHugging Face23Obscure-Entropy /conceptual_captions_jsonimage1M<n<10M0 likes392 downloads2y agoHugging Face24wangxiangyu0814 /TravelUAV_data_jsontext10K<n<100K0 likes362 downloads2y agoHugging Face25sraivante /home-commands-json-v1 Home Commands JSON v1 English and Romanized Hindi/Hinglish home commands mapped to structured JSON. This release contains the exact source and processed snapshot associated with Superfast Tiny Home Robotics JSON 1M v1. It also includes a separately authored evaluation challenge set. Native-script Hindi appears only in a small later diagnostic, not as a supported training language claim. Publisher: sraivante. Release date: 2026-09-25. Version: v1.0.2. Configurations… See the full description on the dataset page: https://huggingface.co/datasets/sraivante/home-commands-json-v1.texttext-generation1M<n<10M0 likes339 downloads14d agoHugging Face26anilkeshwani /jsonl-mls-hubert_large_ll60k-layer_22tabular1M<n<10M0 likes295 downloads1y agoHugging Face27wirthal1990-tech /USDA-Phytochemical-Database-JSON Ethno-API v2.4.0 — Public Sample Hugging Face hosts a 400-row public sample of Ethno-API v2.4.0: a cleaned and enriched phytochemical data-engineering project derived from the USDA Dr. Duke source data.The full project contains 76,907 records, 2,313 plant species, 24,746 unique chemical entities, and a 16-field public schema with PubMed, ClinicalTrials.gov, ChEMBL, PatentsView, PubChem CID/SMILES, and partner-assisted CID/IUPAC resolution fields.QA-gated public dataset… See the full description on the dataset page: https://huggingface.co/datasets/wirthal1990-tech/USDA-Phytochemical-Database-JSON.tabulartext-retrievaln<1K1 likes293 downloads2mo agoHugging Face28MedinaArmando /cfr-rag-jsonECFR from 06/2025 texttext-generation100K<n<1M0 likes291 downloads1y agoHugging Face29terminusresearch /pseudo-camera-10k-structured-json pseudo-camera-10k, structured JSON captions The 9,997 training images from bghira/pseudo-camera-10k, recaptioned into the structured JSON caption schema that Ideogram 4 consumes. The images are unchanged: free photographs from world class photographers, Lanczos-resized so the shorter edge is 1024px, nothing upsampled. The original dataset carries short CogVLM prose captions. This one replaces them with one JSON object per image describing the scene at three levels: an overall… See the full description on the dataset page: https://huggingface.co/datasets/terminusresearch/pseudo-camera-10k-structured-json.imagetext-to-image10K<n<100K0 likes291 downloads1mo agoHugging Face30AscendKernelGen /Ascend-COT-v2-json AscendKernelGen/Ascend-COT-v2-json AscendKernelGen/Ascend-CoT-v2-json contains a subset of the full Ascend-CoT dataset, which will be released in stages. The Ascend-CoT Dataset is a high-quality, domain-specific dataset that incorporates Chain-of-Thought (CoT) reasoning derived from real-world kernel implementations. It combines three types of reasoning: documentation-based reasoning, code-centric reasoning extracted from actual NPU kernel code, and general reasoning chains that… See the full description on the dataset page: https://huggingface.co/datasets/AscendKernelGen/Ascend-COT-v2-json.texttext-generation10K<n<100K3 likes287 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.