Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01albertvillanova /tests-raw-jsonltext10K<n<100K1 likes31k downloads5y agoHugging Face02hf-internal-testing /raw_jsonltext10K<n<100K0 likes28k downloads5y agoHugging Face03permutans /wdc-common-crawl-embedded-jsonldtext10B<n<100B4 likes16k downloads2y agoHugging Face04endomorphosis /Caselaw_Access_Project_JSON The Caselaw Access Project In collaboration with Ravel Law, Harvard Law Library digitized over 40 million U.S. court decisions consisting of 6.7 million cases from the last 360 years into a dataset that is widely accessible to use. Access a bulk download of the data through the Caselaw Access Project API (CAPAPI): https://case.law/caselaw/ Find more information about accessing state and federal written court decisions of common law through the bulk data service documentation here:… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/Caselaw_Access_Project_JSON.text-generation1M<n<10M3 likes12k downloads2y agoHugging Face05hf-internal-testing /ner-jsonltext10K<n<100K0 likes7.3k downloads1y agoHugging Face06introspector /ocaml-opam-ppxlib-json-astversion https://git-lfs.github.com/spec/v1 oid sha256:b797e216eb8d6720d794ca5d07a01019cac2d79df456fbcc69b6407497cc267c size 473 0 likes6k downloads2y agoHugging Face07kaczmarj /wsinfer-model-zoo-jsonThis is the registry of models in the WSInfer Model Zoo. See https://wsinfer.readthedocs.io/en/latest/ and https://github.com/SBU-BMI/wsinfer-zoo for more information. textn<1K1 likes4.9k downloads3y agoHugging Face08chupei /format-jsonltextn<1K0 likes4.6k downloads2y agoHugging Face09chupei /format-jsontextn<1K0 likes4.6k downloads2y agoHugging Face10datasets-examples /doc-formats-jsonl-1 [doc] formats - jsonl - 1 This dataset contains one jsonl file at the root. textn<1K0 likes3.9k downloads3y agoHugging Face11epfl-dlab /JSONSchemaBench JSONSchemaBench JSONSchemaBench is a benchmark of real-world JSON schemas designed to evaluate structured output generation for Large Language Models (LLMs). It contains approximately 10,000 JSON schemas, capturing diverse constraints and complexities. import datasets from datasets import load_dataset def main(): # Inspect the available subsets of the datasetall_subsets = datasets.get_dataset_config_names("epfl-dlab/JSONSchemaBench") print("Available subsets:"… See the full description on the dataset page: https://huggingface.co/datasets/epfl-dlab/JSONSchemaBench.texttext-generation10K<n<100K12 likes3.5k downloads2y agoHugging Face12jsonhash /LLVIP LLVIP 数据集 [中文] [English] 这里存储了[LLVIP 数据集]的备份 和 其 [COCO标注格式的标注] 下载 数据集:https://huggingface.co/datasets/UserNae3/LLVIP/blob/main/LLVIP.zip COCO格式标注:https://huggingface.co/datasets/UserNae3/LLVIP/blob/main/coco_annotations.7z 版权 版权链接: https://github.com/bupt-ai-cz/LLVIP?tab=readme-ov-file#license 1 likes2.6k downloads2y agoHugging Face13ChatCRG /Federal-Register-2016-2025-JSON1K<n<10K0 likes1.4k downloads6mo agoHugging Face14flax-sentence-embeddings /stackexchange_title_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M9 likes1.2k downloads4y agoHugging Face15flax-sentence-embeddings /stackexchange_titlebody_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M11 likes1.2k downloads4y agoHugging Face16scilons /SciLaD-all-json-v1 SciLaD (JSON) SciLaD is a novel, large-scale dataset of scientific language constructed entirely using open-source frameworks and publicly available data sources. It comprises a curated English split containing over 10 million scientific publications and a multilingual, unfiltered TEI XML split including more than 35 million publications. We also publish the extensible pipeline for generating SciLaD. Dataset Details In this repository we share the full… See the full description on the dataset page: https://huggingface.co/datasets/scilons/SciLaD-all-json-v1.text10M<n<100M0 likes1.1k downloads2mo agoHugging Face17flax-sentence-embeddings /stackexchange_title_body_jsonljsonl.gz format from https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_xml Each line contains a dict in the format: {"text": ["title", "body"], "tags": ["tag1", "tag2"]} The following parameters have been used for filtering: min_title_len = 20 min_body_len = 20 max_body_len = 4096 min_score = 0 If a stackexchange contained less than 10k questions (after filtering), it is written to the small_stackexchanges.jsonl.gz file. This is a dump of the files from… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_title_body_jsonl.text1M<n<10M2 likes1k downloads5y agoHugging Face18Arun63 /sharegpt-quizz-generation-json-output ShareGPT-Formatted Dataset for Quizz Generation in Structured JSON Output Dataset Description This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate quizz in structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios. Usage This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-quizz-generation-json-output.texttext-generationn<1K1 likes912 downloads2y agoHugging Face19flax-sentence-embeddings /stackexchange_titlebody_best_and_down_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering100K<n<1M13 likes783 downloads4y agoHugging Face20Arun63 /sharegpt-structured-output-json ShareGPT-Formatted Dataset for Structured JSON Output Dataset Description This dataset is formatted in the ShareGPT style and is designed for fine-tuning large language models (LLMs) to generate structured JSON outputs. It consists of multi-turn conversations where each response follows a predefined JSON schema, making it ideal for training models that need to produce structured data in natural language scenarios. Usage This dataset can be used to train LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Arun63/sharegpt-structured-output-json.texttext-generationn<1K7 likes686 downloads2y agoHugging Face21NousResearch /json-mode-evaltextn<1K44 likes667 downloads3y agoHugging Face22flax-sentence-embeddings /stackexchange_math_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M20 likes531 downloads4y agoHugging Face23SNUMPR /realfred_jsonDataset used in the paper ReALFRED: An Embodied Instruction Following Benchmark in Photo-Realistic Environments. 1 likes524 downloads2y agoHugging Face24jsonhash /FLIR_aligned 对齐的 FLIR 数据集 [中文] [English] FLIR 数据集由 FLIR 公司发布,双光谱目标检测数据集一般是使用来自 Zhang等人的修改的版本Paper, Dataset 这里存储了[张等人对齐的FLIR数据集]的备份 和 其 [转换成COCO标注格式的数据集] 下载 【推荐】转换成的COCO标注格式:https://huggingface.co/datasets/UserNae3/FLIR_aligned/resolve/main/flir_align.7z?download=trueZhang等人原始发布下载:https://drive.google.com/file/d/1xHDMGl6HJZwtarNWkEV3T4O9X4ZQYz2Y/view 对Zhang等人发布的备份:https://huggingface.co/datasets/UserNae3/FLIR_aligned/resolve/main/aligned.zip?download=true 版权 版权链接:… See the full description on the dataset page: https://huggingface.co/datasets/jsonhash/FLIR_aligned.4 likes489 downloads2y agoHugging Face25dixantp /desktop-accessibility-screenshot-json-dumpsimagen<1K1 likes475 downloads1y agoHugging Face26picollect /danbooru_jsonimage1M<n<10M1 likes458 downloads2y agoHugging Face27lenadan /otel-test-snippet-jsonl ⚠️ TEST DATASET - DO NOT USE FOR PRODUCTION This is a small test snippet for internal validation purposes only. This dataset contains a subset of OpenTelemetry traces from various LLM inference benchmarks. It is intended for testing dataset infrastructure and should NOT be used for research, benchmarking, or production purposes. Dataset Structure The dataset contains OpenTelemetry traces organized by: Benchmark: appworld, tau2_telecom Agent Framework: openai_solo… See the full description on the dataset page: https://huggingface.co/datasets/lenadan/otel-test-snippet-jsonl.text-generationn<1K0 likes458 downloads5mo agoHugging Face28aitf-its-tim3-dfk /aitf-dfk3-vlm-dataset-jsonlimage10K<n<100K0 likes436 downloads4mo agoHugging Face29isaiahbjork /web-ui-grounding-jsontext100K<n<1M1 likes417 downloads2y agoHugging Face30cpral /step35-en2pl-conv-pass4-jsonlconversations: 1,251,034 chat-template tokens (role+content, incl. special tokens): 2,664,206,408 reasoning_content tokens (not covered by chat template, counted separately): 6,662,763,429 avg tokens/conversation: 2129.6 used tokenizer: APT4 100K<n<1M0 likes411 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.