Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01albertvillanova /tests-raw-jsonltext10K<n<100K1 likes31k downloads5y agoHugging Face02hf-internal-testing /raw_jsonltext10K<n<100K0 likes28k downloads5y agoHugging Face03hf-internal-testing /ner-jsonltext10K<n<100K0 likes7.3k downloads1y agoHugging Face04kaczmarj /wsinfer-model-zoo-jsonThis is the registry of models in the WSInfer Model Zoo. See https://wsinfer.readthedocs.io/en/latest/ and https://github.com/SBU-BMI/wsinfer-zoo for more information. textn<1K1 likes4.9k downloads3y agoHugging Face05chupei /format-jsonltextn<1K0 likes4.6k downloads2y agoHugging Face06chupei /format-jsontextn<1K0 likes4.6k downloads2y agoHugging Face07datasets-examples /doc-formats-jsonl-1 [doc] formats - jsonl - 1 This dataset contains one jsonl file at the root. textn<1K0 likes3.9k downloads3y agoHugging Face08flax-sentence-embeddings /stackexchange_title_body_jsonljsonl.gz format from https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_xml Each line contains a dict in the format: {"text": ["title", "body"], "tags": ["tag1", "tag2"]} The following parameters have been used for filtering: min_title_len = 20 min_body_len = 20 max_body_len = 4096 min_score = 0 If a stackexchange contained less than 10k questions (after filtering), it is written to the small_stackexchanges.jsonl.gz file. This is a dump of the files from… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_title_body_jsonl.text1M<n<10M2 likes1k downloads5y agoHugging Face09dixantp /desktop-accessibility-screenshot-json-dumpsimagen<1K1 likes475 downloads1y agoHugging Face10picollect /danbooru_jsonimage1M<n<10M1 likes458 downloads2y agoHugging Face11aitf-its-tim3-dfk /aitf-dfk3-vlm-dataset-jsonlimage10K<n<100K0 likes436 downloads4mo agoHugging Face12cpral /step35-en2pl-conv-pass4-jsonlconversations: 1,251,034 chat-template tokens (role+content, incl. special tokens): 2,664,206,408 reasoning_content tokens (not covered by chat template, counted separately): 6,662,763,429 avg tokens/conversation: 2129.6 used tokenizer: APT4 100K<n<1M0 likes411 downloads3mo agoHugging Face13Obscure-Entropy /conceptual_captions_jsonimage1M<n<10M0 likes392 downloads2y agoHugging Face14wangxiangyu0814 /TravelUAV_data_jsontext10K<n<100K0 likes362 downloads2y agoHugging Face15sraivante /home-commands-json-v1 Home Commands JSON v1 English and Romanized Hindi/Hinglish home commands mapped to structured JSON. This release contains the exact source and processed snapshot associated with Superfast Tiny Home Robotics JSON 1M v1. It also includes a separately authored evaluation challenge set. Native-script Hindi appears only in a small later diagnostic, not as a supported training language claim. Publisher: sraivante. Release date: 2026-09-25. Version: v1.0.2. Configurations… See the full description on the dataset page: https://huggingface.co/datasets/sraivante/home-commands-json-v1.texttext-generation1M<n<10M0 likes339 downloads14d agoHugging Face16anilkeshwani /jsonl-mls-hubert_large_ll60k-layer_22tabular1M<n<10M0 likes295 downloads1y agoHugging Face17MedinaArmando /cfr-rag-jsonECFR from 06/2025 texttext-generation100K<n<1M0 likes291 downloads1y agoHugging Face18cpral /poziomka-fun-v9-jsonltext100K<n<1M0 likes263 downloads1mo agoHugging Face19LeonOverload /primo-rl-json PRIMO RL Data Stage-2 (GRPO reinforcement learning) training annotations for PRIMO R1 (paper). Unlike the SFT data, these records carry no chain-of-thought traces — RL optimizes against a verifiable progress reward, so only the ground-truth answer is needed. That is the point of the method: the model discovers its own reasoning instead of imitating someone else's. 328,454 records across 6 subsets. Annotations only (1.7 GB); videos are in primo-video-media. Subsets… See the full description on the dataset page: https://huggingface.co/datasets/LeonOverload/primo-rl-json.textvideo-text-to-text100K<n<1M0 likes254 downloads1mo agoHugging Face20taoroalin /code_contests_slim_jsontabular1K<n<10K0 likes249 downloads2y agoHugging Face21cpral /step35-en2pl-conv-pass2-jsonl1M<n<10M0 likes227 downloads5mo agoHugging Face22happynew111 /MATH_BS_BCE_train_json1M<n<10M0 likes218 downloads1y agoHugging Face23dataunitylab /json-schema-storeThis contains a set of schemas obtained via the JSON Schema Store catalog. textn<1K2 likes212 downloads2y agoHugging Face24TruongSinhAI /deepcad_prompt_jsontext100K<n<1M0 likes202 downloads2y agoHugging Face25SimbaMaw1547 /south-african-monolingual-corpora-jsonl South African Languages Pretraining Dataset This dataset contains pretraining text data for 9 South African languages, compiled from multiple sources including CC100, Glot500, mC4, ParaCrawl, and various corpora collections. The datasets were gathered as part of the University of Cape Town's SALLM project. Where data was gathered from multiple sources, extensive filtering and deduplication was conducted to ensure dataset integrity Languages Included Language… See the full description on the dataset page: https://huggingface.co/datasets/SimbaMaw1547/south-african-monolingual-corpora-jsonl.text1M<n<10M0 likes201 downloads1y agoHugging Face26sraivante /home-commands-json-v3 Home Commands JSON v3 English and Hinglish (Romanised Hindi) home-automation instructions mapped to a single JSON device command, with optional conditional rules. This is the exact training mix behind Superfast Tiny Home Robotics JSON 1M v3, plus the generators that produced it. {"instruction": "agar room ka temperature 42 se upar jaye to bedroom ka heater band kar do", "output": "{\"activity\":\"heating\",\"subject\":\"bedroom_room_heater\",\"action\":\"OFF\"… See the full description on the dataset page: https://huggingface.co/datasets/sraivante/home-commands-json-v3.texttext-generation100K<n<1M0 likes186 downloads14d agoHugging Face27dataunitylab /json-schema JSON Schema Dataset This dataset consists of a collection of JSON Schema documents collected from GitHub by searching using the Sourcegraph API. Step 1: Find a list of JSON Schema paths The Sourcegraph code search API is used to find files with a .json extension and containing {\n "$schema": "https://json-schema.org/". This is somewhat restrictive, but still manages to find a large number of schemas. pipenv run python slurp.py --outfile repos.csv Step 2:… See the full description on the dataset page: https://huggingface.co/datasets/dataunitylab/json-schema.text10K<n<100K2 likes184 downloads2y agoHugging Face28ayousanz /midi-classical-music-toio-json MIDI Classical Music drengskapur/midi-classical-musicのデータセットをtoioの soundコマンドで再生しやすいように以下のフォーマットのjsonに変換したデータを含めたデータセット data format [ { "track_name": "ALBENIZ: Aragon Op 47/6", "priority": 1, "notes": [ { "note_number": 77, "start_time_ms": 0, "duration_units": 26 }, { }, }, { "track_name": "apurdam@pcug.org.au", "priority": 2, "notes": [ { "note_number": 53, "start_time_ms": 0… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/midi-classical-music-toio-json.text10K<n<100K2 likes176 downloads2y agoHugging Face29Emulated-Inc /json-schema-instances-training-pool JSON schema and instance training pool Real JSON Schemas from the public collections named below, read at the pinned revisions given there, each paired where possible with documents that satisfy it, laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one shape, 20004 rows, one JSON object per line, with these fields. Field What it holds id a row identifier unique within this file prompt the request a model would… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/json-schema-instances-training-pool.texttext-generation10K<n<100K1 likes173 downloads28d agoHugging Face30minpeter /hermes-function-calling-v1-jsonl Hermes Function-Calling V1 This dataset is the compilation of structured output and function calling data used in the Hermes 2 Pro series of models. This repository contains a structured output dataset with function-calling conversations, json-mode, agentic json-mode and structured extraction samples, designed to train LLM models in performing function calls and returning structured output based on natural language instructions. The dataset features various conversational scenarios… See the full description on the dataset page: https://huggingface.co/datasets/minpeter/hermes-function-calling-v1-jsonl.texttext-generation10K<n<100K1 likes168 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.