datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tests-raw-jsonlraw_jsonlner-jsonlwsinfer-model-zoo-jsonThis is the registry of models in the WSInfer Model Zoo.
See https://wsinfer.readthedocs.io/en/latest/ and https://github.com/SBU-BMI/wsinfer-zoo for more information.
format-jsonlformat-jsondoc-formats-jsonl-1
[doc] formats - jsonl - 1
This dataset contains one jsonl file at the root.
stackexchange_title_body_jsonljsonl.gz format from https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_xml
Each line contains a dict in the format: {"text": ["title", "body"], "tags": ["tag1", "tag2"]}
The following parameters have been used for filtering: min_title_len = 20 min_body_len = 20 max_body_len = 4096 min_score = 0
If a stackexchange contained less than 10k questions (after filtering), it is written to the small_stackexchanges.jsonl.gz file.
This is a dump of the files from… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_title_body_jsonl.desktop-accessibility-screenshot-json-dumpsdanbooru_jsonaitf-dfk3-vlm-dataset-jsonlstep35-en2pl-conv-pass4-jsonlconversations: 1,251,034
chat-template tokens (role+content, incl. special tokens): 2,664,206,408
reasoning_content tokens (not covered by chat template, counted separately): 6,662,763,429
avg tokens/conversation: 2129.6
used tokenizer: APT4
conceptual_captions_jsonTravelUAV_data_jsonhome-commands-json-v1
Home Commands JSON v1
English and Romanized Hindi/Hinglish home commands mapped to structured JSON.
This release contains the exact source and processed snapshot associated
with Superfast Tiny Home Robotics JSON 1M v1.
It also includes a separately authored evaluation challenge set. Native-script
Hindi appears only in a small later diagnostic, not as a supported training
language claim.
Publisher: sraivante. Release date: 2026-09-25. Version: v1.0.2.
Configurations… See the full description on the dataset page: https://huggingface.co/datasets/sraivante/home-commands-json-v1.jsonl-mls-hubert_large_ll60k-layer_22cfr-rag-jsonECFR from 06/2025
poziomka-fun-v9-jsonlprimo-rl-json
PRIMO RL Data
Stage-2 (GRPO reinforcement learning) training annotations for PRIMO R1 (paper). Unlike the SFT data, these records carry no chain-of-thought traces — RL optimizes against a verifiable progress reward, so only the ground-truth answer is needed. That is the point of the method: the model discovers its own reasoning instead of imitating someone else's.
328,454 records across 6 subsets. Annotations only (1.7 GB); videos are in primo-video-media.
Subsets… See the full description on the dataset page: https://huggingface.co/datasets/LeonOverload/primo-rl-json.code_contests_slim_jsonstep35-en2pl-conv-pass2-jsonlMATH_BS_BCE_train_jsonjson-schema-storeThis contains a set of schemas obtained via the JSON Schema Store catalog.
deepcad_prompt_jsonsouth-african-monolingual-corpora-jsonl
South African Languages Pretraining Dataset
This dataset contains pretraining text data for 9 South African languages, compiled from multiple sources including CC100, Glot500, mC4, ParaCrawl, and various corpora collections.
The datasets were gathered as part of the University of Cape Town's SALLM project. Where data was gathered from multiple sources, extensive filtering and deduplication was conducted to ensure dataset integrity
Languages Included
Language… See the full description on the dataset page: https://huggingface.co/datasets/SimbaMaw1547/south-african-monolingual-corpora-jsonl.home-commands-json-v3
Home Commands JSON v3
English and Hinglish (Romanised Hindi) home-automation instructions mapped to a
single JSON device command, with optional conditional rules. This is the
exact training mix behind
Superfast Tiny Home Robotics JSON 1M v3,
plus the generators that produced it.
{"instruction": "agar room ka temperature 42 se upar jaye to bedroom ka heater band kar do",
"output": "{\"activity\":\"heating\",\"subject\":\"bedroom_room_heater\",\"action\":\"OFF\"… See the full description on the dataset page: https://huggingface.co/datasets/sraivante/home-commands-json-v3.json-schema
JSON Schema Dataset
This dataset consists of a collection of JSON Schema documents collected from GitHub by searching using the Sourcegraph API.
Step 1: Find a list of JSON Schema paths
The Sourcegraph code search API is used to find files with a .json extension and containing {\n "$schema": "https://json-schema.org/".
This is somewhat restrictive, but still manages to find a large number of schemas.
pipenv run python slurp.py --outfile repos.csv
Step 2:… See the full description on the dataset page: https://huggingface.co/datasets/dataunitylab/json-schema.midi-classical-music-toio-json
MIDI Classical Music
drengskapur/midi-classical-musicのデータセットをtoioの soundコマンドで再生しやすいように以下のフォーマットのjsonに変換したデータを含めたデータセット
data format
[
{
"track_name": "ALBENIZ: Aragon Op 47/6",
"priority": 1,
"notes": [
{
"note_number": 77,
"start_time_ms": 0,
"duration_units": 26
},
{
},
},
{
"track_name": "apurdam@pcug.org.au",
"priority": 2,
"notes": [
{
"note_number": 53,
"start_time_ms": 0… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/midi-classical-music-toio-json.json-schema-instances-training-pool
JSON schema and instance training pool
Real JSON Schemas from the public collections named below, read at the pinned revisions given there,
each paired where possible with documents that satisfy it, laid out twice. Train on either layer or
on both.
pool.jsonl
Every source rewritten into one shape, 20004 rows, one JSON object per line, with these fields.
Field
What it holds
id
a row identifier unique within this file
prompt
the request a model would… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/json-schema-instances-training-pool.hermes-function-calling-v1-jsonl
Hermes Function-Calling V1
This dataset is the compilation of structured output and function calling data used in the Hermes 2 Pro series of models.
This repository contains a structured output dataset with function-calling conversations, json-mode, agentic json-mode and structured extraction samples, designed to train LLM models in performing function calls and returning structured output based on natural language instructions. The dataset features various conversational scenarios… See the full description on the dataset page: https://huggingface.co/datasets/minpeter/hermes-function-calling-v1-jsonl.
