datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TinyStoriesDataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary.
Described in the following paper: https://arxiv.org/abs/2305.07759.
The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M.
Additional resources:
tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/roneneldan/TinyStories.tiny-supervised-datasettiny-shakespeare
Data source
Downloaded via Andrej Karpathy's nanogpt repo from this link
Data Format
The entire dataset is split into train (90%) and test (10%).
All rows are at most 1024 tokens, using the Llama 2 tokenizer.
All rows are split cleanly so that sentences are whole and unbroken.
tiny-shakespeareesb-datasets-earnings22-validation-tiny-filteredA filtered (<=30s duration) slice (512 samples) of the Earnings22 dataset.
def add_duration(sample):
y, sr = sample['audio']["array"], sample['audio']["sampling_rate"]
sample['duration_ms']=librosa.get_duration(y=y, sr=sr) * 1000
return sample
tedlium = load_dataset("esb/datasets", "earnings22", split='validation', trust_remote_code=True)
# compute duration to filter
tedlium = tedlium.map(add_duration)
tedlium = tedlium.select(range(512))
# Whisper max supported duration
tedlium… See the full description on the dataset page: https://huggingface.co/datasets/D4nt3/esb-datasets-earnings22-validation-tiny-filtered.tiny-scalesThis repo contains the tinyHLE dataset, a list of items to use as a subset of the Humanity's Last Exam benchmark in order to make evaluation more efficient.
The repo contains two files:
tiny_hle.json: a file containing a list of question IDs and weights for three different sample sizes (0.5%, 1.0%, 2.0%)
clean_scales_embedding_hle.parquet: a file containing embeddings representing each item of the HLE benchmark along 16 cognitive scales dimensions, used to create the subsets
Since these are… See the full description on the dataset page: https://huggingface.co/datasets/ambean-tr/tiny-scales.tinyGSM8k
tinyGSM8K
Welcome to tinyGSM8K! This dataset serves as a concise version of the GSM8K dataset, offering a subset of 100 data points selected from the original compilation.
tinyGSM8K is designed to enable users to efficiently estimate the performance of a large language model (LLM) with reduced dataset size, saving computational resources
while maintaining the essence of the GSM8K evaluation.
Features
Compact Dataset: With only 100 data points, tinyGSM8K… See the full description on the dataset page: https://huggingface.co/datasets/tinyBenchmarks/tinyGSM8k.tinyMMLU
tinyMMLU
Welcome to tinyMMLU! This dataset serves as a concise version of the MMLU dataset, offering a subset of 100 data points selected from the original compilation.
tinyMMLU is designed to enable users to efficiently estimate the performance of a large language model (LLM) with reduced dataset size, saving computational resources
while maintaining the essence of the MMLU evaluation.
Features
Compact Dataset: With only 100 data points, tinyMMLU provides a swift… See the full description on the dataset page: https://huggingface.co/datasets/tinyBenchmarks/tinyMMLU.trivia_qa_tiny
Dataset Card for Dataset Name
Dataset Summary
This dataset contains 100 samples from trivia_qa dataset. It is used mainly for testing purposes.
Languages
English.
Dataset Structure
Data Instances
Total data size: 8Kb.
Data Fields
question: string feature, containing question to be answered.
`answer: string feature, answer to the question.
Data Splits
Only test split, that contains 100 rows, is supported.
Tiny_Theory_of_Mind
Tiny Theory of Mind
Tiny Theory of Mind is our first attempt at evaluating theory of mind capabilities in small language models. The benchmark covers a wide variety of ToM topics, ranging in difficulties that, for humans, would be appropriate for Pre-K through 6th grade.
The benchmark is designed primarily for base-model continuation log-likelihood scoring. It does not require instruction following, chain-of-thought, or generated explanations. Random-choice accuracy is 25%.… See the full description on the dataset page: https://huggingface.co/datasets/AxiomicLabs/Tiny_Theory_of_Mind.product-photography-v1-tiny-prompts-tasks-collage-filtereddanish-asr-unified-hviske-v5-tiny
danish-asr-unified — two-model labels and a quality manifest
Transcriptions, per-token confidences, and a per-row quality verdict for every
row of syvai/danish-asr-unified
(3,414,589 rows, 8 sources).
Two independently trained models labelled the whole corpus:
model
architecture
vocabulary
syvai/hviske-v5-tiny
encoder-decoder
16,384 BPE
3dio-ai/svale-110M
RNN-T (Parakeet)
44 characters
Both decode greedily. Shards mirror the source data/train-*.parquet by name… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified-hviske-v5-tiny.tinyHellaswag
tinyHellaswag
Welcome to tinyHellaswag! This dataset serves as a concise version of the hellaswag dataset, offering a subset of 100 data points selected from the original compilation.
tinyHellaswag is designed to enable users to efficiently estimate the performance of a large language model (LLM) with reduced dataset size, saving computational resources
while maintaining the essence of the hellaswag evaluation.
Features
Compact Dataset: With only 100 data points… See the full description on the dataset page: https://huggingface.co/datasets/tinyBenchmarks/tinyHellaswag.tiny-aya-l2-thinker-multilingual-reasoning
Tiny Aya L2 Multilingual Reasoning (44 languages)
Translated multilingual reasoning traces used to train Tiny Aya L2-Thinker.
Each example has the prompt, thinking, and answer in the same non-English language alongside the original texts in English.
Data source
Prompts from AM-DeepSeek-R1-0528-Distilled
Thinking traces and outputs distilled from gpt-oss-120b
Translated with command-a-translate and DeepSeek-V3
Languages (44)
Language
Train… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/tiny-aya-l2-thinker-multilingual-reasoning.tiny-instruct
tiny-instruct-v1
This dataset is collated from multiple other open-source datasets (de-duplicated). This has a total of ~6M rows each with an instruction and response (single-turn converstion).
Code Datasets:
CodeAlpaca_20K
CodeExercise-Python-27k
Evol-Instruct-Code-80k-v1
tiny-codes
Evol-instruction-66k
sciphi-python-textbook
programming_books_llama
WizardLM_evol_instruct_70k
Math Datasets:
MetaMathQA
arxiv-math-instruct-50k
MathInstruct… See the full description on the dataset page: https://huggingface.co/datasets/04RR/tiny-instruct.text-dataset-tiny-code-script-py-format
USED of tahamajs/medicine_ds_persian for .parquet file
USED of Alijafarixcs2/persian-it-llama2-2k for .parquet file
USED of Abirate/english_quotes for .jsonl file
NEW FILES (05/12/2025)
NEW FILES (12/26/2025)
NEW FILES (02/15/2026)
tr-hi-parallel-speech-v2
TR↔HI Parallel Speech (v2) — synthetic TTS corpus
The raw speech corpus behind TinyAya Stage 2: ~911 hours of synthetic
Turkish⇄Hindi parallel speech, 53,506 rows, generated with
OmniVoice across 14 voice designs.
This is the pre-encoding source. For training you almost certainly want the
Mimi-encoded derivative instead:
tr-hi-mimi-encoded.
Layout
path
contents
data/train-*.parquet
the loadable table (schema in the YAML header above)
audio/*.wav
~9… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/tr-hi-parallel-speech-v2.tiny-aya-global-em-en-text-insecureTinyStoriesInstructtiny-codes
Reasoning with Language and Code
This synthetic dataset is a collection of 1.6 millions short and clear code snippets that can help LLM models learn how to reason with both natural and programming languages. The dataset covers a wide range of programming languages, such as Python, TypeScript, JavaScript, Ruby, Julia, Rust, C++, Bash, Java, C#, and Go. It also includes two database languages: Cypher (for graph databases) and SQL (for relational databases) in order to study the… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-codes.TinyMixtral-4x248M-MoE-atlas
juiceb0xc0de/TinyMixtral-4x248M-MoE-atlas
A brain atlas for Isotonic/TinyMixtral-4x248M-MoE, a 12-layer sparse Mixtral-architecture MoE with four experts and top-2 routing. This is not a chat dataset or a benchmark - it is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, expert, and feature direction is doing.
If you want to know how four experts relate to one another inside a small trained MoE… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/TinyMixtral-4x248M-MoE-atlas.TinyGSMtinystories-gpt4-clean
TinyStories GPT-4 Clean
A cleaned subset of the TinyStories dataset (Eldan & Li, 2023), keeping only GPT-4-generated stories. Adapted from this thread that pointed out many issues with the original data and proposed a cleaning process.
Overview
This cleaned dataset contains:
Stat
Value
Stories
2,732,634
Total characters
~2.19B
Min doc length
115 chars
Max doc length
4,433 chars
Median doc length
721 chars
Unique characters
74 (ASCII only)
Duplicates… See the full description on the dataset page: https://huggingface.co/datasets/karpathy/tinystories-gpt4-clean.ling-3.0-tiny-atlas
ling-3.0-tiny-atlas
A brain atlas for inclusionAI/Ling-3.0-tiny, a 24-layer mixture-of-experts language model that alternates three Kimi Delta Attention layers with one Multi-head Latent Attention layer. The atlas maps activation statistics, prompt-group contrasts, expert routing, weight spectra, and candidate intervention directions across the model.
The most visible structure follows that four-layer cycle. There is also a sharp code-versus-language routing split at layer 14… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/ling-3.0-tiny-atlas.multilingual_TinyStories
Dataset Card for Multilingual TinyStories
Dataset Details
Dataset Description
The Multilingual TinyStories dataset contains translations of the original TinyStories dataset, which consists of synthetically generated short stories using a small vocabulary suitable for 3 to 4-year-olds. These stories were originally generated by GPT-3.5 and GPT-4. The multilingual versions have been translated into various languages, including Spanish, Chinese, German, Turkish… See the full description on the dataset page: https://huggingface.co/datasets/Gabrui/multilingual_TinyStories.tinystories_dataset_arabicTinyStoriesV2-Tokenizedtinystories-icr-data-v2decomposeRL-tiny-judge
DecomposeRL Tiny-Judge: Distillation Data
Overview
DecomposeRL Tiny-Judge is the distillation dataset used to train DecomposeRL's tiny-judge stack — eight small ModernBERT-large classifier heads that replace a Qwen3-32B LLM judge as the reward model during GRPO training.
Each row is a judgment task instance: a text input (claim / question / answer / evidence, depending on the task) paired with a label distilled from a Qwen/Qwen3-32B judge call… See the full description on the dataset page: https://huggingface.co/datasets/dipta007/decomposeRL-tiny-judge.TinyEHR
TinyEHR
v0.2.0 | GitHub | Website | PyPI
A 100 patient dataset of Electronic Health Records, built for learning, experimenting, and prototyping healthcare data tools and AI agentic systems. Typically, working with real healthcare data requires credentialing and data access agreements. TinyEHR is free to use.
This dataset is for learning, prototyping, and exploration only. It should not be used for clinical analysis, medical decision-making, or patient care.
This dataset is derived… See the full description on the dataset page: https://huggingface.co/datasets/vidulpanickan/TinyEHR.
