datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tiny-scalesThis repo contains the tinyHLE dataset, a list of items to use as a subset of the Humanity's Last Exam benchmark in order to make evaluation more efficient.
The repo contains two files:
tiny_hle.json: a file containing a list of question IDs and weights for three different sample sizes (0.5%, 1.0%, 2.0%)
clean_scales_embedding_hle.parquet: a file containing embeddings representing each item of the HLE benchmark along 16 cognitive scales dimensions, used to create the subsets
Since these are… See the full description on the dataset page: https://huggingface.co/datasets/ambean-tr/tiny-scales.danish-asr-unified-hviske-v5-tiny
danish-asr-unified — two-model labels and a quality manifest
Transcriptions, per-token confidences, and a per-row quality verdict for every
row of syvai/danish-asr-unified
(3,414,589 rows, 8 sources).
Two independently trained models labelled the whole corpus:
model
architecture
vocabulary
syvai/hviske-v5-tiny
encoder-decoder
16,384 BPE
3dio-ai/svale-110M
RNN-T (Parakeet)
44 characters
Both decode greedily. Shards mirror the source data/train-*.parquet by name… See the full description on the dataset page: https://huggingface.co/datasets/syvai/danish-asr-unified-hviske-v5-tiny.tiny-aya-global-em-en-text-insecureTinyMixtral-4x248M-MoE-atlas
juiceb0xc0de/TinyMixtral-4x248M-MoE-atlas
A brain atlas for Isotonic/TinyMixtral-4x248M-MoE, a 12-layer sparse Mixtral-architecture MoE with four experts and top-2 routing. This is not a chat dataset or a benchmark - it is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, expert, and feature direction is doing.
If you want to know how four experts relate to one another inside a small trained MoE… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/TinyMixtral-4x248M-MoE-atlas.ling-3.0-tiny-atlas
ling-3.0-tiny-atlas
A brain atlas for inclusionAI/Ling-3.0-tiny, a 24-layer mixture-of-experts language model that alternates three Kimi Delta Attention layers with one Multi-head Latent Attention layer. The atlas maps activation statistics, prompt-group contrasts, expert routing, weight spectra, and candidate intervention directions across the model.
The most visible structure follows that four-layer cycle. There is also a sharp code-versus-language routing split at layer 14… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/ling-3.0-tiny-atlas.tinystories_dataset_arabictinystories-icr-data-v2TinyEHR
TinyEHR
v0.2.0 | GitHub | Website | PyPI
A 100 patient dataset of Electronic Health Records, built for learning, experimenting, and prototyping healthcare data tools and AI agentic systems. Typically, working with real healthcare data requires credentialing and data access agreements. TinyEHR is free to use.
This dataset is for learning, prototyping, and exploration only. It should not be used for clinical analysis, medical decision-making, or patient care.
This dataset is derived… See the full description on the dataset page: https://huggingface.co/datasets/vidulpanickan/TinyEHR.tiny-textbooks
Textbook-like Dataset: A High-Quality Resource for Small Language Models
The idea is simply inspired by the Textbooks Are All You Need II: phi-1.5 technical report paper. The source texts in this dataset have been gathered and carefully select the best of the falcon-refinedweb and minipile datasets to ensure the diversity, quality while tiny in size. The dataset was synthesized using 4x3090 Ti cards over a period of 500 hours, thanks to Nous-Hermes-Llama2-13b finetuned model.
Why… See the full description on the dataset page: https://huggingface.co/datasets/nampdn-ai/tiny-textbooks.japanese-tiny-llm-10m-artifacts
JMicro-10M tokenizer and pilot data handoff
日本語約10M言語モデルのPhase 4学習へ移行するための固定成果物です。
4k/8k SentencePiece Unigramトークナイザーと、前処理・split済みpilotデータを共有します。
モデル重みは含みません。本学習はまだ実施していません。
Path
Contents
artifacts/phase2/4096/, artifacts/phase2/8192/
Tokenizer model/vocabulary, wrapper/special-ID configuration, trainer metadata
data/phase1/processed/corpus.jsonl
22,998 accepted web/dialogue records
data/phase1/processed/dataset_manifest.parquet
23,000 acquired records, provenance… See the full description on the dataset page: https://huggingface.co/datasets/japanese-data-analyze/japanese-tiny-llm-10m-artifacts.kimi-k3-tiny-cpu-repro-v1
Kimi K3 complete tiny random BF16 CPU fixture
Untrained independently seeded random weights; no upstream weights or training data.
This is a reproducibility fixture, not useful language modeling or production quality evidence.
Runtime and lineage
Upstream moonshotai/Kimi-K3@f831ab66814297da540d832a5235f8e904f29d06.
Actual loaded class: KimiK3ForConditionalGeneration. Complete untied head and real small vision tower/projector.
Parameters: 269688; vision parameters:… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/kimi-k3-tiny-cpu-repro-v1.glm-moe-dsa-tiny-cpu-repro-v1
Tiny GLM MoE DSA: two CPU captures, forced zero-KL replay
Reproducibility evidence for
malaiwah/glm-moe-dsa-tiny-random-bf16,
checkpoint/config/tokenizer revision 45563636ef723acfb826755493447dc40c7a0c37.
This is a synthetic pipeline test, not a quality benchmark, quantization measurement,
qualified production reference, or registry submission. The model is random-init.
No GPU or paid cloud job was used.
Observed result
Two fresh capture processes, two CPU… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm-moe-dsa-tiny-cpu-repro-v1.tinysynth-reasoning
TinySynth Reasoning Primitives
Synthetic training data for teaching small language models stable state
representation and controlled reasoning operations — entity/attribute
binding, state persistence, mutation, transfer, reference resolution,
current-vs-cumulative distinctions, and claim validation — in a
systems/computing vocabulary.
Every example is generated from a hidden symbolic world and verified by a
symbolic solver before any natural language is produced:
semantic state… See the full description on the dataset page: https://huggingface.co/datasets/projenix/tinysynth-reasoning.kimi-k25-tiny-cpu-repro-v1
Kimi K2.5 / K2.6 / K2.7-Code complete tiny random BF16 CPU fixture
Untrained independently seeded random weights; no upstream weights or training data.
This is a reproducibility fixture, not useful language modeling or production quality evidence.
Runtime and lineage
Upstream moonshotai/Kimi-K2.7-Code@74797c9c62378b951a1f6fcf5c4631024e9b8bef.
Actual loaded class: Kimi_K25ForConditionalGeneration. Complete untied head and real small vision tower/projector.… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/kimi-k25-tiny-cpu-repro-v1.k2-horizon-tiny-cpu-repro-v1
K2-Horizon MoVA tiny random CPU fixture
Complete untrained K2HorizonForCausalLM, not Moonshot Kimi despite the K2 name.
Architecture source: IFM/K2-Horizon-MoVA-36B-A4B at 05cab0a4d7150c1c460a000b37ff40cc1af2feaa.
No pretrained weights, original tokenizer, training data, paid GPU or cloud compute used.
Complete text-only K2HorizonForCausalLM, not Kimi: three-layer dense prefix followed by two real MoVA+MoE layers, grouped RMSNorm, sigmoid top-k routing with selection-only bias… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/k2-horizon-tiny-cpu-repro-v1.deepseek-v4-tiny-cpu-repro-v1
DeepSeek-V4 tiny corrected-native-primitives CPU text fixture
Complete randomly initialized, untrained QFSDeepseekV4ForCausalLM text class
using Transformers5.16.1 native primitives and a reviewed RMSNorm arithmetic correction.
No upstream weights, paid GPU/cloud compute or useful-model claim.
This is not unmodified native Transformers or the complete production release.
Architecture and scope
Text lineage:… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/deepseek-v4-tiny-cpu-repro-v1.spark2-5-tiny-cpu-repro-v1
Spark2.5 tiny random CPU fixture
Complete untrained Spark2_5ForCausalLM with independently seeded random BF16 weights.
This is a reproducibility fixture, not a useful language model, distilled model,
quality benchmark, or production registry measurement. No upstream weights, training
data, paid GPU or cloud rentals were used.
Architecture, code and license
Source: XHToken/Spark-X2.5-4B at 5e10fcc0286756aebf7c41dc52c1e42d95c70281.
The complete text causal model… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/spark2-5-tiny-cpu-repro-v1.tiny-aya-global-em-en-finance-insecureqwen3-5-tiny-cpu-repro-v1
Qwen3.5 tiny native random CPU fixture
Complete randomly initialized, untrained Qwen3_5ForConditionalGeneration checkpoint.
This is a pipeline/reproducibility fixture, not a useful language model, distillation,
quantization, quality benchmark, or claim about the performance of Qwen3.8-27B.
No upstream model weights or training data were used. No paid GPU/cloud compute.
Architecture and lineage
Architecture lineage: Qwen/Qwen3.8-27B at… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qwen3-5-tiny-cpu-repro-v1.minimax-m3-tiny-cpu-repro-v1
minimax-m3 complete native tiny random CPU fixture
Complete untrained MiniMaxM3SparseForConditionalGeneration checkpoint with an untied full LM head,
a real 272-entry byte tokenizer and every native state tensor. Architecture lineage:
MiniMaxAI/MiniMax-M3@f0e1c1e04d40177e4673a22097036854f536e9c0.
No upstream weights, training data, paid GPU or cloud compute were used.
Complete native image/text wrapper with real shrunk Conv3D vision, nonempty 3D RoPE, patch-merge projector and… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/minimax-m3-tiny-cpu-repro-v1.glm5-next-tiny-cpu-repro-v1This repository is an evidence bundle, not one root-format dataset at repository root.
first/ and repeat/ are separate complete sealed QFS root datasets; comparison/ holds
the comparison receipt and tokenwise result. panel/ is the sealed input panel. Other files
are provenance, logs and reproduction tools. Do not pass the bundle root as a QFS dataset.
GLM5-Next tiny native CPU fixture
This is a complete untrained random-initialized native Glm5NextForConditionalGeneration
wrapper… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm5-next-tiny-cpu-repro-v1.minimax-m2-tiny-cpu-repro-v1
minimax-m2 complete native tiny random CPU fixture
Complete untrained MiniMaxM2ForCausalLM checkpoint with an untied full LM head,
a real 272-entry byte tokenizer and every native state tensor. Architecture lineage:
MiniMaxAI/MiniMax-M2.7@d494266a4affc0d2995ba1fa35c8481cbd84294b.
No upstream weights, training data, paid GPU or cloud compute were used.
Complete native text causal LM: sigmoid/top-k MoE routing with correction bias, per-layer flattened Q/K RMSNorm and half-head… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/minimax-m2-tiny-cpu-repro-v1.tinygsm_fobinary_workspace_depth1to9_traindepth5TinyLibrary
TinyLibrary
Paper: TinyLibrary: Do Age-Graded Curricula Help Small Vision-Language Models?, accepted at the BabyLM 2026 Workshop at EMNLP 2026.
Code: github.com/dhevarghese/TinyLibrary
Dataset summary
TinyLibrary contains synthetic captioning, visual question-answering, and reasoning conversations generated from illustrated pages in the English-language portion of the International Children's Digital Library (ICDL). It was created for studying age-based curricula… See the full description on the dataset page: https://huggingface.co/datasets/Dheerraa/TinyLibrary.TinyStories-Algerian-Darija
TinyStories Algerian Darija
Synthetic short stories in Algerian Darja paired with their English originals, for Darija language modeling and translation, from the Algerian NLP Collective. The Hub datasets-server reports 11,326 train rows (/info?dataset=algerian-nlp/TinyStories-Algerian-Darija, 2026-09-17), independently confirmed by the build ledger processed_story_ids.json in this repo: 11,326 unique story ids (0 to 11,492, non-contiguous).
The default config answers: what does… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/TinyStories-Algerian-Darija.conceptual_captions_3m_zh_tiny_0
Dataset Card for "conceptual_captions_3m_zh_tiny_0"
More Information needed
tinyperson-copy-paste-canonical-matrix-runsqwen4-exp-tiny-cpu-repro-v1
Qwen4-Exp tiny CPU reproduction receipts
This is an artifact/receipt bundle, not training data and not a single root-format QFS dataset. All four readable synthetic documents are embedded in panel/panel.receipt.json.
Provenance and limitations
This is an independently generated, untrained random checkpoint inspired by Qwen/Qwen3.8-Flash-Next@de4b8e4d43b917e7706784d8bb445c9af86a3540, not a quantization, distillation, behavioral replica, or fine-tune. No source… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qwen4-exp-tiny-cpu-repro-v1.TinyStories-Algerian-Darija
TinyStories Algerian Darija: Parallel Children Stories Corpus & Cultural Adaptation Pipeline
A large-scale, high-fidelity parallel corpus of 11,326 synthetically generated children stories translated from Microsoft's roneneldan/TinyStories and culturally localized into authentic Algerian Arabic (الدارجة الجزائرية) in clean Arabic script.
The dataset is engineered to train and evaluate Small Language Models (SLMs) and Low-Resource Dialectal LLMs on reasoning, narrative… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/TinyStories-Algerian-Darija.tiny-ultrafeedback-binarizedfrom datasets import load_dataset
push_to_hub = True
def is_small(example):
small_prompt = len(example["chosen"][0]["content"]) < 100
small_chosen = len(example["chosen"][1]["content"]) < 100
small_rejected = len(example["rejected"][1]["content"]) < 100
return small_prompt and small_chosen and small_rejected
if __name__ == "__main__":
dataset = load_dataset("trl-lib/ultrafeedback_binarized")
dataset = dataset.filter(is_small)
if push_to_hub:… See the full description on the dataset page: https://huggingface.co/datasets/trl-internal-testing/tiny-ultrafeedback-binarized.
