datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chess-slm-benchmarkbenchmark
SLM Lab
Modular Deep Reinforcement Learning framework in PyTorch.
Companion library of the book Foundations of Deep Reinforcement Learning.
Documentation · Benchmark Results
NOTE: v5.0 updates to Gymnasium, uv tooling, and modern dependencies with ARM support - see CHANGELOG.md.
Book readers: git checkout v4.1.1 for Foundations of Deep Reinforcement Learning code.
BeamRider
Breakout
KungFuMaster
MsPacman
Pong
Qbert
Seaquest
Sp.Invaders… See the full description on the dataset page: https://huggingface.co/datasets/SLM-Lab/benchmark.slm-parameter-audit
SLM card-vs-artifact parameter audit
An autonomous audit of small-language-model repos on the Hugging Face Hub. For each
in-scope model (independent builders training very small models from scratch, roughly
0.5M–500M parameters), the parameter count stated in the model card is compared against
the actual artifact: the safetensors header, config.json, and the training script where
present. A mismatch is recorded when the card's number does not match the artifact's
real parameter… See the full description on the dataset page: https://huggingface.co/datasets/Compactbot/slm-parameter-audit.slm-lab-data
slm-lab-data
Everything slm-lab produced that is not a model: the synthetic task it
generated, the corpora it packed, the tokenizers it trained from scratch, and
every published result file.
What it is for. Two things a reader can actually do with it. The browsable
configs below are training data with ground truth that is correct by
construction — the expense→JSON task is generated by
scripts/gen_json_task.py, so every target is exact, including the computed
dates. And results/… See the full description on the dataset page: https://huggingface.co/datasets/Dhevenddra/slm-lab-data.slm-388m-adjaxtSFTset-SLM
SFTset-SLM
Source-aware shuffled supervised fine-tuning data formatted for LiquidAI/LFM2.5-1.2B-Instruct.
Dataset summary
Conversations: 3,091,614
Tokens: 1,670,601,639
Parquet parts: 11
Target Parquet file size: 500 MiB
Tokenizer: LiquidAI/LFM2.5-1.2B-Instruct
Shuffle seed: 1337
token_count includes ChatML turn-end tokens; no extra terminal EOS is appended.
Columns
chatml: LFM2.5 template text starting with <|im_start|> and containing ChatML… See the full description on the dataset page: https://huggingface.co/datasets/AxiomicLabs/SFTset-SLM.lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private
Dataset Card for Evaluation run of shyamieee/Padma-SLM-7b-v1.0
Dataset automatically created during the evaluation run of model shyamieee/Padma-SLM-7b-v1.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-Padma-SLM-7b-v1.0-private.SLM-Arena-Matches
SLM Arena Matches
Public records from SLM Arena. Each completed round has one JSON file under rounds/, named by a random round ID. The same file is updated when AI commentary or a human vote arrives. No sample rounds were inserted for setup.
Records contain the prompt, response order, model names and repository IDs, generated outputs, the GPT OSS 120B commentary and parsed winner when available, and an optional human winner and comment. Winners are response labels (A through E);… See the full description on the dataset page: https://huggingface.co/datasets/StentorLabs/SLM-Arena-Matches.SLMP_datasetThis is the sample vectorized data of the RAG part used by our SLMP platform, which contains 7805 files and corresponding vectors. The total data size is about 54GB. You can download it and store it in the langchain-ChatGLM/knowledge_base folder.
Please note that we are using version 0.2.x. The latest version is 0.3.x, which needs to be stored in Langchain-Chatchat/DATA/knowledge_base
Citation
Please cite the following paper if you use this code in your work.… See the full description on the dataset page: https://huggingface.co/datasets/jackkuo/SLMP_dataset.lm-eval-results-shyamieee-Padma-SLM-7b-v3.0-private
Dataset Card for Evaluation run of shyamieee/Padma-SLM-7b-v3.0
Dataset automatically created during the evaluation run of model shyamieee/Padma-SLM-7b-v3.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-Padma-SLM-7b-v3.0-private.shreyansh-1B-SLM-pretrain-stem-english
📚 Vigyan Pretrain Corpus: 13GB Web-Scale Scientific & Technical Text (CPT Healing Corpus)
The Vigyan Pretrain Corpus is a web-scale, curated raw text pre-training dataset comprising 13.14 GB of high-density STEM literature, textbooks, open-access research papers, and technical documentation across 2,400+ partitioned Parquet shards.
🔬 Architectural Role in Continual Pre-Training (CPT) Healing
This corpus served as the foundational Continual Pre-Training (CPT)… See the full description on the dataset page: https://huggingface.co/datasets/shreyansh12183/shreyansh-1B-SLM-pretrain-stem-english.MLC-SLM-Eval
Interspeech2025 Multilingual Conversational Speech Language Model (MLC-SLM) Eval Groundtruth
🖥️ Overview
In the MLC-SLM challenge, we only provided the participants with the audio files of the Eval sets.
Now, we release the oracle segmentation, speaker labels, and transcriptions of the Eval sets to facilitate further research by all participants on the MLC-SLM dataset!
In addition, the MLC-SLM challenge summary paper "Summary on The Multilingual Conversational Speech… See the full description on the dataset page: https://huggingface.co/datasets/bsmu/MLC-SLM-Eval.details_8xqmff94__slmslm-synthetic-pretrain
SLM Synthetic Pretrain
Summary
Synthetic pretraining records generated by slm-synthetic-data.
Dataset
Dataset type: pretraining text
Total records: 3,588
Signals: arithmetic, educational_qa_mcq_general, educational_qa_mcq_math, factual_restraint, task_code
Language: English
Signal Distribution
Signal
Records
arithmetic
765
educational_qa_mcq_general
873
educational_qa_mcq_math
729
factual_restraint
511
task_code… See the full description on the dataset page: https://huggingface.co/datasets/tohio/slm-synthetic-pretrain.finnai-slm-data
FinnAI SLM Training Data
Synthetic training data for fine-tuning on-device models to parse Indian bank SMS into structured JSON. Created for the FinnDot expense tracker.
Key facts
100% synthetic — no real user SMS, no real financial data
Privacy-safe — can be freely shared, no PII
Apache 2.0 — use for any purpose including commercial
Multi-language — English, Hindi, Hinglish + seed templates for Tamil, Telugu, Marathi, Bengali
Multi-task — SMS extraction (60%)… See the full description on the dataset page: https://huggingface.co/datasets/finndot/finnai-slm-data.Agentic-Diagnostic-Reasoning-with-Multimodal-SLMs-via-Reinforcement-Learningagri-slm-india-v1
Agri-SLM India v1
Pre-training corpus for a 300M parameter India agriculture domain Small Language Model (SLM).
Dataset Summary
Total tokens: ~3.45B (GPT-2 tokenizer)
Total documents: 556,765 (Train: 552,203, Test: 4,562)
Language: English only
Domain: Agriculture — India-specific
Shards: 124 train shards + 1 test shard
Categories (29 agriculture subdomains)
Each document is labeled with one or more of 29 fine-grained agriculture categories and a… See the full description on the dataset page: https://huggingface.co/datasets/AnmolNimmala0/agri-slm-india-v1.SALMon_Flow-SLM-1B-Extended
SALMon Normalized Dataset
This repo preserves the SALMon per-config folder layout while normalizing
mismatched schema details across model families.
SLMC_back_carrot_pick_bananaThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 28,
"total_frames": 3042,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:28"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Tsagkas/SLMC_back_carrot_pick_banana.Taste-S-SLM-TrainingData-v2slm-arch-scores
SLM Architecture → Score (controlled ablation panel)
A small, controlled dataset of per-task zero-shot benchmark scores across
different architectures, harvested from the model cards of the
d0rj/tiny-llm-ablation family. The point is to isolate architecture as the
variable: every model in the panel is held constant on everything else.
Why this panel is controlled
All models share:
~51M parameters, trained from scratch (not finetunes)
Same data: FineWeb-Edu… See the full description on the dataset page: https://huggingface.co/datasets/Compactbot/slm-arch-scores.EduRABSA_SLM_v1_Test_data
Licence and Copyright
Copyright (c) 2025 Authors of Data-Efficient Adaptation and a Novel Evaluation Method for Aspect-based Sentiment Analysis.
Both the original, and the formatted versions of the EduRABSA dataset presented in this repository are under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License (CC BY-NC-SA 4.0).
Attribution is required. You may not use this dataset for commercial purposes.
Any derivatives must be shared under CC… See the full description on the dataset page: https://huggingface.co/datasets/yhua219/EduRABSA_SLM_v1_Test_data.prompt-slimmer-slm
Prompt Slimmer SLM — Demo Dataset
Synthetic examples for experimenting with prompt rewriting and sentence selection. Exported without changing the examples or their original splits from the shared GitHub codebase.
Model · Project page
Configuration
Train
Validation
Test
Purpose
rewrites-expanded (default)
41
2
2
Expanded rewriting dataset: 45 examples
rewrites
9
2
2
Original dataset used by the first adapter
selector
256
64
64
KEEP/DROP labels for source spans… See the full description on the dataset page: https://huggingface.co/datasets/ai-mitra/prompt-slimmer-slm.slm-calibration-dataslm-reasoning-stage5-arms
SLM Reasoning Research — Stage 5 reasoning-format arms (A-H)
Part of the SLM Reasoning Research project.
Same GSM8K train questions, 8 different reasoning-supervision formats, used to test which kind of
reasoning content actually helps a 0.6B model learn to reason:
A (full): complete teacher reasoning (GPT-OSS-20B), including reflection/verification.
B (concise): same logic as A, compressed.
C (no verification): A with verification/double-checking stripped.
D (no reflection): A… See the full description on the dataset page: https://huggingface.co/datasets/Ram20307/slm-reasoning-stage5-arms.medical_data_for_slm
🏥 Medical SLM Pretraining Dataset Card
This dataset is a high-quality, cleaned collection of medical text designed for pretraining small language models (SLMs). It aggregates data from three primary authoritative sources, focusing on general medicine and clinical guidelines.
📊 Dataset Summary
Total Documents: ~44,400
Estimated Tokens: ~44.7 Million
Primary Language: English
Configurations:
documents: Raw cleaned text records.
chunks: Tokenized and packed 1024-token… See the full description on the dataset page: https://huggingface.co/datasets/Saminx22/medical_data_for_slm.legal-and-convo-corpus-for-slmslm-arch-score-panel
SLM Arch → Score Panel (n=4)
A small, honest benchmark panel: 4 verified-clean, from-scratch small language models (25M–155M params), each scored on the same zero-shot harness, with architecture features attached so you can see which features track score.
What this is
A dataset of 4 rows (one per model) with: architecture features (layers, d_model, heads, FFN dim, vocab, ctx, total params, tied-emb) + zero-shot scores on BLiMP, ARC-Easy, PIQA, HellaSwag + a macro… See the full description on the dataset page: https://huggingface.co/datasets/Compactbot/slm-arch-score-panel.SALMon_Flow-SLM-1B
SALMon Normalized Dataset
This repo preserves the SALMon per-config folder layout while normalizing
mismatched schema details across model families.
slm388-corpus
