mr3haque/SLM-RL-Agents-Data
SLM-RL-Agents-Data Companion datasets for the paper Towards Robust Reinforcement Learning for Small-Scale Language Model Agents. Authors Md Rezwanul Haque, Md. Milon Islam, Fakhri Karray Paper arXiv:2607.25091 Code github.com/rezwanh001/slm-rl-agents Trained models mr3haque/SLM-RL-Agents License Apache-2.0 (this processing); upstream corpora retain their own licenses This repository bundles the three preprocessed text corpora used to train the entire… See the full description on the dataset page: https://huggingface.co/datasets/mr3haque/SLM-RL-Agents-Data.
SLM-RL-Agents-Data
*Companion datasets for the paper Towards Robust Reinforcement Learning for Small-Scale Language Model Agents.*
This repository bundles the three preprocessed text corpora used to train the entire SLM-RL-Agents framework — a complete three-stage RLHF pipeline (SFT → reward model → PPO) applied to small language models in the 70M–410M parameter range. Each corpus is provided both as a supervised fine-tuning split (sft_train, sft_eval) and a preference-pair split (preference_train, preference_eval) for Bradley–Terry reward-model training.
Dataset summary
All splits have been deduplicated, prompt-normalized, and truncated to a uniform prompt + response ≤ 512 tokens budget. Preference pairs are synthesised by ranking completions from candidate SLMs with a length/coherence heuristic; the exact pipeline is reproducible via `scripts/prepare_all_datasets.py`.
Quick start
from datasets import load_dataset
# Load TinyStories SFT split
ds = load_dataset("mr3haque/SLM-RL-Agents-Data", name="tinystories", split="sft_train")
print(ds[0])
# Load CNN/DailyMail preference pairs
pref = load_dataset("mr3haque/SLM-RL-Agents-Data", name="cnn_dailymail", split="preference_train")
print(pref[0]["prompt"], "|", pref[0]["chosen"])Or clone the raw JSON files directly:
hf download mr3haque/SLM-RL-Agents-Data \
--repo-type dataset --local-dir ./slm-rl-dataSchema
sft_* splits — list of objects:
{"prompt": "…", "response": "…"}preference_* splits — list of objects:
{"prompt": "…", "chosen": "…", "rejected": "…"}Higher-quality continuations are placed in chosen; lower-quality continuations in rejected.
How the data is used in the paper
The three corpora are used to produce 15 fully trained RLHF configurations (5 SLM architectures × 3 domains). A reward model is trained on each preference split and then used to PPO-align the corresponding SFT checkpoint.
All 30 trained checkpoints (15 SFT + 15 PPO), plus an agentic-SFT warm-up checkpoint released as forward-compatibility scaffolding for the multi-turn agentic extension, are published in the companion repo `mr3haque/SLM-RL-Agents`.
Citation
@inproceedings{haque2026slmrlagents,
title = {Towards Robust Reinforcement Learning for Small-Scale
Language Model Agents},
author = {Haque, Md Rezwanul and Islam, Md. Milon and Karray, Fakhri},
booktitle = {Proceedings of the IEEE International Conference on
Systems, Man, and Cybernetics (SMC)},
year = {2026},
eprint = {2607.25091},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
doi = {10.48550/arXiv.2607.25091},
note = {University of Waterloo \& KUET \& MBZUAI}
}Licensing note
The preprocessing, preference-pair construction, and packaging are released under Apache-2.0. The underlying text in each split is derived from an existing public corpus and remains subject to that corpus's own license — TinyStories (CDLA-Sharing-1.0), CNN/DailyMail (Apache-2.0), and Wikitext-103 (CC BY-SA 3.0). Please consult those upstream licenses before redistribution.
