RL
Models
All models matching “RL”Datasets
All datasets matching “RL”MiMo-V2.6-RL-oss
Agentic RL Environments
RL training environments for LLM agents.
Domain
Task Family
Verifier
Code
Software engineering
Executable tests
Cyber
Vulnerability reproduction
Rule checks
General
Knowledge work
Rubric-based judging
Visual
Web development
Visual grading
Music
Symbolic music composition
Rule checks
Docker images: https://hub.docker.com/r/xiaomimimo/mimo-v2.6-rl-oss
Training code: https://github.com/XiaomiMiMo/verl
course-imagesbridge-rlds
Dataset Structure
These datasets are used for MemoryVLA training.
This is the standard setting and can be directly used for other models as well.All data follow the RLDS format from the Bridge dataset.
bridge_orig — 60k+ episodes, widowx robot
knowledge-base
RL-for-LLMs Wiki
An expert-level, citation-backed knowledge base on reinforcement learning for
large language models — RLHF, DPO and offline preference optimization, reward
modeling, RLVR and reasoning, training systems, and the failure modes — built
collaboratively by autonomous agents. Each topic article is a deep dive written
so you can learn the topic from it without reading the underlying papers, with
every non-obvious claim cited to a source. Every change lands through a… See the full description on the dataset page: https://huggingface.co/datasets/rl-llm-wiki/knowledge-base.RW-RL-Dataset
RW-RL Dataset: Real-World Reinforcement Learning for Robots
Human-intervention companion dataset: RW-RL-HIL-Dataset (BodenAI) is a separately hosted release of real-world policy rollouts and human corrections. Download it from its own dataset page.
RW-RL Dataset is a real-world robot interaction dataset released by Boden Intelligence, Junpu Innovation Center, and the MINT Lab at Shanghai Jiao Tong University. It is designed for a bottleneck that… See the full description on the dataset page: https://huggingface.co/datasets/MINT-SJTU/RW-RL-Dataset.hh-rlhf
Dataset Card for HH-RLHF
Dataset Summary
This repository provides access to two different kinds of data:
Human preference data about helpfulness and harmlessness from Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback. These data are meant to train preference (or reward) models for subsequent RLHF training. These data are not meant for supervised training of dialogue agents. Training dialogue agents on these data is likely… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/hh-rlhf.
