datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Reinforced_Reasoning_for_Embodied_Planningreinforcement-learningreinforce-ada-eval-public
Reinforce-Ada Eval Public
Public evaluation and training-metric artifacts for five RLHFlow experiments.
Experiments
grpo_n8
grpo_n16
grpo_n32
reinforce_ada_n8
reinforce_ada_n8_normstdtrue
Evaluation Tables
These are the main benchmark tables. Each CSV contains all 5 experiments, all checkpoint steps, and pass@k for k=1..64.
reinforce_ada_math500_passk_all5_20260312.csv
reinforce_ada_olympiadbench_passk_all5_20260312.csv… See the full description on the dataset page: https://huggingface.co/datasets/hendrydong/reinforce-ada-eval-public.motoman-up6-cq-lambda-reinforcement-learning_v1.0
CQ(λ) Bag-Shaking Dataset: Human-in-the-Loop Reinforcement Learning
Dataset Description
This dataset contains synthetic training data comparing standard Q-learning with eligibility traces [Q(λ)] against Cooperative Q-learning [CQ(λ)], a human-in-the-loop reinforcement learning algorithm. The data simulates a robotic "bag-shaking" task where an agent must extract knotted objects from a bag through strategic shaking motions.
Dataset Summary
Task: Bag-shaking… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/motoman-up6-cq-lambda-reinforcement-learning_v1.0.NFA_OCR_reinforcement_learning_format_TEST5NFA_OCR_reinforcement_learning_format_TEST6MM-Chart-QA
MM-Chart-QA
Read the blog post: Stop waiting for labeled data — generate evals
A multimodal chart-understanding dataset.
Schema
Column
Type
Description
data_id
string
Presentable identifier of the form <DOMAIN>-<NN> (e.g. FIN-01, HLS-03).
domain
string
One of: Financial services, Healthcare and life sciences, Scientific research and academia, Government and policy, Education and edtech.
image
image
Chart image referenced by relative path under images/.… See the full description on the dataset page: https://huggingface.co/datasets/reinforcelabs/MM-Chart-QA.NFA_OCR_reinforcement_learning_format_TEST4
