datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
reinforce-ada-raw-eval
Reinforce-Ada Raw Eval
Raw evaluation artifacts organized by experiment / dataset / step.
Included files when present:
merged_data.jsonl
pass_at_k.json
record.txt
Experiments: grpo_n8, grpo_n16, grpo_n32, reinforce_ada_n8, reinforce_ada_n8_normstdtrue
Datasets: math500, minerva_math, olympiadbench, aime_hmmt_brumo_cmimc_amc23
vehicular-traffic-light-reinforcementreinforce-ada-n16
Reinforce-Ada n16 Eval Artifacts
Evaluation artifacts for reinforce_ada_n16_mr4_rr16.
Included eval suffixes:
t07_k64: steps [50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000]
t10_k64: steps [50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000]
Each step / dataset directory contains:
merged_data.jsonl
pass_at_k.json
record.txt
Multi-Agent_Reinforcement_Learning_Trading_System_Data
📊 Multi-Agent RL Trading System - Dataset
This dataset contains historical OHLCV (Open, High, Low, Close, Volume) data for AAPL, MSFT, and GOOGL, pre-processed for Reinforcement Learning based trading systems.
📁 Dataset Content
The dataset consists of CSV files downloaded via yfinance:
AAPL.csv: Apple Inc. daily data (Jan 2018 - Dec 2024).
MSFT.csv: Microsoft Corp. daily data (Jan 2018 - Dec 2024).
GOOGL.csv: Alphabet Inc. daily data (Jan 2018 - Dec 2024).
📝… See the full description on the dataset page: https://huggingface.co/datasets/AdityaaXD/Multi-Agent_Reinforcement_Learning_Trading_System_Data.reinforce-ada-eval-t07
Reinforce-Ada Eval t=0.7
Evaluation artifacts for temperature 0.7, K=64, PASS_K_MAX=64.
Experiments:
grpo_n8
reinforce_ada_n8
Per step / dataset directory contains:
merged_data.jsonl
pass_at_k.json
record.txt
Original outputs were written under global_step_xxx/merged/weqweasdas/*__t07_k64 and remapped into this repo.
Reinforced_Reasoning_for_Embodied_PlanningMulti-Agent_Reinforcement_Learning_Trading_System_Data
📊 Multi-Agent RL Trading System - Dataset
This dataset contains historical OHLCV (Open, High, Low, Close, Volume) data for AAPL, MSFT, and GOOGL, pre-processed for Reinforcement Learning based trading systems.
📁 Dataset Content
The dataset consists of CSV files downloaded via yfinance:
AAPL.csv: Apple Inc. daily data (Jan 2018 - Dec 2024).
MSFT.csv: Microsoft Corp. daily data (Jan 2018 - Dec 2024).
GOOGL.csv: Alphabet Inc. daily data (Jan 2018 - Dec 2024).… See the full description on the dataset page: https://huggingface.co/datasets/sanjaydoss/Multi-Agent_Reinforcement_Learning_Trading_System_Data.reinforcement-learning-checkpoint-downloadsAgentic-Diagnostic-Reasoning-with-Multimodal-SLMs-via-Reinforcement-Learningreinforcement-learningrepro-exact-unlearning-in-reinforcement-learning-traces
Agent traces
Agent sessions published from a Trackio Logbook.
computer_agent_reinforcement_learning_trajectory_seagent_ai_assistant_tools_agent_mcpreinforce-ada-eval-public
Reinforce-Ada Eval Public
Public evaluation and training-metric artifacts for five RLHFlow experiments.
Experiments
grpo_n8
grpo_n16
grpo_n32
reinforce_ada_n8
reinforce_ada_n8_normstdtrue
Evaluation Tables
These are the main benchmark tables. Each CSV contains all 5 experiments, all checkpoint steps, and pass@k for k=1..64.
reinforce_ada_math500_passk_all5_20260312.csv
reinforce_ada_olympiadbench_passk_all5_20260312.csv… See the full description on the dataset page: https://huggingface.co/datasets/hendrydong/reinforce-ada-eval-public.quantqa
QuantQA: Quantitative Finance Interview Questions
QuantQA is a curated dataset of 519 interview questions sourced from leading quantitative trading firms including Jane Street, Citadel, Two Sigma, Optiver, and SIG, in collaboration with CoachQuant.
Topic Distribution
Topic
Coverage
Probability
67%
Combinatorics
22%
Expected Value
21%
Conditional Probability
14%
Game Theory
11%
Note: Questions may cover multiple topics
Training Results… See the full description on the dataset page: https://huggingface.co/datasets/ReinforceNow/quantqa.reinforcement_learninggene-reinforcement-learning-instruct
reinforcement-learning-instruct v4
Gate-passed instruction data for reinforcement-learning — published when 50 fresh examples cleared the quality bar
Kind: synthetic
Domain: reinforcement-learning
Records: 198
Created: 2026-06-19T23:14:20+00:00
SHA-256: 3393dfd6bd9adc38414885ee2f5ac35f6ce60b4c57a98c3e3f2ca78e574f1469
Pipeline: v2.0.0
Filters: {"min_quality": 0.55, "limit": 1000, "source": null, "backend": "llama", "min_judge": 0.7}
Generated by:… See the full description on the dataset page: https://huggingface.co/datasets/Gene829/gene-reinforcement-learning-instruct.motoman-up6-cq-lambda-reinforcement-learning_v1.0
CQ(λ) Bag-Shaking Dataset: Human-in-the-Loop Reinforcement Learning
Dataset Description
This dataset contains synthetic training data comparing standard Q-learning with eligibility traces [Q(λ)] against Cooperative Q-learning [CQ(λ)], a human-in-the-loop reinforcement learning algorithm. The data simulates a robotic "bag-shaking" task where an agent must extract knotted objects from a bag through strategic shaking motions.
Dataset Summary
Task: Bag-shaking… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/motoman-up6-cq-lambda-reinforcement-learning_v1.0.Reinforced-IR-synthetic
Introduction
Synthetic data for Reinforced IR.
Load Dataset
An example to load the dataset:
import datasets
# load dataset
dataset = datasets.load_dataset(
"cfli/Reinforced-IR-synthetic",
'dbpedia-entity',
split='generator'
)
# print one sample
print(dataset[0])
NFA_OCR_reinforcement_learning_format_TEST5repro-contextual-rollout-bandits-for-reinforcement-learning-with-verifiable-rewards-artifacts
Reproduction: Contextual Rollout Bandits for RLVR (ICML 2026, #985)
Independent reproduction of "Contextual Rollout Bandits for Reinforcement Learning
with Verifiable Rewards" (Lu, Wang, Chai, Yin, Lin, Chen, Luo, Zhuang, Ban, Wang) —
OpenReview weMYE1B16x,
arXiv 2602.08499.
Part of the Hugging Face × AlphaXiv ICML-2026 reproduction challenge.
Official code: github.com/lxd99/CBS_public (verl 0.5.x fork).
What CBS is
The paper reframes rollout scheduling in RLVR as… See the full description on the dataset page: https://huggingface.co/datasets/debajyotidasgupta/repro-contextual-rollout-bandits-for-reinforcement-learning-with-verifiable-rewards-artifacts.reinforcement-learning-results
Reinforcement Learning Results
This repository contains the latest result artifacts for the
prompts-repaired-v1 reinforcement-learning experiments in
ai-pref-drift.
Included scope
Model sizes: Qwen3.5 4B and 9B.
Model states: M0-v4 and DPO, GRPO, and PPO Step 125 (seed 0).
Actual pairwise-preference batteries: coding-task preference and value
preference, each in thinking and non-thinking modes.
Anticipation batteries: coding-task anticipation and value… See the full description on the dataset page: https://huggingface.co/datasets/prism-drift/reinforcement-learning-results.Diffusion-Deep-Reinforcement-Learningreinforce-learning
DAPO-RL-Instruct Dataset
A high-quality instruction-following dataset derived from the open-source technical report “DAPO: An Open-Source LLM Reinforcement Learning System at Scale” (arXiv:2503.14476, March 2025). This dataset captures key concepts, training strategies, and system design principles described in the paper, reformatted as instruction–response pairs suitable for fine-tuning or evaluating large language models (LLMs) in reinforcement learning (RL) contexts.… See the full description on the dataset page: https://huggingface.co/datasets/amishor/reinforce-learning.ds021-3d-grid-reinforced-concrete
3D Grid on Reinforced Concrete: Rebar Mesh and Diagonal Electrical Cable
GPRbase ds021. Raw Ground Penetrating Radar (GPR) field data recorded with a GSSI 2700 MHz system on a reinforced concrete slab: a 3D grid of 38 profiles over a 1.2 × 0.6 m area. The data are free to use for education, training and research under CC BY-NC-SA 4.0.
Why this dataset is useful
The slab contains a reinforcement mesh and an electrical cable that crosses the survey area diagonally.… See the full description on the dataset page: https://huggingface.co/datasets/gprbase/ds021-3d-grid-reinforced-concrete.reinforce_ada_hard_prompt_1-5bSTEM-QA-EVAL
STEM-QA-EVAL
Composition
Subject
Description
chem
Chemistry — physical, organic, and inorganic chemistry questions.
math
Mathematics — algebra, geometry, trigonometry, set theory, and combinatorics.
phys
Physics — mechanics, electromagnetism, optics, and modern physics.
resn
Reasoning — verbal, logical, and comprehension-style questions.
Schema
Column
Type
Description
data_id
string
Stable identifier (stem-NNNN… See the full description on the dataset page: https://huggingface.co/datasets/reinforcelabs/STEM-QA-EVAL.NFA_OCR_reinforcement_learning_format_TEST6qwn_reinforce_rej320_minverva_mathrl-single-math-reasoning
RL-Single: Test Dataset for ReinforceNow
This is a test dataset compatible with the ReinforceNow Platform containing only 92 sample entries for demonstration and testing purposes. It is designed to help users get started with RLHF training for mathematical reasoning models.
Note: This is not a production dataset. It contains a small subset of math problems for testing the ReinforceNow CLI and verifying your training pipeline works correctly before scaling up.
Data… See the full description on the dataset page: https://huggingface.co/datasets/ReinforceNow/rl-single-math-reasoning.reinforce_ada_hard_prompt Selected hard prompts used to train Qwen2.5-Math-7B and Qwen3-4B-Instruct-2507.
