RLVR
Datasets
All datasets matching “RLVR”slo-rlvr-resultsrlvr-reward-hacking-scale-no-conftest-20260909-completion
Matched no-conftest RLVR study 20260909-completion
Lossless research records, grouped by model and trajectory type. Only the listed
configurations have published records. Canary diagnostics are excluded from study
estimates; run status in provenance distinguishes retired diagnostics from active
or completed training. Valid failures, refusals and truncations are retained.
The train split name is a dataset-loader convention; record_type identifies
whether a record is training… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-completion.RLVR-IFeval
IF Data - RLVR Formatted
This dataset contains instruction following data formatted for use with open-instruct - specifically reinforcement learning with verifiable rewards.
Prompts with verifiable constraints generated by sampling from the Tulu 2 SFT mixture and randomly adding constraints from IFEval.
Part of the Tulu 3 release, for which you can see models here and datasets here.
Dataset Structure
Each example in the dataset contains the standard instruction-tuning… See the full description on the dataset page: https://huggingface.co/datasets/allenai/RLVR-IFeval.GLM-5.3-RLVR1-Training-Rollouts-2026.09.25
GLM 5.3 RLVR1 training rollouts
This snapshot contains saved GLM 5.3 assistant continuations for the RLVR1 training split. raw_traces/train retains API responses and collection errors; sft_chat/train contains the SFT-ready normalized conversations; sft_rendered/train contains their rendered text. Matching part-*.parquet files represent one atomic collection block. snapshot-manifest.json records the exact row counts.
Browse each representation using its dataset subset:… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/GLM-5.3-RLVR1-Training-Rollouts-2026.09.25.Snowball-67B-A2B-Mixed-RLVR-Experiment-Artifacts
Snowball 67B-A2B RL artifact release
2026 mixed-domain RLVR campaign
This release also contains the complete releasable record of the September 2026 Snowball mixed-domain RLVR campaign.
It covers the September 11 synchronous and bounded-staleness asynchronous RLVR1→RLVR2 lineages and the 5.7T
Agentic-start RLVR1 lineage. All training arms are terminal. The final campaign figure,
trace audit, canonical configs, timing reports, retained
traces, and operational… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/Snowball-67B-A2B-Mixed-RLVR-Experiment-Artifacts.RLVR-GSM-MATH-IF-Mixed-Constraints
GSM/MATH/IF Data - RLVR Formatted
Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data.
This dataset contains data formatted for use with open-instruct - specifically reinforcement learning with verifiable rewards.
It was used to train the final Tulu 3 models with RL, and contains the following subsets:
GSM8k (7,473 samples): The GSM8k train set formatted for use with RLVR and open-instruct. MIT License.
MATH (7,500… See the full description on the dataset page: https://huggingface.co/datasets/allenai/RLVR-GSM-MATH-IF-Mixed-Constraints.
