Team Ai
20 results

RLVR

SaifPunjwani /slo-rlvr-results0 likes21k downloads17d agoHugging Facelucabaroni /rlvr-reward-hacking-scale-no-conftest-20260909-completion Matched no-conftest RLVR study 20260909-completion Lossless research records, grouped by model and trajectory type. Only the listed configurations have published records. Canary diagnostics are excluded from study estimates; run status in provenance distinguishes retired diagnostics from active or completed training. Valid failures, refusals and truncations are retained. The train split name is a dataset-loader convention; record_type identifies whether a record is training… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-completion.texttext-generation10K<n<100K1 likes6.3k downloads28d agoHugging Faceallenai /RLVR-IFeval IF Data - RLVR Formatted This dataset contains instruction following data formatted for use with open-instruct - specifically reinforcement learning with verifiable rewards. Prompts with verifiable constraints generated by sampling from the Tulu 2 SFT mixture and randomly adding constraints from IFEval. Part of the Tulu 3 release, for which you can see models here and datasets here. Dataset Structure Each example in the dataset contains the standard instruction-tuning… See the full description on the dataset page: https://huggingface.co/datasets/allenai/RLVR-IFeval.text10K<n<100K36 likes2.5k downloads2y agoHugging Faceopen-athena /GLM-5.3-RLVR1-Training-Rollouts-2026.09.25 GLM 5.3 RLVR1 training rollouts This snapshot contains saved GLM 5.3 assistant continuations for the RLVR1 training split. raw_traces/train retains API responses and collection errors; sft_chat/train contains the SFT-ready normalized conversations; sft_rendered/train contains their rendered text. Matching part-*.parquet files represent one atomic collection block. snapshot-manifest.json records the exact row counts. Browse each representation using its dataset subset:… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/GLM-5.3-RLVR1-Training-Rollouts-2026.09.25.text10K<n<100K0 likes1.3k downloads13d agoHugging Faceopen-athena /Snowball-67B-A2B-Mixed-RLVR-Experiment-Artifacts Snowball 67B-A2B RL artifact release 2026 mixed-domain RLVR campaign This release also contains the complete releasable record of the September 2026 Snowball mixed-domain RLVR campaign. It covers the September 11 synchronous and bounded-staleness asynchronous RLVR1→RLVR2 lineages and the 5.7T Agentic-start RLVR1 lineage. All training arms are terminal. The final campaign figure, trace audit, canonical configs, timing reports, retained traces, and operational… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/Snowball-67B-A2B-Mixed-RLVR-Experiment-Artifacts.imagen<1K0 likes786 downloads18d agoHugging Faceallenai /RLVR-GSM-MATH-IF-Mixed-Constraints GSM/MATH/IF Data - RLVR Formatted Note that this collection is licensed under ODC-BY-1.0 license; different licenses apply to subsets of the data. This dataset contains data formatted for use with open-instruct - specifically reinforcement learning with verifiable rewards. It was used to train the final Tulu 3 models with RL, and contains the following subsets: GSM8k (7,473 samples): The GSM8k train set formatted for use with RLVR and open-instruct. MIT License. MATH (7,500… See the full description on the dataset page: https://huggingface.co/datasets/allenai/RLVR-GSM-MATH-IF-Mixed-Constraints.text10K<n<100K32 likes704 downloads2y agoHugging Face