Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hendrydong /reinforce-ada-raw-eval Reinforce-Ada Raw Eval Raw evaluation artifacts organized by experiment / dataset / step. Included files when present: merged_data.jsonl pass_at_k.json record.txt Experiments: grpo_n8, grpo_n16, grpo_n32, reinforce_ada_n8, reinforce_ada_n8_normstdtrue Datasets: math500, minerva_math, olympiadbench, aime_hmmt_brumo_cmimc_amc23 text0 likes20k downloads7mo agoHugging Face02trangdo43323 /vehicular-traffic-light-reinforcement0 likes14k downloads55m agoHugging Face03hendrydong /reinforce-ada-n16 Reinforce-Ada n16 Eval Artifacts Evaluation artifacts for reinforce_ada_n16_mr4_rr16. Included eval suffixes: t07_k64: steps [50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000] t10_k64: steps [50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000] Each step / dataset directory contains: merged_data.jsonl pass_at_k.json record.txt text0 likes268 downloads7mo agoHugging Face04AdityaaXD /Multi-Agent_Reinforcement_Learning_Trading_System_Data 📊 Multi-Agent RL Trading System - Dataset This dataset contains historical OHLCV (Open, High, Low, Close, Volume) data for AAPL, MSFT, and GOOGL, pre-processed for Reinforcement Learning based trading systems. 📁 Dataset Content The dataset consists of CSV files downloaded via yfinance: AAPL.csv: Apple Inc. daily data (Jan 2018 - Dec 2024). MSFT.csv: Microsoft Corp. daily data (Jan 2018 - Dec 2024). GOOGL.csv: Alphabet Inc. daily data (Jan 2018 - Dec 2024). 📝… See the full description on the dataset page: https://huggingface.co/datasets/AdityaaXD/Multi-Agent_Reinforcement_Learning_Trading_System_Data.tabulartime-series-forecasting1K<n<10K14 likes260 downloads8mo agoHugging Face05hendrydong /reinforce-ada-eval-t07 Reinforce-Ada Eval t=0.7 Evaluation artifacts for temperature 0.7, K=64, PASS_K_MAX=64. Experiments: grpo_n8 reinforce_ada_n8 Per step / dataset directory contains: merged_data.jsonl pass_at_k.json record.txt Original outputs were written under global_step_xxx/merged/weqweasdas/*__t07_k64 and remapped into this repo. text0 likes253 downloads7mo agoHugging Face06poopoo3882 /Reinforced_Reasoning_for_Embodied_Planningimage1K<n<10K0 likes246 downloads1y agoHugging Face07sanjaydoss /Multi-Agent_Reinforcement_Learning_Trading_System_Data 📊 Multi-Agent RL Trading System - Dataset This dataset contains historical OHLCV (Open, High, Low, Close, Volume) data for AAPL, MSFT, and GOOGL, pre-processed for Reinforcement Learning based trading systems. 📁 Dataset Content The dataset consists of CSV files downloaded via yfinance: AAPL.csv: Apple Inc. daily data (Jan 2018 - Dec 2024). MSFT.csv: Microsoft Corp. daily data (Jan 2018 - Dec 2024). GOOGL.csv: Alphabet Inc. daily data (Jan 2018 - Dec 2024).… See the full description on the dataset page: https://huggingface.co/datasets/sanjaydoss/Multi-Agent_Reinforcement_Learning_Trading_System_Data.tabulartime-series-forecasting1K<n<10K20 likes221 downloads1mo agoHugging Face08open-source-metrics /reinforcement-learning-checkpoint-downloadstextn<1K9 likes179 downloads4y agoHugging Face09n1ghtf4l1 /Agentic-Diagnostic-Reasoning-with-Multimodal-SLMs-via-Reinforcement-Learning15 likes175 downloads11mo agoHugging Face10introvoyz041 /reinforcement-learningimagen<1K6 likes164 downloads1y agoHugging Face11Srishti280992 /repro-exact-unlearning-in-reinforcement-learning-traces Agent traces Agent sessions published from a Trackio Logbook. 11 likes164 downloads2mo agoHugging Face12James4Ever0 /computer_agent_reinforcement_learning_trajectory_seagent_ai_assistant_tools_agent_mcp6 likes153 downloads1y agoHugging Face13hendrydong /reinforce-ada-eval-public Reinforce-Ada Eval Public Public evaluation and training-metric artifacts for five RLHFlow experiments. Experiments grpo_n8 grpo_n16 grpo_n32 reinforce_ada_n8 reinforce_ada_n8_normstdtrue Evaluation Tables These are the main benchmark tables. Each CSV contains all 5 experiments, all checkpoint steps, and pass@k for k=1..64. reinforce_ada_math500_passk_all5_20260312.csv reinforce_ada_olympiadbench_passk_all5_20260312.csv… See the full description on the dataset page: https://huggingface.co/datasets/hendrydong/reinforce-ada-eval-public.documentn<1K0 likes118 downloads6mo agoHugging Face14ReinforceNow /quantqa QuantQA: Quantitative Finance Interview Questions QuantQA is a curated dataset of 519 interview questions sourced from leading quantitative trading firms including Jane Street, Citadel, Two Sigma, Optiver, and SIG, in collaboration with CoachQuant. Topic Distribution Topic Coverage Probability 67% Combinatorics 22% Expected Value 21% Conditional Probability 14% Game Theory 11% Note: Questions may cover multiple topics Training Results… See the full description on the dataset page: https://huggingface.co/datasets/ReinforceNow/quantqa.documentquestion-answeringn<1K0 likes114 downloads9mo agoHugging Face15youinwww /reinforcement_learningtextn<1K7 likes108 downloads1y agoHugging Face16Gene829 /gene-reinforcement-learning-instruct reinforcement-learning-instruct v4 Gate-passed instruction data for reinforcement-learning — published when 50 fresh examples cleared the quality bar Kind: synthetic Domain: reinforcement-learning Records: 198 Created: 2026-06-19T23:14:20+00:00 SHA-256: 3393dfd6bd9adc38414885ee2f5ac35f6ce60b4c57a98c3e3f2ca78e574f1469 Pipeline: v2.0.0 Filters: {"min_quality": 0.55, "limit": 1000, "source": null, "backend": "llama", "min_judge": 0.7} Generated by:… See the full description on the dataset page: https://huggingface.co/datasets/Gene829/gene-reinforcement-learning-instruct.text-generationn<1K6 likes98 downloads4mo agoHugging Face17DBbun /motoman-up6-cq-lambda-reinforcement-learning_v1.0 CQ(λ) Bag-Shaking Dataset: Human-in-the-Loop Reinforcement Learning Dataset Description This dataset contains synthetic training data comparing standard Q-learning with eligibility traces [Q(λ)] against Cooperative Q-learning [CQ(λ)], a human-in-the-loop reinforcement learning algorithm. The data simulates a robotic "bag-shaking" task where an agent must extract knotted objects from a bag through strategic shaking motions. Dataset Summary Task: Bag-shaking… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/motoman-up6-cq-lambda-reinforcement-learning_v1.0.imagen<1K3 likes85 downloads7mo agoHugging Face18cfli /Reinforced-IR-synthetic Introduction Synthetic data for Reinforced IR. Load Dataset An example to load the dataset: import datasets # load dataset dataset = datasets.load_dataset( "cfli/Reinforced-IR-synthetic", 'dbpedia-entity', split='generator' ) # print one sample print(dataset[0]) text100K<n<1M1 likes77 downloads1y agoHugging Face19SeeWye /NFA_OCR_reinforcement_learning_format_TEST5image1K<n<10K4 likes61 downloads1mo agoHugging Face20debajyotidasgupta /repro-contextual-rollout-bandits-for-reinforcement-learning-with-verifiable-rewards-artifacts Reproduction: Contextual Rollout Bandits for RLVR (ICML 2026, #985) Independent reproduction of "Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards" (Lu, Wang, Chai, Yin, Lin, Chen, Luo, Zhuang, Ban, Wang) — OpenReview weMYE1B16x, arXiv 2602.08499. Part of the Hugging Face × AlphaXiv ICML-2026 reproduction challenge. Official code: github.com/lxd99/CBS_public (verl 0.5.x fork). What CBS is The paper reframes rollout scheduling in RLVR as… See the full description on the dataset page: https://huggingface.co/datasets/debajyotidasgupta/repro-contextual-rollout-bandits-for-reinforcement-learning-with-verifiable-rewards-artifacts.1 likes60 downloads3mo agoHugging Face21prism-drift /reinforcement-learning-results Reinforcement Learning Results This repository contains the latest result artifacts for the prompts-repaired-v1 reinforcement-learning experiments in ai-pref-drift. Included scope Model sizes: Qwen3.5 4B and 9B. Model states: M0-v4 and DPO, GRPO, and PPO Step 125 (seed 0). Actual pairwise-preference batteries: coding-task preference and value preference, each in thinking and non-thinking modes. Anticipation batteries: coding-task anticipation and value… See the full description on the dataset page: https://huggingface.co/datasets/prism-drift/reinforcement-learning-results.0 likes54 downloads1mo agoHugging Face22D21W12 /Diffusion-Deep-Reinforcement-Learning4 likes45 downloads4mo agoHugging Face23amishor /reinforce-learning DAPO-RL-Instruct Dataset A high-quality instruction-following dataset derived from the open-source technical report “DAPO: An Open-Source LLM Reinforcement Learning System at Scale” (arXiv:2503.14476, March 2025). This dataset captures key concepts, training strategies, and system design principles described in the paper, reformatted as instruction–response pairs suitable for fine-tuning or evaluating large language models (LLMs) in reinforcement learning (RL) contexts.… See the full description on the dataset page: https://huggingface.co/datasets/amishor/reinforce-learning.textn<1K0 likes40 downloads1y agoHugging Face24gprbase /ds021-3d-grid-reinforced-concretegated 3D Grid on Reinforced Concrete: Rebar Mesh and Diagonal Electrical Cable GPRbase ds021. Raw Ground Penetrating Radar (GPR) field data recorded with a GSSI 2700 MHz system on a reinforced concrete slab: a 3D grid of 38 profiles over a 1.2 × 0.6 m area. The data are free to use for education, training and research under CC BY-NC-SA 4.0. Why this dataset is useful The slab contains a reinforcement mesh and an electrical cable that crosses the survey area diagonally.… See the full description on the dataset page: https://huggingface.co/datasets/gprbase/ds021-3d-grid-reinforced-concrete.tabularn<1K0 likes38 downloads1d agoHugging Face25RLHFlow /reinforce_ada_hard_prompt_1-5btext10K<n<100K0 likes37 downloads1y agoHugging Face26reinforcelabs /STEM-QA-EVAL STEM-QA-EVAL Composition Subject Description chem Chemistry — physical, organic, and inorganic chemistry questions. math Mathematics — algebra, geometry, trigonometry, set theory, and combinatorics. phys Physics — mechanics, electromagnetism, optics, and modern physics. resn Reasoning — verbal, logical, and comprehension-style questions. Schema Column Type Description data_id string Stable identifier (stem-NNNN… See the full description on the dataset page: https://huggingface.co/datasets/reinforcelabs/STEM-QA-EVAL.texttext-generationn<1K0 likes35 downloads4mo agoHugging Face27SeeWye /NFA_OCR_reinforcement_learning_format_TEST6image1K<n<10K0 likes34 downloads1mo agoHugging Face28weqweasdas /qwn_reinforce_rej320_minverva_mathtextn<1K0 likes29 downloads1y agoHugging Face29ReinforceNow /rl-single-math-reasoning RL-Single: Test Dataset for ReinforceNow This is a test dataset compatible with the ReinforceNow Platform containing only 92 sample entries for demonstration and testing purposes. It is designed to help users get started with RLHF training for mathematical reasoning models. Note: This is not a production dataset. It contains a small subset of math problems for testing the ReinforceNow CLI and verifying your training pipeline works correctly before scaling up. Data… See the full description on the dataset page: https://huggingface.co/datasets/ReinforceNow/rl-single-math-reasoning.textn<1K0 likes28 downloads9mo agoHugging Face30RLHFlow /reinforce_ada_hard_prompt Selected hard prompts used to train Qwen2.5-Math-7B and Qwen3-4B-Instruct-2507. text10K<n<100K2 likes27 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.