Team Ai
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AdityaaXD /Multi-Agent_Reinforcement_Learning_Trading_System_Data 📊 Multi-Agent RL Trading System - Dataset This dataset contains historical OHLCV (Open, High, Low, Close, Volume) data for AAPL, MSFT, and GOOGL, pre-processed for Reinforcement Learning based trading systems. 📁 Dataset Content The dataset consists of CSV files downloaded via yfinance: AAPL.csv: Apple Inc. daily data (Jan 2018 - Dec 2024). MSFT.csv: Microsoft Corp. daily data (Jan 2018 - Dec 2024). GOOGL.csv: Alphabet Inc. daily data (Jan 2018 - Dec 2024). 📝… See the full description on the dataset page: https://huggingface.co/datasets/AdityaaXD/Multi-Agent_Reinforcement_Learning_Trading_System_Data.tabulartime-series-forecasting1K<n<10K9 likes254 downloads8mo agoHugging Face02sanjaydoss /Multi-Agent_Reinforcement_Learning_Trading_System_Data 📊 Multi-Agent RL Trading System - Dataset This dataset contains historical OHLCV (Open, High, Low, Close, Volume) data for AAPL, MSFT, and GOOGL, pre-processed for Reinforcement Learning based trading systems. 📁 Dataset Content The dataset consists of CSV files downloaded via yfinance: AAPL.csv: Apple Inc. daily data (Jan 2018 - Dec 2024). MSFT.csv: Microsoft Corp. daily data (Jan 2018 - Dec 2024). GOOGL.csv: Alphabet Inc. daily data (Jan 2018 - Dec 2024).… See the full description on the dataset page: https://huggingface.co/datasets/sanjaydoss/Multi-Agent_Reinforcement_Learning_Trading_System_Data.tabulartime-series-forecasting1K<n<10K16 likes187 downloads1mo agoHugging Face03open-source-metrics /reinforcement-learning-checkpoint-downloadstextn<1K5 likes155 downloads4y agoHugging Face04n1ghtf4l1 /Agentic-Diagnostic-Reasoning-with-Multimodal-SLMs-via-Reinforcement-Learning12 likes147 downloads11mo agoHugging Face05introvoyz041 /reinforcement-learningimagen<1K5 likes142 downloads1y agoHugging Face06Srishti280992 /repro-exact-unlearning-in-reinforcement-learning-traces Agent traces Agent sessions published from a Trackio Logbook. 7 likes134 downloads2mo agoHugging Face07James4Ever0 /computer_agent_reinforcement_learning_trajectory_seagent_ai_assistant_tools_agent_mcp4 likes126 downloads1y agoHugging Face08prism-drift /reinforcement-learning-results Reinforcement Learning Results This repository contains the latest result artifacts for the prompts-repaired-v1 reinforcement-learning experiments in ai-pref-drift. Included scope Model sizes: Qwen3.5 4B and 9B. Model states: M0-v4 and DPO, GRPO, and PPO Step 125 (seed 0). Actual pairwise-preference batteries: coding-task preference and value preference, each in thinking and non-thinking modes. Anticipation batteries: coding-task anticipation and value… See the full description on the dataset page: https://huggingface.co/datasets/prism-drift/reinforcement-learning-results.0 likes111 downloads29d agoHugging Face09youinwww /reinforcement_learningtextn<1K5 likes105 downloads1y agoHugging Face10SeeWye /NFA_OCR_reinforcement_learning_format_TEST5image1K<n<10K4 likes82 downloads1mo agoHugging Face11debajyotidasgupta /repro-contextual-rollout-bandits-for-reinforcement-learning-with-verifiable-rewards-artifacts Reproduction: Contextual Rollout Bandits for RLVR (ICML 2026, #985) Independent reproduction of "Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards" (Lu, Wang, Chai, Yin, Lin, Chen, Luo, Zhuang, Ban, Wang) — OpenReview weMYE1B16x, arXiv 2602.08499. Part of the Hugging Face × AlphaXiv ICML-2026 reproduction challenge. Official code: github.com/lxd99/CBS_public (verl 0.5.x fork). What CBS is The paper reframes rollout scheduling in RLVR as… See the full description on the dataset page: https://huggingface.co/datasets/debajyotidasgupta/repro-contextual-rollout-bandits-for-reinforcement-learning-with-verifiable-rewards-artifacts.1 likes81 downloads3mo agoHugging Face12Gene829 /gene-reinforcement-learning-instruct reinforcement-learning-instruct v4 Gate-passed instruction data for reinforcement-learning — published when 50 fresh examples cleared the quality bar Kind: synthetic Domain: reinforcement-learning Records: 198 Created: 2026-06-19T23:14:20+00:00 SHA-256: 3393dfd6bd9adc38414885ee2f5ac35f6ce60b4c57a98c3e3f2ca78e574f1469 Pipeline: v2.0.0 Filters: {"min_quality": 0.55, "limit": 1000, "source": null, "backend": "llama", "min_judge": 0.7} Generated by:… See the full description on the dataset page: https://huggingface.co/datasets/Gene829/gene-reinforcement-learning-instruct.text-generationn<1K4 likes76 downloads4mo agoHugging Face13DBbun /motoman-up6-cq-lambda-reinforcement-learning_v1.0 CQ(λ) Bag-Shaking Dataset: Human-in-the-Loop Reinforcement Learning Dataset Description This dataset contains synthetic training data comparing standard Q-learning with eligibility traces [Q(λ)] against Cooperative Q-learning [CQ(λ)], a human-in-the-loop reinforcement learning algorithm. The data simulates a robotic "bag-shaking" task where an agent must extract knotted objects from a bag through strategic shaking motions. Dataset Summary Task: Bag-shaking… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/motoman-up6-cq-lambda-reinforcement-learning_v1.0.imagen<1K2 likes69 downloads7mo agoHugging Face14SeeWye /NFA_OCR_reinforcement_learning_format_TEST6image1K<n<10K0 likes35 downloads1mo agoHugging Face15D21W12 /Diffusion-Deep-Reinforcement-Learning3 likes24 downloads3mo agoHugging Face16SeeWye /NFA_OCR_reinforcement_learning_format_TEST4image1K<n<10K0 likes17 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.