Team Ai
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ReinforceNow /quantqa QuantQA: Quantitative Finance Interview Questions QuantQA is a curated dataset of 519 interview questions sourced from leading quantitative trading firms including Jane Street, Citadel, Two Sigma, Optiver, and SIG, in collaboration with CoachQuant. Topic Distribution Topic Coverage Probability 67% Combinatorics 22% Expected Value 21% Conditional Probability 14% Game Theory 11% Note: Questions may cover multiple topics Training Results… See the full description on the dataset page: https://huggingface.co/datasets/ReinforceNow/quantqa.documentquestion-answeringn<1K0 likes114 downloads9mo agoHugging Face02youinwww /reinforcement_learningtextn<1K7 likes108 downloads1y agoHugging Face03cfli /Reinforced-IR-synthetic Introduction Synthetic data for Reinforced IR. Load Dataset An example to load the dataset: import datasets # load dataset dataset = datasets.load_dataset( "cfli/Reinforced-IR-synthetic", 'dbpedia-entity', split='generator' ) # print one sample print(dataset[0]) text100K<n<1M1 likes77 downloads1y agoHugging Face04amishor /reinforce-learning DAPO-RL-Instruct Dataset A high-quality instruction-following dataset derived from the open-source technical report “DAPO: An Open-Source LLM Reinforcement Learning System at Scale” (arXiv:2503.14476, March 2025). This dataset captures key concepts, training strategies, and system design principles described in the paper, reformatted as instruction–response pairs suitable for fine-tuning or evaluating large language models (LLMs) in reinforcement learning (RL) contexts.… See the full description on the dataset page: https://huggingface.co/datasets/amishor/reinforce-learning.textn<1K0 likes40 downloads1y agoHugging Face05ReinforceNow /rl-single-math-reasoning RL-Single: Test Dataset for ReinforceNow This is a test dataset compatible with the ReinforceNow Platform containing only 92 sample entries for demonstration and testing purposes. It is designed to help users get started with RLHF training for mathematical reasoning models. Note: This is not a production dataset. It contains a small subset of math problems for testing the ReinforceNow CLI and verifying your training pipeline works correctly before scaling up. Data… See the full description on the dataset page: https://huggingface.co/datasets/ReinforceNow/rl-single-math-reasoning.textn<1K0 likes28 downloads9mo agoHugging Face06open-llm-leaderboard /SultanR__SmolTulu-1.7b-Reinforced-detailsgated Dataset Card for Evaluation run of SultanR/SmolTulu-1.7b-Reinforced Dataset automatically created during the evaluation run of model SultanR/SmolTulu-1.7b-Reinforced The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/SultanR__SmolTulu-1.7b-Reinforced-details.tabular10K<n<100K0 likes25 downloads2y agoHugging Face07madihalim /v3-reinforcement-engtextn<1K0 likes18 downloads11mo agoHugging Face08TongZheng1999 /iter_1_reinforce_baseline_per_sample_200epoch_strong_init_step_150text10K<n<100K0 likes16 downloads6mo agoHugging Face09TongZheng1999 /iter_1_reinforce_baseline_per_sample_200epoch_strong_init_step_150_filteredtext10K<n<100K0 likes9 downloads6mo agoHugging Face10mencosk /gomodel-tool-reinforced GoModel Tool-Reinforced Dataset A tool-reinforced instruction dataset for training Go coding assistants. It combines Go code tasks with examples that teach a model when and how to use repository and Go development tools. Tools read_file search_code list_functions get_imports write_file go_vet go_build Dataset statistics Split File Examples Bytes train train_tool_reinforced.jsonl 1,732 3,726,144 eval eval_tool_reinforced.jsonl 193 408… See the full description on the dataset page: https://huggingface.co/datasets/mencosk/gomodel-tool-reinforced.texttext-generation1K<n<10K0 likes8 downloads2mo agoHugging Face11Maitreyajayaraj /reinforcement_agent_crash_pairs_v6textn<1K0 likes4 downloads6mo agoHugging Face12jennalee1385 /reinforcementtextn<1K0 likes3 downloads2y agoHugging Face13Yoga26 /nlquad-reinforcedgatedtext1K<n<10K0 likes2 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.