datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quantqa
QuantQA: Quantitative Finance Interview Questions
QuantQA is a curated dataset of 519 interview questions sourced from leading quantitative trading firms including Jane Street, Citadel, Two Sigma, Optiver, and SIG, in collaboration with CoachQuant.
Topic Distribution
Topic
Coverage
Probability
67%
Combinatorics
22%
Expected Value
21%
Conditional Probability
14%
Game Theory
11%
Note: Questions may cover multiple topics
Training Results… See the full description on the dataset page: https://huggingface.co/datasets/ReinforceNow/quantqa.reinforcement_learningReinforced-IR-synthetic
Introduction
Synthetic data for Reinforced IR.
Load Dataset
An example to load the dataset:
import datasets
# load dataset
dataset = datasets.load_dataset(
"cfli/Reinforced-IR-synthetic",
'dbpedia-entity',
split='generator'
)
# print one sample
print(dataset[0])
reinforce-learning
DAPO-RL-Instruct Dataset
A high-quality instruction-following dataset derived from the open-source technical report “DAPO: An Open-Source LLM Reinforcement Learning System at Scale” (arXiv:2503.14476, March 2025). This dataset captures key concepts, training strategies, and system design principles described in the paper, reformatted as instruction–response pairs suitable for fine-tuning or evaluating large language models (LLMs) in reinforcement learning (RL) contexts.… See the full description on the dataset page: https://huggingface.co/datasets/amishor/reinforce-learning.rl-single-math-reasoning
RL-Single: Test Dataset for ReinforceNow
This is a test dataset compatible with the ReinforceNow Platform containing only 92 sample entries for demonstration and testing purposes. It is designed to help users get started with RLHF training for mathematical reasoning models.
Note: This is not a production dataset. It contains a small subset of math problems for testing the ReinforceNow CLI and verifying your training pipeline works correctly before scaling up.
Data… See the full description on the dataset page: https://huggingface.co/datasets/ReinforceNow/rl-single-math-reasoning.SultanR__SmolTulu-1.7b-Reinforced-details
Dataset Card for Evaluation run of SultanR/SmolTulu-1.7b-Reinforced
Dataset automatically created during the evaluation run of model SultanR/SmolTulu-1.7b-Reinforced
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/SultanR__SmolTulu-1.7b-Reinforced-details.v3-reinforcement-engiter_1_reinforce_baseline_per_sample_200epoch_strong_init_step_150iter_1_reinforce_baseline_per_sample_200epoch_strong_init_step_150_filteredgomodel-tool-reinforced
GoModel Tool-Reinforced Dataset
A tool-reinforced instruction dataset for training Go coding assistants. It combines Go code tasks with examples that teach a model when and how to use repository and Go development tools.
Tools
read_file
search_code
list_functions
get_imports
write_file
go_vet
go_build
Dataset statistics
Split
File
Examples
Bytes
train
train_tool_reinforced.jsonl
1,732
3,726,144
eval
eval_tool_reinforced.jsonl
193
408… See the full description on the dataset page: https://huggingface.co/datasets/mencosk/gomodel-tool-reinforced.reinforcement_agent_crash_pairs_v6reinforcementnlquad-reinforced
