anthonym21/rlcd-decision-v1
RLCD Decision Dataset (v1) Typed decision questions for training and evaluating models that answer with a calibrated probability distribution over declared options instead of generated text. Built for RLCD — reinforcement learning for calibrated decisions (see also the trained export anthonym21/qwen3-0.6b-rlcd-decision). Every row is one typed question over a context: a choice question over unordered options, a score question over ordered levels, or a noul (yes/no) question. The… See the full description on the dataset page: https://huggingface.co/datasets/anthonym21/rlcd-decision-v1.
RLCD Decision Dataset (v1)
Typed decision questions for training and evaluating models that answer with a calibrated probability distribution over declared options instead of generated text. Built for RLCD — reinforcement learning for calibrated decisions (see also the trained export anthonym21/qwen3-0.6b-rlcd-decision).
Every row is one typed question over a context: a choice question over unordered options, a score question over ordered levels, or a noul (yes/no) question. The RLCD training loop treats each row as a bandit arm — the environment reveals only whether the sampled option was correct, never the label — and the reward r = c - p_a pushes the sampled probability toward calibration.
Schema
One JSON object per line; CRLF line endings preserved from the release build.
Splits
Train primitives: 34,667 choice / 18,667 score / 10,666 noul. Validation and test: 4,334 / 2,333 / 1,333 each. stats.json (included) has the full per-source breakdown.
Provenance
Built with rlcd.data build --per-source 8000 --seed 0 on 2026-09-17 from the code at commit 57a179b. These are the exact files used for every published run in the repo's results tables; a later rebuild produced different bytes, so use these for comparable runs.
Source datasets (question converters in `rlcd/data.py`):
The triage rows are fully synthetic (original work). All other rows are transformed subsets of the upstream datasets above; credit for the underlying texts belongs to the upstream sources, and their terms (some share-alike) apply to those portions. This repo is distributed as CC-BY-4.0 as a convenience tag; if your use is sensitive to the upstream terms, follow the links and check them.
Checksums (md5)
train.jsonl cdfee4c9792751cf5b22668eb3f9dc33
val.jsonl 4b95911ffc76ed1789f7989f623d0a5a
test.jsonl 15c33b165d70639d8bd7d23d624908f4Intended use
Research on decision-making LLMs under bandit/outcome-only feedback, probability calibration (ECE, Brier), and confidence-aware classification. Not a benchmark of world knowledge: every split is in-distribution for the sources above and the questions are template-generated.
Citation
@software{maio2026eve_rlcd,
title = {eve-rlcd: reinforcement learning for calibrated decisions},
author = {Anthony Maio},
url = {https://github.com/anthony-maio/eve-rlcd},
year = {2026}
}