Team Ai
Datasetpublic

anthonym21/rlcd-decision-v1

RLCD Decision Dataset (v1) Typed decision questions for training and evaluating models that answer with a calibrated probability distribution over declared options instead of generated text. Built for RLCD — reinforcement learning for calibrated decisions (see also the trained export anthonym21/qwen3-0.6b-rlcd-decision). Every row is one typed question over a context: a choice question over unordered options, a score question over ordered levels, or a noul (yes/no) question. The… See the full description on the dataset page: https://huggingface.co/datasets/anthonym21/rlcd-decision-v1.

sourceHugging Facecc-by-4.0updated 12d agoView on Hugging Face
0likes103downloads
Dataset Card

RLCD Decision Dataset (v1)

Typed decision questions for training and evaluating models that answer with a calibrated probability distribution over declared options instead of generated text. Built for RLCD — reinforcement learning for calibrated decisions (see also the trained export anthonym21/qwen3-0.6b-rlcd-decision).

Every row is one typed question over a context: a choice question over unordered options, a score question over ordered levels, or a noul (yes/no) question. The RLCD training loop treats each row as a bandit arm — the environment reveals only whether the sampled option was correct, never the label — and the reward r = c - p_a pushes the sampled probability toward calibration.

Schema

fieldtypemeaning
primitivestringchoice (unordered), score (ordered levels), noul (yes/no)
contextstringthe state / passage the question is about
questionstringthe question, in words
choiceslist[string]the declared options (2–26)
orderedboolwhether the options are ordered levels (score questions)
answerintindex into choices of the correct option
sourcestringwhich of the 8 sources the row came from
idstringstable row id <source>-<n>[-<aspect>]

One JSON object per line; CRLF line endings preserved from the release build.

Splits

splitrowssourcesper-source
train64,00088,000
validation8,00081,000
test8,00081,000

Train primitives: 34,667 choice / 18,667 score / 10,666 noul. Validation and test: 4,334 / 2,333 / 1,333 each. stats.json (included) has the full per-source breakdown.

Provenance

Built with rlcd.data build --per-source 8000 --seed 0 on 2026-09-17 from the code at commit 57a179b. These are the exact files used for every published run in the repo's results tables; a later rebuild produced different bytes, so use these for comparable runs.

Source datasets (question converters in `rlcd/data.py`):

sourceupstream datasetupstream licensetransformation
bitextbitext/Bitext-customer-support-llm-chatbot-training-datasetCDLA-Sharing-1.0intent classification over sampled label subsets
banking77legacy-datasets/banking77 (mirror of PolyAI/banking77)CC-BY-4.0intent classification over sampled label subsets
ag_newsfancyzhx/ag_newsunspecified upstreamtopic classification, fixed 4 options
mnlinyu-mll/multi_nliCC-BY-3.0 / CC-BY-SA-3.0 (mixed)premise–hypothesis relation, fixed 3 options
sst5SetFit/sst5unspecified upstream (SST derivatives)5-level sentiment, ordered
yelpYelp/yelp_review_fullYelp Dataset terms5-level star rating, ordered, text truncated to 1,500 chars
boolqgoogle/boolqCC-BY-SA-3.0yes/no question over passage
triagesynthetic generator in rlcd/data.py (original)—enterprise support tickets with department / priority / escalation questions

The triage rows are fully synthetic (original work). All other rows are transformed subsets of the upstream datasets above; credit for the underlying texts belongs to the upstream sources, and their terms (some share-alike) apply to those portions. This repo is distributed as CC-BY-4.0 as a convenience tag; if your use is sensitive to the upstream terms, follow the links and check them.

Checksums (md5)

train.jsonl  cdfee4c9792751cf5b22668eb3f9dc33
val.jsonl    4b95911ffc76ed1789f7989f623d0a5a
test.jsonl   15c33b165d70639d8bd7d23d624908f4

Intended use

Research on decision-making LLMs under bandit/outcome-only feedback, probability calibration (ECE, Brier), and confidence-aware classification. Not a benchmark of world knowledge: every split is in-distribution for the sources above and the questions are template-generated.

Citation

bibtex
@software{maio2026eve_rlcd,
  title = {eve-rlcd: reinforcement learning for calibrated decisions},
  author = {Anthony Maio},
  url = {https://github.com/anthony-maio/eve-rlcd},
  year = {2026}
}