datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jev-playground-rlcd-v0
openjev-rlcd-v0
Synthetic calibrated-decision dataset for training Jev-style System One classifiers plan). States are generated programmatically per domain; typed questions follow the openjev /v1/systemone contract; reference distributions are exact by construction, the property a proper-scoring-rule / calibration objective needs.
Format
JSONL, one state per line:
{
"domain": "support_tickets",
"state": "...",
"questions": {
"urgent": {"type": "noul"… See the full description on the dataset page: https://huggingface.co/datasets/soyrsoyr/jev-playground-rlcd-v0.RLCDAlignBench
RLCDAlignBench
Paper: Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures (arXiv:2609.29429)
Code: github.com/sumleo/RLCDAlignBench · Project page: sumleo.github.io/RLCDAlignBench
RLCDAlignBench measures whether a detector can tell when a language model's output is an alignment failure.
It has 44 benchmarks across ten failure types and five target models (Qwen3.5-2B, Phi-4-mini, Gemma-2-2B, Llama-3.2-3B, Olmo-3-7B), for… See the full description on the dataset page: https://huggingface.co/datasets/sumleo/RLCDAlignBench.rlcd-decision-v1
RLCD Decision Dataset (v1)
Typed decision questions for training and evaluating models that answer with a calibrated probability distribution over declared options instead of generated text. Built for RLCD — reinforcement learning for calibrated decisions (see also the trained export anthonym21/qwen3-0.6b-rlcd-decision).
Every row is one typed question over a context: a choice question over unordered options, a score question over ordered levels, or a noul (yes/no) question. The… See the full description on the dataset page: https://huggingface.co/datasets/anthonym21/rlcd-decision-v1.rlcd
Dataset Card for "rlcd"
More Information needed
eve-rlcd-runsRLCD-generated-preference-data-split
Dataset Card for "RLCD-generated-preference-data-split"
More Information needed
Qwev-RLCD-DataRLCD-SFT-conversationsMIT License
Dataset Card for "RLCD-SFT-conversations"
More Information needed
rlcd-decision-atlas
rlcd-decision-atlas — a typed-decision dataset
This is a dataset for training and evaluating models that make decisions rather than write prose. Every example is one situation, one question, and a short list of declared options — and exactly one of those options is correct. The model's whole job is to put a probability on each option and commit to one. Nothing here asks for free text.
It is assembled from 114 public sources plus one synthetic generator, each converted into the… See the full description on the dataset page: https://huggingface.co/datasets/goutam/rlcd-decision-atlas.RLCD-generated-preference-data
Dataset Card for "RLCD-generated-preference-data"
More Information needed
C2-Nemo-Delta-RLCD-PreferenceMy attempt at cheesing the lack of roleplay preference data through a combination of delta tuning and RLCD.
Unfortunately there are several notable issues, namely: Unbalanced em-dash usage, Nemo 253B and 4B not having exactly the same sytle, RLCD being iffy as Nemo 253B and 4B don't have a great idea of what constitutes bad and good rp, and the 4B being a bit too not atrociously bad enough for delta tuning.
Obviously, as this is lightly filtered C2 there's a good amount of nsfl and nsfw.
RLCD-SFT-dataset
Dataset Card for "RLCD-SFT-dataset"
More Information needed
