datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RLCDAlignBench
RLCDAlignBench
Paper: Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures (arXiv:2609.29429)
Code: github.com/sumleo/RLCDAlignBench · Project page: sumleo.github.io/RLCDAlignBench
RLCDAlignBench measures whether a detector can tell when a language model's output is an alignment failure.
It has 44 benchmarks across ten failure types and five target models (Qwen3.5-2B, Phi-4-mini, Gemma-2-2B, Llama-3.2-3B, Olmo-3-7B), for… See the full description on the dataset page: https://huggingface.co/datasets/sumleo/RLCDAlignBench.RLCD-generated-preference-data-split
Dataset Card for "RLCD-generated-preference-data-split"
More Information needed
RLCD-generated-preference-data
Dataset Card for "RLCD-generated-preference-data"
More Information needed
