Team Ai
Datasetpublicgated

sumleo/RLCDAlignBench

RLCDAlignBench Paper: Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures (arXiv:2609.29429) Code: github.com/sumleo/RLCDAlignBench · Project page: sumleo.github.io/RLCDAlignBench RLCDAlignBench measures whether a detector can tell when a language model's output is an alignment failure. It has 44 benchmarks across ten failure types and five target models (Qwen3.5-2B, Phi-4-mini, Gemma-2-2B, Llama-3.2-3B, Olmo-3-7B), for… See the full description on the dataset page: https://huggingface.co/datasets/sumleo/RLCDAlignBench.

sourceHugging Facecc-by-nc-4.0updated 16d agoView on Hugging Face
1likes142downloads
file_map.csvDownload Raw Back to root

This repository is gated, so its file contents are only served once you have accepted the publisher's terms at Hugging Face. Open it at the source above.