sumleo/RLCDAlignBench
RLCDAlignBench Paper: Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures (arXiv:2609.29429) Code: github.com/sumleo/RLCDAlignBench · Project page: sumleo.github.io/RLCDAlignBench RLCDAlignBench measures whether a detector can tell when a language model's output is an alignment failure. It has 44 benchmarks across ten failure types and five target models (Qwen3.5-2B, Phi-4-mini, Gemma-2-2B, Llama-3.2-3B, Olmo-3-7B), for… See the full description on the dataset page: https://huggingface.co/datasets/sumleo/RLCDAlignBench.
This repository belongs to sumleo on Hugging Face.
Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
