logicBombExe/double_logging_system_to_detect_an_unaligned_llm
Double Logging to Detect an Unaligned LLM Can comparing a model's own action log with an automatic system log detect harmful behaviour? The detector sees only the two logs and is scored on detection rate and false alarms. Status: in design. No data yet. Code GitHub Earlier study INSIDER_LLM_DETECTION_BENCHMARK (archived)
043
Double Logging to Detect an Unaligned LLM
Can comparing a model's own action log with an automatic system log detect harmful behaviour? The detector sees only the two logs and is scored on detection rate and false alarms.
Status: in design. No data yet.
