Team Ai
Datasetpublic

logicBombExe/double_logging_system_to_detect_an_unaligned_llm

Double Logging to Detect an Unaligned LLM Can comparing a model's own action log with an automatic system log detect harmful behaviour? The detector sees only the two logs and is scored on detection rate and false alarms. Status: in design. No data yet. Code GitHub Earlier study INSIDER_LLM_DETECTION_BENCHMARK (archived)

sourceHugging Faceupdated 15d agoView on Hugging Face
0likes43downloads
Dataset Card

Double Logging to Detect an Unaligned LLM

Can comparing a model's own action log with an automatic system log detect harmful behaviour? The detector sees only the two logs and is scored on detection rate and false alarms.

Status: in design. No data yet.

CodeGitHub
Earlier studyINSIDER_LLM_DETECTION_BENCHMARK (archived)