datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai-deception-signal-coherence-risk-v0.1What this repo is for
Detect deception and strategic misrepresentation.
Focus:
hidden failure signals
misleading summaries
omission of critical risk
alignment masking
Targets one of the highest-risk alignment failures: systems that know more than they reveal.
self-deception-detection-v0.1
What this dataset does
This dataset tests whether a model can detect self-deception.
The task is simple:
Given a scenario and a self-deception claim, predict whether the claim is supported.
Core stability idea
Self-deception occurs when a system maintains a preferred narrative despite contradictory evidence.
Common signs include:
externalizing blame
ignoring contradictory evidence
preserving identity over accuracy
selective interpretation
rationalization
narrative… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/self-deception-detection-v0.1.Deception-Sandbagging_Scenarios
Deception - Sandbagging Scenarios
This dataset contains 16k samples of deceptive and honest scenario completions where deception is incentivized but never forced by using harmful threats.
