datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
deception-probes-activations
Deception Probes Activations
Pre-extracted residual-stream activations for training and evaluating deception
detection probes on LLMs. Each example contains per-token hidden states from a
specific transformer layer, saved in bfloat16 safetensors format.
License
This dataset contains activations derived from multiple sources with different licenses.
See the LICENSE file for full details.
Component
Source
License
Apollo Probe Pairs (statements)
Azaria & Mitchell… See the full description on the dataset page: https://huggingface.co/datasets/xycoord/deception-probes-activations.synthetic-math-toolcall-deception
Synthetic Math Tool-Call Deception
200 paired multi-turn math-assistant trajectories (400 rows) for evaluating deception
detectors on mid-trajectory tool-call misreporting.
Each trajectory: a system prompt instructs the model to compute via an execute_python
tool under a stated tool-call limit, and requires every call to carry a running
call_index argument (1 for the first call, 2 for the second, …). The platform enforcing
the limit is said to only see the reported call_index… See the full description on the dataset page: https://huggingface.co/datasets/annahbanannah/synthetic-math-toolcall-deception.deception-activationsdeception-activationsdeception_taxonomy_paperdeception-probing-tutorial
Deception probing tutorial — Gemma-2-9B-IT activations
Precomputed residual-stream activations for a hands-on replication of
Natarajan et al. (2026), One Probe Won't Catch Them All (arXiv:2602.01425),
which builds on Goldowsky-Dill et al. (2025), Detecting Strategic Deception with
Linear Probes.
The point of shipping activations rather than a model: everything scientifically
interesting in both papers happens downstream of the forward pass. With these
vectors the whole tutorial… See the full description on the dataset page: https://huggingface.co/datasets/Rutabin/deception-probing-tutorial.DeceptionBench
DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios
Paper: DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios
Code: https://github.com/Aries-iai/DeceptionBench
Overview
DeceptionBench is a comprehensive framework designed to systematically evaluate deceptive behaviors in large language models (LLMs). As LLMs achieve remarkable proficiency across diverse tasks, emergent behaviors like… See the full description on the dataset page: https://huggingface.co/datasets/skyai798/DeceptionBench.DeceptionBench
DeceptionBench: A Comprehensive Benchmark for Evaluating Deceptive Behaviors in Large Language Models
🔍 Overview
DeceptionBench is the first systematic benchmark designed to assess deceptive behaviors in Large Language Models (LLMs). As modern LLMs increasingly rely on chain-of-thought (CoT) reasoning, they may exhibit deceptive alignment - situations where models appear aligned while covertly pursuing misaligned goals.
This benchmark addresses a critical gap in AI… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/DeceptionBench.llm-deception-trajectories
LLM Deception Trajectories
Hidden-state trajectories from 11 transformer architectures processing matched truthful/deceptive prompt pairs across 20 deception categories.
Dataset Description
This dataset captures the internal processing trajectories of large language models as they generate responses to truthful vs. deceptive prompts. Each trajectory records the hidden state at every transformer layer, enabling analysis of how deception manifests in model… See the full description on the dataset page: https://huggingface.co/datasets/dSLLab/llm-deception-trajectories.llm-deception-amongusdeception-warning-study-runs
Deception Warning Study — run-level benchmark results
This dataset contains run-level rows for the controlled benchmark on warning placement for web agents under deceptive interfaces (ShopLane / WorkHub tasks).
Contents
File
Description
run_level.parquet
Hub-friendly columnar format (recommended)
run_level.jsonl
One JSON object per run
run_level.csv
Same data as CSV
export_meta.json
Export metadata: column list, row count, schema version
Current… See the full description on the dataset page: https://huggingface.co/datasets/deceptive-web/deception-warning-study-runs.human_deceptionInterrogation_Dataset_for_AI_Deception_Detection
Interrogation Dataset for AI Deception Detection
Overview
This dataset is designed for training AI models in deception detection, behavioral analysis, and tactical decision-making during criminal interrogations.
It contains 1600 entries (INT-0001 to INT-1600) in JSONL format, covering various criminal scenarios such as financial crimes, murder, fraud, burglary, physical assault, and molestation.
The dataset reflects realistic law enforcement contexts across diverse global settings… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Interrogation_Dataset_for_AI_Deception_Detection.gemma-4-e2b-deception-behavior-completions
Gemma-4-E2B deception & behavior completions
Consolidated 910-row corpus of (scenario prompt + Gemma-4-E2B-generated completion) pairs from earlier mechanistic-interpretability experiments. Each row captures the prompt the model saw and the text it actually produced; for a subset, Claude-Haiku-4-5 judge verdicts and SAE-feature labels are included.
The corpus is meant to be used as activation-extraction input for downstream interpretability work — Natural Language Autoencoder (NLA)… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-deception-behavior-completions.MM-DeceptionBench
🎭 MM-DeceptionBench
A Multimodal Benchmark for Evaluating Deceptive Behaviors in Vision-Language Models
📖 Overview
MM-DeceptionBench is a comprehensive benchmark designed to stress-test Multimodal Large Language Models (MLLMs) for strategic deception in visually grounded contexts. It captures nuanced deceptive behaviors that emerge when models interact with images and text, spanning diverse real-world scenarios.
✨ Key Highlights
🔢… See the full description on the dataset page: https://huggingface.co/datasets/sitong-fang/MM-DeceptionBench.dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-5-relabel-v5
dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-5-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-5 with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label), relabeled… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-5-relabel-v5.train-deceptiondev-varied-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-1dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3-relabel-v5
dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3 with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label), relabeled… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3-relabel-v5.eval-deception-backdoordev-instructed-deception-Qwen3.5-27B-c-mo-qwen3.5-27b-relabel-v5
dev-instructed-deception-Qwen3.5-27B-c-mo-qwen3.5-27b-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-c-mo-qwen3.5-27b with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label), relabeled (v5 !=… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-c-mo-qwen3.5-27b-relabel-v5.deception_obfuscation_deepseek_v3_behavioral_iter2dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6-relabel-v5
dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6-relabel-v5 — v5 relabel + split
Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6 with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the
method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a
train/test/validation split column.
Added columns: deceptive (v5 label; official fallback where excluded), official (original dev
label), relabeled… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6-relabel-v5.dev-varied-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6merged-train-deceptiondetected-solid-deceptionqwen3.5-9b-deception-probe
Qwen3.5-9B deception probe (belief-relative v5 labels)
Qwen/Qwen3.5-9B answering the aletheias dev-set questions (true/false instructed + open
varied trivia) under three system prompts: neutral (20× resample, no system prompt),
honest ("Answer the following question truthfully."), lie ("Answer the following question
with a lie."). Temperature 0.7, reasoning disabled.
Labels (v5, belief-relative). Each question's 20 neutral answers define the model's own
belief; a… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/qwen3.5-9b-deception-probe.train-deception-backdoordeception_obfuscation_nemotron_30b_behavioral_v4_1272hidden-goal-model-organism-deception-dataset-gemma3-27b-v1
AlignmentResearch/hidden-goal-model-organism-deception-dataset-gemma3-27b-v1
Private dataset of on-policy model-organism transcripts labelled
honest/deceptive, for lie-detection research.
Do not redistribute.
Columns
model — HuggingFace repo id of the model organism that generated the transcript.
messages — the conversation in ChatML format; the last message is the assistant
turn that is being labelled.
deceptive — bool; whether the last assistant message is a… See the full description on the dataset page: https://huggingface.co/datasets/AlignmentResearch/hidden-goal-model-organism-deception-dataset-gemma3-27b-v1.
