datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
deception-probes-activations
Deception Probes Activations
Pre-extracted residual-stream activations for training and evaluating deception
detection probes on LLMs. Each example contains per-token hidden states from a
specific transformer layer, saved in bfloat16 safetensors format.
License
This dataset contains activations derived from multiple sources with different licenses.
See the LICENSE file for full details.
Component
Source
License
Apollo Probe Pairs (statements)
Azaria & Mitchell… See the full description on the dataset page: https://huggingface.co/datasets/xycoord/deception-probes-activations.synthetic-math-toolcall-deception
Synthetic Math Tool-Call Deception
200 paired multi-turn math-assistant trajectories (400 rows) for evaluating deception
detectors on mid-trajectory tool-call misreporting.
Each trajectory: a system prompt instructs the model to compute via an execute_python
tool under a stated tool-call limit, and requires every call to carry a running
call_index argument (1 for the first call, 2 for the second, …). The platform enforcing
the limit is said to only see the reported call_index… See the full description on the dataset page: https://huggingface.co/datasets/annahbanannah/synthetic-math-toolcall-deception.deception-activationsdeception-activationsdeception-probing-tutorial
Deception probing tutorial — Gemma-2-9B-IT activations
Precomputed residual-stream activations for a hands-on replication of
Natarajan et al. (2026), One Probe Won't Catch Them All (arXiv:2602.01425),
which builds on Goldowsky-Dill et al. (2025), Detecting Strategic Deception with
Linear Probes.
The point of shipping activations rather than a model: everything scientifically
interesting in both papers happens downstream of the forward pass. With these
vectors the whole tutorial… See the full description on the dataset page: https://huggingface.co/datasets/Rutabin/deception-probing-tutorial.DeceptionBench
DeceptionBench: A Comprehensive Benchmark for Evaluating Deceptive Behaviors in Large Language Models
🔍 Overview
DeceptionBench is the first systematic benchmark designed to assess deceptive behaviors in Large Language Models (LLMs). As modern LLMs increasingly rely on chain-of-thought (CoT) reasoning, they may exhibit deceptive alignment - situations where models appear aligned while covertly pursuing misaligned goals.
This benchmark addresses a critical gap in AI… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/DeceptionBench.DeceptionBench
DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios
Paper: DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios
Code: https://github.com/Aries-iai/DeceptionBench
Overview
DeceptionBench is a comprehensive framework designed to systematically evaluate deceptive behaviors in large language models (LLMs). As LLMs achieve remarkable proficiency across diverse tasks, emergent behaviors like… See the full description on the dataset page: https://huggingface.co/datasets/skyai798/DeceptionBench.EDR_Telemetry_SampleThis dataset contains raw Endpoint Detection & Response (EDR) telemetry captured during controlled Deception.Pro malware sandbox operations on an enterprise Active Directory network. Unlike most malware sandboxes — which detonate samples for roughly 30 minutes — our operations run for hours or days per analysis, capturing the full arc of adversary behavior. The data represents a full-fidelity snapshot of system activity recorded while threat actors interacted with a live deception environment… See the full description on the dataset page: https://huggingface.co/datasets/DeceptionPro/EDR_Telemetry_Sample.energy-cost-deception-llm
ECDL: Energy Cost of Deception in LLMs
Objective
This dataset records exploratory tests of language-model responses to false instructions and contextual information. The project examines answer changes and token log-probabilities, and asks how they relate to error correction and computational cost.
The project name describes a research direction. These probability and response records do not by themselves establish an energy cost of deception.
Plausible… See the full description on the dataset page: https://huggingface.co/datasets/levgogo/energy-cost-deception-llm.llm-deception-trajectories
LLM Deception Trajectories
Hidden-state trajectories from 11 transformer architectures processing matched truthful/deceptive prompt pairs across 20 deception categories.
Dataset Description
This dataset captures the internal processing trajectories of large language models as they generate responses to truthful vs. deceptive prompts. Each trajectory records the hidden state at every transformer layer, enabling analysis of how deception manifests in model… See the full description on the dataset page: https://huggingface.co/datasets/dSLLab/llm-deception-trajectories.Sahal-DeceptionBlobdeception-localization
Counterfactual Deception Localization
This dataset contains synthetic counterfactual localization data for studying when language models become committed to truthful or deceptive behavior during reasoning.
Each example starts from a model-generated reasoning trace in a strategic-deception environment. The trace is split into sentence prefixes. At selected sentence boundaries, the prefix is fixed and the same model is asked to sample multiple possible continuations. Those… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-neurips-2026-ED/deception-localization.llm-deception-amongusCCP-deception
CCP-Deception (may_27 splits)
One row per (model, language, question_id) that passes the split filter.
Full per-language traces under experiments (chinese, prefill, eval_awareness, phrasing).
Lie splits
split
n
filter
kimi_chinese_lie
102
kimi: strict 3/3 committed lie
kimi_english_lie
75
kimi: strict 3/3 committed lie
qwen_chinese_lie
104
qwen: ≥2/3 committed lie
qwen_english_lie
109
qwen: ≥2/3 committed lie
Truth splits… See the full description on the dataset page: https://huggingface.co/datasets/dipikakhullar/CCP-deception.deception-activationsdeception-activations-70bdeception-behavioral-multimodel
Multi-Model Deception Behavioral Activation Dataset
Activation vectors from three language models during deceptive vs honest text generation, collected using V3 behavioral sampling.
Key Results
Model
Params
d_model
Peak Balanced Acc
AUROC
Samples (dec/hon)
nanochat-d32
1.88B
2048
86.9%
0.923
650:677
Llama 3.2-1B
1.3B
2048
76.2%
0.820
103:67
nanochat-d20
561M
1280
66.1%
0.713
132:128
Signal strength scales with model size. All p < 0.01.… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/deception-behavioral-multimodel.deception-probing-tutorial-lite
Deception probing tutorial — Gemma-2-9B-IT activations (lite)
Precomputed residual-stream activations for a hands-on replication of
Natarajan et al. (2026), One Probe Won't Catch Them All (arXiv:2602.01425),
which builds on Goldowsky-Dill et al. (2025), Detecting Strategic Deception with
Linear Probes.
The point of shipping activations rather than a model: everything scientifically
interesting in both papers happens downstream of the forward pass. With these
vectors the whole… See the full description on the dataset page: https://huggingface.co/datasets/Rutabin/deception-probing-tutorial-lite.Sahal-DeceptionBlobExamining_Deception_GemmaScope_Examplesdeception-warning-study-runs
Deception Warning Study — run-level benchmark results
This dataset contains run-level rows for the controlled benchmark on warning placement for web agents under deceptive interfaces (ShopLane / WorkHub tasks).
Contents
File
Description
run_level.parquet
Hub-friendly columnar format (recommended)
run_level.jsonl
One JSON object per run
run_level.csv
Same data as CSV
export_meta.json
Export metadata: column list, row count, schema version
Current… See the full description on the dataset page: https://huggingface.co/datasets/deceptive-web/deception-warning-study-runs.DeceptionDecoded
Data for “Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models” (ICLR 2026)
This repository provides the DeceptionDecoded benchmark for intent-aware multimodal misinformation detection (MMD).
DeceptionDecoded Dataset
DeceptionDecoded contains 12,000 multimodal news samples, evenly distributed across six intent classes. These classes correspond to different forms of text- and image-based misleadingness, as shown in… See the full description on the dataset page: https://huggingface.co/datasets/jiayingwu19/DeceptionDecoded.human_deceptionInterrogation_Dataset_for_AI_Deception_Detection
Interrogation Dataset for AI Deception Detection
Overview
This dataset is designed for training AI models in deception detection, behavioral analysis, and tactical decision-making during criminal interrogations.
It contains 1600 entries (INT-0001 to INT-1600) in JSONL format, covering various criminal scenarios such as financial crimes, murder, fraud, burglary, physical assault, and molestation.
The dataset reflects realistic law enforcement contexts across diverse global settings… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Interrogation_Dataset_for_AI_Deception_Detection.deception_taxonomy_papercontextual-deception-detection
ConDec: Contextual Deception Detection Benchmark
Detecting Technically-True-but-Misleading Claims in Scientific ML Papers
Overview
ConDec is a benchmark for detecting contextual deception — statements in scientific ML papers that are literally true but systematically misleading due to omitted context, cherry-picked results, or other forms of pragmatic manipulation.
Unlike fact verification, which checks whether claims are supported by evidence, contextual deception… See the full description on the dataset page: https://huggingface.co/datasets/t6harsh/contextual-deception-detection.train-deceptiondeceptiongemma-4-e2b-deception-behavior-completions
Gemma-4-E2B deception & behavior completions
Consolidated 910-row corpus of (scenario prompt + Gemma-4-E2B-generated completion) pairs from earlier mechanistic-interpretability experiments. Each row captures the prompt the model saw and the text it actually produced; for a subset, Claude-Haiku-4-5 judge verdicts and SAE-feature labels are included.
The corpus is meant to be used as activation-extraction input for downstream interpretability work — Natural Language Autoencoder (NLA)… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-deception-behavior-completions.dev-varied-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-1
