Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01xycoord /deception-probes-activations Deception Probes Activations Pre-extracted residual-stream activations for training and evaluating deception detection probes on LLMs. Each example contains per-token hidden states from a specific transformer layer, saved in bfloat16 safetensors format. License This dataset contains activations derived from multiple sources with different licenses. See the LICENSE file for full details. Component Source License Apollo Probe Pairs (statements) Azaria & Mitchell… See the full description on the dataset page: https://huggingface.co/datasets/xycoord/deception-probes-activations.texttext-classification1M<n<10M1 likes51k downloads5mo agoHugging Face02annahbanannah /synthetic-math-toolcall-deception Synthetic Math Tool-Call Deception 200 paired multi-turn math-assistant trajectories (400 rows) for evaluating deception detectors on mid-trajectory tool-call misreporting. Each trajectory: a system prompt instructs the model to compute via an execute_python tool under a stated tool-call limit, and requires every call to carry a running call_index argument (1 for the first call, 2 for the second, …). The platform enforcing the limit is said to only see the reported call_index… See the full description on the dataset page: https://huggingface.co/datasets/annahbanannah/synthetic-math-toolcall-deception.tabulartext-classificationn<1K0 likes6.4k downloads3mo agoHugging Face03AISC-Linear-Probe-Gen /deception-activationstabular10K<n<100K0 likes1.5k downloads9mo agoHugging Face04lasrprobegen /deception-activationstabular10K<n<100K2 likes1.2k downloads10mo agoHugging Face05AISC-Linear-Probe-Gen /deception_taxonomy_papertext10K<n<100K1 likes616 downloads6mo agoHugging Face06Rutabin /deception-probing-tutorial Deception probing tutorial — Gemma-2-9B-IT activations Precomputed residual-stream activations for a hands-on replication of Natarajan et al. (2026), One Probe Won't Catch Them All (arXiv:2602.01425), which builds on Goldowsky-Dill et al. (2025), Detecting Strategic Deception with Linear Probes. The point of shipping activations rather than a model: everything scientifically interesting in both papers happens downstream of the forward pass. With these vectors the whole tutorial… See the full description on the dataset page: https://huggingface.co/datasets/Rutabin/deception-probing-tutorial.textfeature-extraction1K<n<10K0 likes367 downloads2mo agoHugging Face07skyai798 /DeceptionBench DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios Paper: DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios Code: https://github.com/Aries-iai/DeceptionBench Overview DeceptionBench is a comprehensive framework designed to systematically evaluate deceptive behaviors in large language models (LLMs). As LLMs achieve remarkable proficiency across diverse tasks, emergent behaviors like… See the full description on the dataset page: https://huggingface.co/datasets/skyai798/DeceptionBench.texttext-classificationn<1K4 likes303 downloads11mo agoHugging Face08PKU-Alignment /DeceptionBench DeceptionBench: A Comprehensive Benchmark for Evaluating Deceptive Behaviors in Large Language Models 🔍 Overview DeceptionBench is the first systematic benchmark designed to assess deceptive behaviors in Large Language Models (LLMs). As modern LLMs increasingly rely on chain-of-thought (CoT) reasoning, they may exhibit deceptive alignment - situations where models appear aligned while covertly pursuing misaligned goals. This benchmark addresses a critical gap in AI… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/DeceptionBench.texttext-classificationn<1K4 likes283 downloads1y agoHugging Face09dSLLab /llm-deception-trajectories LLM Deception Trajectories Hidden-state trajectories from 11 transformer architectures processing matched truthful/deceptive prompt pairs across 20 deception categories. Dataset Description This dataset captures the internal processing trajectories of large language models as they generate responses to truthful vs. deceptive prompts. Each trajectory records the hidden state at every transformer layer, enabling analysis of how deception manifests in model… See the full description on the dataset page: https://huggingface.co/datasets/dSLLab/llm-deception-trajectories.tabulartext-classification10K<n<100K0 likes133 downloads3mo agoHugging Face10h-gajdov /llm-deception-amongustext0 likes94 downloads2mo agoHugging Face11deceptive-web /deception-warning-study-runs Deception Warning Study — run-level benchmark results This dataset contains run-level rows for the controlled benchmark on warning placement for web agents under deceptive interfaces (ShopLane / WorkHub tasks). Contents File Description run_level.parquet Hub-friendly columnar format (recommended) run_level.jsonl One JSON object per run run_level.csv Same data as CSV export_meta.json Export metadata: column list, row count, schema version Current… See the full description on the dataset page: https://huggingface.co/datasets/deceptive-web/deception-warning-study-runs.textn<1K0 likes45 downloads5mo agoHugging Face12compl-ai /human_deceptiontextn<1K0 likes44 downloads1y agoHugging Face13darkknight25 /Interrogation_Dataset_for_AI_Deception_Detection Interrogation Dataset for AI Deception Detection Overview This dataset is designed for training AI models in deception detection, behavioral analysis, and tactical decision-making during criminal interrogations. It contains 1600 entries (INT-0001 to INT-1600) in JSONL format, covering various criminal scenarios such as financial crimes, murder, fraud, burglary, physical assault, and molestation. The dataset reflects realistic law enforcement contexts across diverse global settings… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Interrogation_Dataset_for_AI_Deception_Detection.texttext-classification1K<n<10K1 likes44 downloads1y agoHugging Face14Solshine /gemma-4-e2b-deception-behavior-completions Gemma-4-E2B deception & behavior completions Consolidated 910-row corpus of (scenario prompt + Gemma-4-E2B-generated completion) pairs from earlier mechanistic-interpretability experiments. Each row captures the prompt the model saw and the text it actually produced; for a subset, Claude-Haiku-4-5 judge verdicts and SAE-feature labels are included. The corpus is meant to be used as activation-extraction input for downstream interpretability work — Natural Language Autoencoder (NLA)… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/gemma-4-e2b-deception-behavior-completions.tabulartext-generationn<1K0 likes30 downloads5mo agoHugging Face15sitong-fang /MM-DeceptionBench 🎭 MM-DeceptionBench A Multimodal Benchmark for Evaluating Deceptive Behaviors in Vision-Language Models 📖 Overview MM-DeceptionBench is a comprehensive benchmark designed to stress-test Multimodal Large Language Models (MLLMs) for strategic deception in visually grounded contexts. It captures nuanced deceptive behaviors that emerge when models interact with images and text, spanning diverse real-world scenarios. ✨ Key Highlights 🔢… See the full description on the dataset page: https://huggingface.co/datasets/sitong-fang/MM-DeceptionBench.imagevisual-question-answering1K<n<10K0 likes26 downloads11mo agoHugging Face16reinthal /dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-5-relabel-v5 dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-5-relabel-v5 — v5 relabel + split Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-5 with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a train/test/validation split column. Added columns: deceptive (v5 label; official fallback where excluded), official (original dev label), relabeled… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-5-relabel-v5.tabularn<1K0 likes26 downloads3mo agoHugging Face17Avyay10 /train-deceptiontabular10K<n<100K0 likes25 downloads2y agoHugging Face18aletheias-quest /dev-varied-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-1tabularn<1K0 likes24 downloads3mo agoHugging Face19reinthal /dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3-relabel-v5 dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3-relabel-v5 — v5 relabel + split Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3 with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a train/test/validation split column. Added columns: deceptive (v5 label; official fallback where excluded), official (original dev label), relabeled… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-3-relabel-v5.tabularn<1K0 likes24 downloads3mo agoHugging Face20Avyay10 /eval-deception-backdoortext1K<n<10K0 likes23 downloads2y agoHugging Face21reinthal /dev-instructed-deception-Qwen3.5-27B-c-mo-qwen3.5-27b-relabel-v5 dev-instructed-deception-Qwen3.5-27B-c-mo-qwen3.5-27b-relabel-v5 — v5 relabel + split Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-c-mo-qwen3.5-27b with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a train/test/validation split column. Added columns: deceptive (v5 label; official fallback where excluded), official (original dev label), relabeled (v5 !=… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-c-mo-qwen3.5-27b-relabel-v5.tabularn<1K0 likes23 downloads3mo agoHugging Face22Reih02 /deception_obfuscation_deepseek_v3_behavioral_iter2tabular1K<n<10K0 likes22 downloads7mo agoHugging Face23reinthal /dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6-relabel-v5 dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6-relabel-v5 — v5 relabel + split Copy of aletheias-quest/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6 with the v5 belief-relative label (see reinthal/aletheias-dev-relabel-v5 for the method: 20x neutral resample -> DeepSeek-V4-Flash judge, no canonicalization) and a train/test/validation split column. Added columns: deceptive (v5 label; official fallback where excluded), official (original dev label), relabeled… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/dev-instructed-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6-relabel-v5.tabularn<1K0 likes22 downloads3mo agoHugging Face24aletheias-quest /dev-varied-deception-Qwen3.5-27B-a-mo-qwen3.5-27b-6tabularn<1K0 likes21 downloads3mo agoHugging Face25Avyay10 /merged-train-deceptiontext10K<n<100K0 likes20 downloads2y agoHugging Face26arianaazarbal /detected-solid-deceptiontabular10K<n<100K0 likes20 downloads1y agoHugging Face27reinthal /qwen3.5-9b-deception-probe Qwen3.5-9B deception probe (belief-relative v5 labels) Qwen/Qwen3.5-9B answering the aletheias dev-set questions (true/false instructed + open varied trivia) under three system prompts: neutral (20× resample, no system prompt), honest ("Answer the following question truthfully."), lie ("Answer the following question with a lie."). Temperature 0.7, reasoning disabled. Labels (v5, belief-relative). Each question's 20 neutral answers define the model's own belief; a… See the full description on the dataset page: https://huggingface.co/datasets/reinthal/qwen3.5-9b-deception-probe.tabular10K<n<100K0 likes20 downloads3mo agoHugging Face28Avyay10 /train-deception-backdoortext1K<n<10K0 likes19 downloads2y agoHugging Face29Reih02 /deception_obfuscation_nemotron_30b_behavioral_v4_1272tabular1K<n<10K0 likes19 downloads5mo agoHugging Face30AlignmentResearch /hidden-goal-model-organism-deception-dataset-gemma3-27b-v1gated AlignmentResearch/hidden-goal-model-organism-deception-dataset-gemma3-27b-v1 Private dataset of on-policy model-organism transcripts labelled honest/deceptive, for lie-detection research. Do not redistribute. Columns model — HuggingFace repo id of the model organism that generated the transcript. messages — the conversation in ChatML format; the last message is the assistant turn that is being labelled. deceptive — bool; whether the last assistant message is a… See the full description on the dataset page: https://huggingface.co/datasets/AlignmentResearch/hidden-goal-model-organism-deception-dataset-gemma3-27b-v1.textn<1K0 likes18 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.