Team Ai
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lamm-mit /gemma4-interpretability Gemma materials-science interpretability research archive Research records supporting Reading and Steering Materials Science-Mechanism Representations in an Open-Weight Language Model, Markus J. Buehler. Release identifier: paper-revision-2026-09-06. This archive supplies the original observations, supporting state arrays, exact prompts, protocols, intervention records, statistics, analysis source, and generated research figures. It includes the original 4B readout and geometry… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/gemma4-interpretability.2 likes1.2k downloads1mo agoHugging Face02Enderchef /Mechanic-Interpretability-Research-Datatabularn<1K1 likes357 downloads2mo agoHugging Face03NeurIPsMay1234 /openpi-interpretability-data openpi-interpretability-data Interpretability artifacts (activations, conceptors, linear steering vectors, sparse autoencoder vectors and checkpoints) extracted from open vision-language-action (VLA) policy models on the LIBERO, MetaWorld, and RoboCasa benchmarks. This dataset accompanies an anonymous submission and is shared for double-blind peer review. Models and benchmarks Model Family Benchmarks pi0_5 (pi05) π-series VLA LIBERO pi0_fast (pi0fast)… See the full description on the dataset page: https://huggingface.co/datasets/NeurIPsMay1234/openpi-interpretability-data.10B<n<100B0 likes340 downloads5mo agoHugging Face04pcam-interpretability /dino_vit_attnmaps0 likes56 downloads1y agoHugging Face05bedderautomation /mechanistic-interpretability-skills Mechanistic Interpretability Skills for Claude Code The first skill set for LLM refusal geometry extraction, boundary surface mapping, and self-referential mechanistic analysis. Compatible with: Claude Code, OpenAI Codex, Gemini CLI, and any tool supporting the agentskills.io standard. Skills 1. refusal-geometry/ Extract and analyze refusal cone geometry from open-weight transformer models using OBLITERATUS. Capabilities: 6-stage extraction pipeline (model… See the full description on the dataset page: https://huggingface.co/datasets/bedderautomation/mechanistic-interpretability-skills.1 likes54 downloads7mo agoHugging Face06Xu-AI4Science /MARRI-interpretability MARRI headline model — test-set interpretability export Per-sample interpretability/diagnostics for the MARRI headline configuration (dyn ft-rnafm, seed 42 — RNA-FM fine-tuned, RNet2D frozen, dynamic negative resampling), evaluated on the fixed held-out test split (n=16,058 pairs, seed 42). Code and training/evidence logs: https://github.com/GainGod-Xu/MARRI Files marri_ft-rnafm_dyn_bs4_seed42_test_results_interpretability.h5 — per-sample diagnostics:… See the full description on the dataset page: https://huggingface.co/datasets/Xu-AI4Science/MARRI-interpretability.0 likes43 downloads11d agoHugging Face07jaygala24 /reasoning-models-interpretability-artifacts Reasoning Models Interpretability Artifacts This dataset contains intermediate artifacts for studying reasoning traces in open-weight language models. It includes annotated-trace hidden representations and spectral metrics computed over reasoning-step categories. The artifacts are intended for analysis and sharing, not for direct datasets.load_dataset(...) loading as a tabular dataset. Contents annotated_traces_reprs/ <model>/ config.json index.json… See the full description on the dataset page: https://huggingface.co/datasets/jaygala24/reasoning-models-interpretability-artifacts.1K<n<10K0 likes34 downloads5mo agoHugging Face08burnssa /judge-distillation-medical-interpretability Judge-Distillation Medical Misalignment Interpretability Dataset A complete artifact bundle for the Phase 2 judge-distillation experiments described in judge_distillation/RESULTS.md. Includes training datasets, source per-prompt activations (the underlying drift_pct labels), Gemma Scope SAE feature attributions, hidden-state captures, and transfer-test corpora & scores across all five versions (v1–v5). What's in here Training datasets Three versions of… See the full description on the dataset page: https://huggingface.co/datasets/burnssa/judge-distillation-medical-interpretability.1K<n<10K0 likes29 downloads5mo agoHugging Face09Znreza /simulation-interpretability-dataset Qwen3.5-2B-Base Blind Spots Dataset A curated dataset documenting systematic failure modes ("blind spots") discovered in Qwen/Qwen3.5-2B-Base through structured probing experiments. Dataset Description This dataset contains 12 carefully selected examples where Qwen3.5-2B-Base exhibits predictable, reproducible failures across three major categories: Category Examples Key Finding Authority-Induced Sycophancy 4 Model accepts false claims when framed with… See the full description on the dataset page: https://huggingface.co/datasets/Znreza/simulation-interpretability-dataset.textn<1K0 likes23 downloads7mo agoHugging Face10maximuspowers /llm-interpretability-v1text1K<n<10K0 likes22 downloads1y agoHugging Face11jub-aer /ConsistencyBench-interpretability ConsistencyBench-Interpretability Extension White-box mechanistic analysis of logical inconsistency using Qwen/Qwen2.5-1.5B-Instruct (local, full activation access) as a dedicated interpretability testbed, distinct from the 17-model black-box leaderboard. Contents layer_probe_results.csv - per-layer logistic-regression probe accuracy for decoding "will this response be inconsistent?" directly from residual-stream activations activation_patching.csv - literal… See the full description on the dataset page: https://huggingface.co/datasets/jub-aer/ConsistencyBench-interpretability.0 likes19 downloads3mo agoHugging Face12fineset-io /mechanistic-interpretability-papers Mechanistic Interpretability Papers — FineSet A research-paper dataset on Mechanistic Interpretability Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-12. It is not auto-updated. Research on Mechanistic Interpretability Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/mechanistic-interpretability-papers.tabulartext-classificationn<1K1 likes18 downloads4mo agoHugging Face13SoumilB7 /Emotional_Interpretability Components Dataset : Emotional_perspectives : Response to a given context under 27 emotional lenses Description Synthetic dataset created to mimic emotional responses primarily made for alignment and interpretability research More details will be listed on github soon license: mit text10K<n<100K1 likes17 downloads1y agoHugging Face14introvoyz041 /interpretability_augmentation0 likes17 downloads3mo agoHugging Face15pcam-interpretability /pcam_heatmaps0 likes10 downloads1y agoHugging Face16maximuspowers /llm-interpretability-v1-messagestext1K<n<10K0 likes9 downloads1y agoHugging Face17pcam-interpretability /train_data0 likes6 downloads1y agoHugging Face18HannahJIANG /interpretabilitygated0 likes4 downloads2mo agoHugging Face19Maksim-KOS /ai_interpretability_hack0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.