confound
lstm.CEBaB_confounding.food_service_positive.absa.5-class.seed_42lstm.CEBaB_confounding.uniform.sa.5-class.seed_43lstm.CEBaB_confounding.uniform.sa.5-class.seed_42lstm.CEBaB_confounding.uniform.absa.5-class.seed_43lstm.CEBaB_confounding.uniform.sa.5-class.seed_44gpt2.CEBaB_confounding.price_food_ambiance_negative.sa.5-class.seed_44lstm.CEBaB_confounding.price_food_ambiance_negative.absa.5-class.seed_43roberta-base.CEBaB_confounding.observational.sa.5-class.seed_44
Datasets
All datasets matching “confound”length-confound-benchmark
Cached benchmark data — length-confound audit
Pre-computed residual-stream hidden states and derived features for auditing
hallucination detectors. The audit re-runs in under an hour once downloaded,
versus roughly 30 hours of forward passes to rebuild from scratch.
Contents
{model}_{dataset}_rtraj_features.npz — 17 conditions.
Keys: labels, responses, questions, proj_h_reasoning,
proj_a_reasoning, proj_m_reasoning, reasoning_dim.… See the full description on the dataset page: https://huggingface.co/datasets/AnonyJterwe/length-confound-benchmark.daily-paper-2026-10-04-grader-confound-agentic-rankings
The Grader Confound: Measuring How Grader Choice Inverts the Measured Cost-Quality Rankings of Agentic Tool-Call Arms on Self-Hosted H200
TL;DR — Graders are instruments, not free observations. We give a cell-level formal model of exactly when switching between a deterministic AST grader, a surface-pattern gate, and an LLM judge inverts the measured cost-quality rankings of agentic tool-call arms, price the judge in tokens (the grade tax), and pre-register a three-grader audit… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-10-04-grader-confound-agentic-rankings.multimodal_confounding
Dataset Card
Semi-synthetic dataset with multimodal confounding.
The dataset is generated according to the description in DoubleMLDeep: Estimation of Causal Effects with Multimodal Data.
Dataset Details
Dataset Description & Usage
The dataset is a semi-synthetic dataset as a benchmark for treatment effect estimation with multimodal confounding. The outcome
variable Y is generated according to a partially linear model
Y=θ0D1+g1(X)+ε
Y = \theta_0 D_1 + g_1(X) +… See the full description on the dataset page: https://huggingface.co/datasets/DoubleML/multimodal_confounding.clinical-quad-safety-signal-latency-reporting-lag-conmed-confound-v0.1Clarus Clinical Quad Coupling Safety Signal Latency Reporting Lag Conmed Confound v0.1
What this dataset isThis dataset tests whether a model can detect latent safety signals when four interacting nodes create uncertainty.
Quad coupling nodes
Emerging safety event pattern
Reporting or entry latency
Concomitant medication or behavior confound
Governance decision timing such as DSMB, batch release, or safety review
Input
One vignette
OutputReturn strict JSON only.
Required output… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-safety-signal-latency-reporting-lag-conmed-confound-v0.1.CEBaB_train_confounding_uniformCEBaB_train_confounding_food_service_positive
repro-care-confounder-aware-aggregationrepro-causal-effect-identifiability-in-the-presence-of-latent-confounders-without-auxiliary-varirepro-causal-effect-identifiability-in-the-presence-of-latent-confounders-without-auxiliary-varicausal-identifiability-latent-confounders-reprorepro-addressing-instrument-outcome-confounding-in-mendelian-randomization
