datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
factprobe-replication-negatives-allnames-v1
Plausible wrong answers, asked under every name (42,267,800 rows)
Status: final — 40 of 40 runs. Models present:
13b, 7b. Training stages present: s1, s2, s3, s4, s5.
What this fixes
A real fact is put to the model under the full cross product of the two
people's name lists, and counts as recognised if any one combination gets a
Yes. That is He et al.'s rule. Until 2026-08-26 the wrong answer it was
compared against was asked under one name per person, so the… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-negatives-allnames-v1.cna-vulnerability-census-replication
Empirical Vulnerability Census (1999–2026, $N = 385,524$), Cybernetic Queueing Instability, and CISA BOD 26-04 Remediation Deficit
Deterministic Empirical Replication Package & Econometric Audits
Principal Investigator: Gia Bao Huynh (Jun Huynh)ORCID: 0009-0008-2372-5852Affiliation: Independent Scholar / Ho Chi Minh City, VietnamLive Interactive Simulator: Cybernetic Queueing Instability Simulator (M/G/1)
🏛️ Executive Summary & Theoretical… See the full description on the dataset page: https://huggingface.co/datasets/giabaohuynhasu/cna-vulnerability-census-replication.factprobe-replication-traj-b2-stage1-v1
factprobe-replication-traj-b2-stage1-v1
Checkpoint-trajectory probing: b2 few-shot recognition on 13 log-spaced stage-1 checkpoints per model, spouse+sibling, both templates, original alternating demos, fp16 pinned. Identity columns model_tag/revision/tokens_b injected from filenames. Complete checkpoint files only.
Dataset Info
Rows: 26315536
Columns: 18
Columns
Column
Type
Description
relation
Value('string')
No description provided… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-traj-b2-stage1-v1.gz3d-aion-replication-attempt
GalaxyZoo 3D Segmentation Benchmark
An attempted replication of the GZ3D segmentation dataset used in section 7.2.3 in the AION-1. This dataset contains volunteer segmentations for the following channels: center, star, spiral, bar. It also contains the RGB images from Legacy Survey and the tokens from the AION-1 image tokenizer of the Legacy Survey image bands. Notably, these are pre–AION-1 transformer encoder, but post image-codec encoder/quantization. The reason why it's an… See the full description on the dataset page: https://huggingface.co/datasets/astronolan/gz3d-aion-replication-attempt.tempest-replication
TEMPEST Replication Dataset
Multi-turn adversarial attack results on 10 frontier LLMs.
Dataset Description
This dataset contains results from replicating the TEMPEST multi-turn jailbreak framework across 10 frontier language models, each evaluated on 100 harmful behaviors from JailbreakBench.
Key Findings
ASR range: 42-100% - All models vulnerable
No scale-safety correlation (r=-0.12)
Thinking mode helps: Kimi K2 Thinking (42%) vs standard (97%)
Files… See the full description on the dataset page: https://huggingface.co/datasets/richardyoung/tempest-replication.reach-of-the-state-replication
Replication Data for: The Reach of the State
This dataset is the official replication package for:
Chang, Charles, and Yuhua Wang. 2024. "The Reach of the State." Comparative Political Studies 57(8): 1243–1275. https://doi.org/10.1177/00104140231194057
It is a byte-for-byte mirror of the replication package archived on Harvard Dataverse
(doi:10.7910/DVN/CSBRZX, license CC0 1.0). If you use this package, please cite the
original paper above. The canonical archival version… See the full description on the dataset page: https://huggingface.co/datasets/fantasien/reach-of-the-state-replication.sshfighter-17-head-router-delay-replication-v1
SSH Fighter 17-Head Frozen-Expert Router v1 — Delay Stress
Seventeen character-specific heads route among nine immutable Agent Gym
policies. The package contains every three-seed checkpoint, averaged inference
weights, the grouped training tensor cache, and fresh-seed exact-engine
evaluation rows. This package is the predeclared six-scenario role-delay stress matrix.
Fresh exact-engine evaluation
condition
points rate
wins
losses
draws
switches/match… See the full description on the dataset page: https://huggingface.co/datasets/LisaMegaWatts/sshfighter-17-head-router-delay-replication-v1.factprobe-replication-b2-olmo7b-stage1-v1
factprobe-replication-b2-olmo7b-stage1-v1
Arm b2 of the FactProbe replication (He et al. arXiv:2503.22362): few-shot completion probing of the BASE OLMo-2-1124-7B final stage-1 checkpoint on symmetric-relation fact recognition (spouse P26, sibling P3373), forward and backward directions, all alias variations. One row per prompt; p_yes/p_no are next-token probability mass over Yes/No (question) or True/False (statement) single-token variants.
Dataset Info
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-b2-olmo7b-stage1-v1.factprobe-replication-asked-negatives-v1
Asking a model for a plausible wrong answer (100,776 questions)
Two prompts, put to gpt-5.1 for every subject surface form in the spouse
and sibling data, ten independent samples each:
(a) "Who is the spouse of X? Just the name, no explanation needed."
(b') "Name a person who could reasonably be mistaken for the spouse of X,
but is not the spouse of X. Just the name, no explanation needed."
50,388 surface forms across 10,592 entities, times two prompts, is
100,776 questions… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-asked-negatives-v1.svb-stablecoins-replication-data
Data for the paper:
Hernandez Cruz, W., Xu, J., Tasca, P., & Campajola, C. (2024). Impact of Adverse Disclosures on Liquidity of Stablecoin Markets. arXiv:2407.11716.
Overview
This dataset holds the small inputs of the paper. The two large inputs are in svb-stablecoins-hourly-tick-liquidity (the tick-level liquidity) and svb-stablecoins-hourly-lp-positions (the positions of each liquidity provider, LP). The paper estimates a difference-in-differences (DiD)… See the full description on the dataset page: https://huggingface.co/datasets/ExponentialScience/svb-stablecoins-replication-data.factprobe-replication-b2-olmo13b-stage1-v1
factprobe-replication-b2-olmo13b-stage1-v1
Arm b2 of the FactProbe replication (He et al. arXiv:2503.22362): few-shot completion probing of the BASE OLMo-2-1124-7B final stage-1 checkpoint on symmetric-relation fact recognition (spouse P26, sibling P3373), forward and backward directions, all alias variations. One row per prompt; p_yes/p_no are next-token probability mass over Yes/No (question) or True/False (statement) single-token variants.
Dataset Info
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-b2-olmo13b-stage1-v1.inside-out-replication-v2-probe-scores
inside-out-replication-v2-probe-scores
Internal (probe) scores: logistic-regression probe on hidden states, best layer chosen by dev K. Trained probe .pkl files and per-layer selection JSON attached to this repo.
Dataset Info
Rows: 1574024
Columns: 7
Columns
Column
Type
Description
question_id
Value('string')
Question identifier
answer
Value('string')
Answer string (full)
label
Value('string')
Judge label CORRECT/INCORRECT
probe_score… See the full description on the dataset page: https://huggingface.co/datasets/latkes/inside-out-replication-v2-probe-scores.factprobe-replication-hardneg-knows-labels-v1
factprobe-replication-hardneg-knows-labels-v1
Paired 'knows the fact' judgements of OLMo-2 (7B, 13B) against GOOD hard negatives, across the training ladder. One row per (model, stage, relation, phrasing, true pair): P(Yes) on the true pair, best P(Yes) on the hardest hard negative, and beats_all (does the model prefer the true partner). Forward direction; spouse (P26) and sibling (P3373).
Dataset Info
Rows: 284260
Columns: 10
Columns
Column… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-hardneg-knows-labels-v1.factprobe-replication-SUPERSEDED-olmomix-counts-v1
SUPERSEDED - do not use
Renamed 2026-08-25. Use latkes/factprobe-replication-stage1-counts-canonical-v1 instead.
It was produced by querying the infini-gram service, which indexes the corpus with the Llama-2 tokenizer, rather than by counting the corpus as OLMo token sequences. It also predates the canonical-name repair.
It is kept only so earlier numbers can be traced to where they came from. Nothing current should read it.
factprobe-replication-olmomix-counts-v1… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-SUPERSEDED-olmomix-counts-v1.inside-out-replication-v2-pamqfix-scores-full
Inside-Out Replication V2 — full corrected raw scores (all 12 cells)
This is the dataset to recompute K and K* for P(True), P(a|q),
P_norm(a|q) (and the sensitivity variants) across all 3 models × 4
relations (12 "cells"), test split.
One row per unique (question, answer) candidate — the greedy answer + the
1,000 temperature-1 samples are deduplicated to unique strings, then scored
once each. ~1.57M rows. Pipeline: judge-bug-fixed labels (pure unique-verdict
"Scheme A") + A.8.1… See the full description on the dataset page: https://huggingface.co/datasets/latkes/inside-out-replication-v2-pamqfix-scores-full.factprobe-replication-arm-a-llama31-70b-instruct-v1
factprobe-replication-arm-a-llama31-70b-instruct-v1
Arm (a) sanity replication of He et al. arXiv:2503.22362: their code, their Zenodo triples, their prompts (chat template, 36 alias variations, greedy). One row per (pair, direction, alias-variation) with the full generated text. Bucketing/analysis is done downstream (their analyse_experiment.py); OLMo cells reproduce their Tables 2/4 to +/-0.006.
Dataset Info
Rows: 5175168
Columns: 10
Columns… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-arm-a-llama31-70b-instruct-v1.p2-etf-trendfolios-replication-datafactprobe-replication-stage1-fact-counts-v1
Correction, 2026-08-26
An earlier version of this card said the corpus states facts asymmetrically
in a way that tracks entity frequency, and gave 70.9% as the figure. That
number is right, and the generalisation drawn from it was wrong.
It holds for spouse and for no other relation. Recomputed across all four:
relation
pairs with a fact sentence
written more often with the MORE frequent entity first
spouse
1,265
70.9%
sibling
203
36.5% — the opposite
twinned town… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-stage1-fact-counts-v1.eh-j-space-layer-contrast-replication-qwen3-4b
j-space-layer-contrast-replication-qwen3-4b -- aggregate exhaust
Aggregate-only: every file committed under this experiment's analysis-committed/ tree (dose-response tables, direction fits, gate AUROCs, manifests, and any other analysis artifact), copied byte-for-byte. No source question text, aliases, or per-row generation text -- analysis-committed/ never carries those.
HF repo: professorsynapse/eh-j-space-layer-contrast-replication-qwen3-4b
Provenance… See the full description on the dataset page: https://huggingface.co/datasets/professorsynapse/eh-j-space-layer-contrast-replication-qwen3-4b.factprobe-replication-generation-spouse-13b-base-v1
factprobe-replication-generation-spouse-13b-base-v1
Free-form spouse generation: for every P26 subject NAME form, the base model was asked 'Who is the {spouse} of ? Answer with just the name:' with 4 in-context demos, and sampled 5 times with nucleus sampling (top_p=0.95, temperature=1.0, max_tokens=64, stop at newline). 28,815 subject names. Companion to the P(Yes) probing datasets — this is what the model GENERATES, not a yes/no score.
Dataset Info
Rows: 30796… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-generation-spouse-13b-base-v1.inside-out-replication-results-v1
inside-out-replication-results-v1
Full Inside-Out replication: 3 models x 4 relations x 450 test questions x 1000 samples. Includes P(a|q), P_norm, P(True) V0/V1/V2, and probe scores.
Dataset Info
Rows: 1523595
Columns: 20
Columns
Column
Type
Description
question_id
Value('string')
Unique question identifier
answer
Value('string')
Model-generated answer
label
Value('string')
Judge verdict: CORRECT or INCORRECT
log_p_a_q
Value('float64')
Log… See the full description on the dataset page: https://huggingface.co/datasets/latkes/inside-out-replication-results-v1.inside-out-replication-v2-gemma-attn-ablation-large-v1
Inside-Out Replication V2 — Gemma 2x2 ablation at N=100 (confirmation)
Confirmation of today's N=25 negative result on google/gemma-2-9b-it / P26 / test, run at 4x larger N (100 question_ids → 35,612 candidate rows per
arm) to reduce subset noise. Same 2x2 design as the N=25 canary:
attn_implementation × query_pre_attn_scalar.
Headline result at N=100
| metric | arm0 (sdpa,256) | arm1 (eager,256) | arm2 (sdpa,224) | arm3 (eager,224) | max |Δ| |
|---|---|---|---|---|---|
|… See the full description on the dataset page: https://huggingface.co/datasets/latkes/inside-out-replication-v2-gemma-attn-ablation-large-v1.factprobe-replication-stage1-cooccurrence-v1
factprobe-replication-stage1-cooccurrence-v1
How many documents of the OLMo-2 pretraining corpus contain BOTH names of a probed pair. Computed over all 1,117 token files (15.50 TB, 3.875 trillion tokens) for the 2,172,383 name pairs the model was probed about; 313,576 of them share at least one document. Replaces an earlier version measured on the 192 GB mid-training mix alone, where 80-92% of pairs never co-occurred and the quantity behaved as a yes/no flag rather than a graded… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-stage1-cooccurrence-v1.factprobe-replication-hardneg-knows-labels-surface-v1
factprobe-replication-hardneg-knows-labels-surface-v1
NAME-level (surface-form) paired 'knows the fact' judgements of OLMo-2 against hard negatives, across the training ladder. One row per subject NAME form (not the entity aggregate). subject_name held fixed; object side is the any-of max. Forward direction; spouse (P26) and sibling (P3373).
Dataset Info
Rows: 1299300
Columns: 11
Columns
Column
Type
Description
tag
Value('string')
7b or… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-hardneg-knows-labels-surface-v1.factprobe-replication-exposure-knows-trajectory-v1
factprobe-replication-exposure-knows-trajectory-v1
Does EXACT per-checkpoint cumulative corpus exposure predict whether OLMo-2 KNOWS a fact? Two measures per (checkpoint, relation): PAIRED (P(Yes|true) > P(Yes|hard-negative), the 'beats' defs) and MARGINAL (P(Yes|true), r2_p_true + per-feature Spearman). Checkpoints: after-phase1, base(s1+s2), SFT, DPO, RLVR + 7 pre-merge s2-anneal ingredients; both models; both relations (P26 spouse, P3373 sibling); surface/name level; all 7… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-exposure-knows-trajectory-v1.math500-bon-prm-replication
Best-of-N Weighted Baseline with PRM — Replicating DeepMind's Test-Time Compute Scaling
Replication of the Best-of-N Weighted baseline from:
"Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters"
(Snell, Lee, Xu, Kumar — 2024) — arxiv:2408.03314
Paper Summary
The paper studies how to optimally scale inference-time computation in LLMs. The key finding: using a compute-optimal test-time strategy can improve efficiency by 4× compared… See the full description on the dataset page: https://huggingface.co/datasets/ramu3405/math500-bon-prm-replication.factprobe-replication-SUPERSEDED-stage1-counts-olmotok-v1
SUPERSEDED - do not use
Renamed 2026-08-25. Use latkes/factprobe-replication-stage1-counts-canonical-v1 instead.
It was counted with a name list missing each entity's canonical Wikidata name for 10,471 of 42,235 entities (24.8%). "Barack Obama" is not in it, and occurs 38,978,811 times in the corpus it claims to count.
It is kept only so earlier numbers can be traced to where they came from. Nothing current should read it.… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-SUPERSEDED-stage1-counts-olmotok-v1.factprobe-replication-stage1-counts-canonical-v1
factprobe-replication-stage1-counts-canonical-v1
Occurrences of every probed entity name in the OLMo-2 PRETRAINING corpus (olmo-mix-1124, 1,117 token files, 14.10 TiB), counted as OLMo token sequences with a word boundary required at each end, both written forms kept separate. SUPERSEDES factprobe-replication-stage1-counts-olmotok-v1, which was missing each entity's canonical name: the released triples name entities by their Wikidata aliases, a field that by construction… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-stage1-counts-canonical-v1.inside-out-replication-canary-v1
inside-out-replication-canary-v1
Canary run: 5 questions per relation, 50 samples, Llama-3-8B. Full pipeline E2E test.
Dataset Info
Rows: 482
Columns: 11
Columns
Column
Type
Description
relation
Value('string')
Wikidata relation (P26=spouse, P264=label, P176=manufacturer, P50=author)
question_id
Value('string')
Unique question identifier
question
Value('string')
Entity-centric question text
gold_answer
Value('string')
Ground truth answer from… See the full description on the dataset page: https://huggingface.co/datasets/latkes/inside-out-replication-canary-v1.inside-out-replication-v2-corrected-metrics-v1
Inside-Out Replication V2 — CORRECTED metrics (Option C)
K/K* for 3 models x 4 relations x scoring methods, from the corrected
pipeline: judge-postprocessing bug fixed (pure unique-verdict Scheme A),
paper-faithful Option-C train (greedy-correct + >=1 judged-incorrect),
corrected labels (flips match the no-GPU baseline exactly), paper-consistent
best layers (llama L11 / mistral L12 / gemma L24). Full test sets n~370-449.
Probe vs best-external K-gap (avg over 4… See the full description on the dataset page: https://huggingface.co/datasets/latkes/inside-out-replication-v2-corrected-metrics-v1.
