datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
replicationbenchReplicationBench
ReplicationBench
arXiv: ReplicationBench: Can AI Agents Replicate Astrophysics Research Papers?
GitHub: https://github.com/Christine8888/replicationbench-release
Dataset Description
The ReplicationBench dataset contains 111 astrophysics research replication tasks, spanning complete replications of 20 research papers. The dataset includes:
Original and masked manuscript text
Metadata (title, abstract, publication info, etc.)
Pointers to datasets and dataset access… See the full description on the dataset page: https://huggingface.co/datasets/ChristineYe8/ReplicationBench.jmh-grpo-replication-data
JMH-Bench GRPO replication data
Data for replicating the BSc thesis Reinforcement Learning with Verifiable Execution Rewards
for Synthesizing Regression-Detecting Microbenchmarks. The code is in the replication repository
RL4JMH-thesis-submission. The trained models are
run 9 (evaluated in the
thesis) and run 8.
Path
Contents
classpaths/
Mutant-bearing jars and dependency jars of the 246 training libraries, as used by GRPO run 9 (1,764 jars, 1.2 GB)… See the full description on the dataset page: https://huggingface.co/datasets/bookxd/jmh-grpo-replication-data.replicationbench
ReplicationBench Dataset
Dataset Description
A benchmark to evaluate AI agents in astrophysics research through replicating existing research papers.
Dataset Structure
Data Splits
ReplicationBench (source: epxert): Core expert-written benchmark
ReplicationBench-Plus (source: showyourwork): Extension dataset generated through hybrid LLM-expert system
Data Configurations
Each split contains two types of data:
metadata: Paper metadata and… See the full description on the dataset page: https://huggingface.co/datasets/replicationbench-submission/replicationbench.dpst-replication
DPST Replication
Replication package for the EMNLP 2025 paper:
Leveraging Semantic Triples for Private Document Generation with Local Differential Privacy GuaranteesStephen Meisenbacher, Maulik Chevli, Florian Mattheshttps://aclanthology.org/2025.emnlp-main.455/
All credit goes to the original authors. This repository provides a replication environment for running the method, based on the official released code: https://github.com/sjmeis/DPST.… See the full description on the dataset page: https://huggingface.co/datasets/weijunl/dpst-replication.oai-hf-incident-replication
OpenAI–Hugging Face Incident Replication
Raw agent transcripts behind the write-up
"OpenAI–Hugging Face: A Reproduction and Lessons for Alignment".
They capture attempts to elicit the four pivotal misaligned behaviours from the
May–July 2026 incident that Hugging Face described here:
Inappropriate writes to shared infrastructure (a mock model registry).
Requesting help from other agents to get around a blocker.
Sharing solutions/exploits with peers, including an… See the full description on the dataset page: https://huggingface.co/datasets/stewy33/oai-hf-incident-replication.factprobe-replication-negatives-allnames-v1
Plausible wrong answers, asked under every name (42,267,800 rows)
Status: final — 40 of 40 runs. Models present:
13b, 7b. Training stages present: s1, s2, s3, s4, s5.
What this fixes
A real fact is put to the model under the full cross product of the two
people's name lists, and counts as recognised if any one combination gets a
Yes. That is He et al.'s rule. Until 2026-08-26 the wrong answer it was
compared against was asked under one name per person, so the… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-negatives-allnames-v1.lifemem-replication-src
LifeMem replication (arXiv:2608.19621)
Replication of Mitigating Identity Essentialism in LLM Agents with Longitudinal Life Trajectories (Wang, Zhou, Du, Su, Cao, Pan, Ai, Wu, Zhang, Liu). Official code: halsayxi/LifeMem.
The paper's claim: static demographic prompts make LLM survey agents essentialist — within-group answers collapse, SES clusters separate. LifeMem stores life events in a hippocampal retriever (top-K=5, α=0.9) and consolidates them into a per-agent LoRA adapter.… See the full description on the dataset page: https://huggingface.co/datasets/mtorres98/lifemem-replication-src.crimsonred-paper-replication
CrimsonRed — Cross-Architecture Emotion-Prime Steering Replication
Replication of the emotion-prime steering protocol from arXiv:2607.18691 (NSM
semantic primes as explanans for emotion in LLMs), extended across four
architectures. Generated by scripts/paper_faithful_steering.py in the
CrimsonRed project.
The finding
The paper's core claim — that semantic-prime recipe directions steer emotion more
strongly than Scherer appraisal directions — replicates… See the full description on the dataset page: https://huggingface.co/datasets/musicakamusic/crimsonred-paper-replication.cna-vulnerability-census-replication
Empirical Vulnerability Census (1999–2026, $N = 385,524$), Cybernetic Queueing Instability, and CISA BOD 26-04 Remediation Deficit
Deterministic Empirical Replication Package & Econometric Audits
Principal Investigator: Gia Bao Huynh (Jun Huynh)ORCID: 0009-0008-2372-5852Affiliation: Independent Scholar / Ho Chi Minh City, VietnamLive Interactive Simulator: Cybernetic Queueing Instability Simulator (M/G/1)
🏛️ Executive Summary & Theoretical… See the full description on the dataset page: https://huggingface.co/datasets/giabaohuynhasu/cna-vulnerability-census-replication.frechesTest_replication
Clean Data Directory for Replication Archive
This directory contains only the data files that are actually used by the replication code.
Directory Structure
data_clean/in/
Input data files used by the analysis:
sipp_raw/: Raw SIPP data files (12 zip files)
Contains Survey of Income and Program Participation data from 2018-2023 only
Files follow pattern: pu{year}_dta.zip and rw{year}_dta.zip
sipp_unzipped/: Pre-unzipped SIPP data files (6 .dta… See the full description on the dataset page: https://huggingface.co/datasets/minhhuynhcong/frechesTest_replication.HE_Ashimwe_222127212_CRIL_REPLICATION_DATASETmeituna-longcat-replications
MeituNa LongCat Open-Ended Paper Replication Portfolio
Generated: 2026-07-31T12:15:27.895345+00:00
Multi-instance Meituan LongCat-2.0 workers produce multi-file open-source
replication scaffolds (method breakdown, code modules, experiment plans, scientific notes)
grounded in real arXiv paper text — not single-file stubs.
Run metadata
Parents: 8
Papers: 16
Model: LongCat-2.0
Passes: ['replication_plan', 'code_modules', 'experiments_and_critique'… See the full description on the dataset page: https://huggingface.co/datasets/Mieaz/meituna-longcat-replications.details_keyfan__vicuna-chinese-replication-v1.1
Dataset Card for Evaluation run of keyfan/vicuna-chinese-replication-v1.1
Dataset Summary
Dataset automatically created during the evaluation run of model keyfan/vicuna-chinese-replication-v1.1 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_keyfan__vicuna-chinese-replication-v1.1.factprobe-replication-traj-b2-stage1-v1
factprobe-replication-traj-b2-stage1-v1
Checkpoint-trajectory probing: b2 few-shot recognition on 13 log-spaced stage-1 checkpoints per model, spouse+sibling, both templates, original alternating demos, fp16 pinned. Identity columns model_tag/revision/tokens_b injected from filenames. Complete checkpoint files only.
Dataset Info
Rows: 26315536
Columns: 18
Columns
Column
Type
Description
relation
Value('string')
No description provided… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-traj-b2-stage1-v1.gz3d-aion-replication-attempt
GalaxyZoo 3D Segmentation Benchmark
An attempted replication of the GZ3D segmentation dataset used in section 7.2.3 in the AION-1. This dataset contains volunteer segmentations for the following channels: center, star, spiral, bar. It also contains the RGB images from Legacy Survey and the tokens from the AION-1 image tokenizer of the Legacy Survey image bands. Notably, these are pre–AION-1 transformer encoder, but post image-codec encoder/quantization. The reason why it's an… See the full description on the dataset page: https://huggingface.co/datasets/astronolan/gz3d-aion-replication-attempt.tempest-replication
TEMPEST Replication Dataset
Multi-turn adversarial attack results on 10 frontier LLMs.
Dataset Description
This dataset contains results from replicating the TEMPEST multi-turn jailbreak framework across 10 frontier language models, each evaluated on 100 harmful behaviors from JailbreakBench.
Key Findings
ASR range: 42-100% - All models vulnerable
No scale-safety correlation (r=-0.12)
Thinking mode helps: Kimi K2 Thinking (42%) vs standard (97%)
Files… See the full description on the dataset page: https://huggingface.co/datasets/richardyoung/tempest-replication.frechesTest_replication
Clean Data Directory for Replication Archive
This directory contains only the data files that are actually used by the replication code.
Directory Structure
data_clean/in/
Input data files used by the analysis:
sipp_raw/: Raw SIPP data files (12 zip files)
Contains Survey of Income and Program Participation data from 2018-2023 only
Files follow pattern: pu{year}_dta.zip and rw{year}_dta.zip
sipp_unzipped/: Pre-unzipped SIPP data files (6 .dta files + 1… See the full description on the dataset page: https://huggingface.co/datasets/dvdijcke/frechesTest_replication.spacr-example-replication
spaCR example data: replication
Parasites per vacuole, for spaCR's Replication Assay module. Every parasite object of twelve control wells of the Toxoplasma MTOC screen, as the screen's own Mask and Measure runs segmented and measured them.
Download it from inside spaCR with Load test data... in the
Replication module, or with spacr-download replication. The
settings file ships beside the data, so after loading, Run is the next step.
Size… See the full description on the dataset page: https://huggingface.co/datasets/einarolafsson/spacr-example-replication.reach-of-the-state-replication
Replication Data for: The Reach of the State
This dataset is the official replication package for:
Chang, Charles, and Yuhua Wang. 2024. "The Reach of the State." Comparative Political Studies 57(8): 1243–1275. https://doi.org/10.1177/00104140231194057
It is a byte-for-byte mirror of the replication package archived on Harvard Dataverse
(doi:10.7910/DVN/CSBRZX, license CC0 1.0). If you use this package, please cite the
original paper above. The canonical archival version… See the full description on the dataset page: https://huggingface.co/datasets/fantasien/reach-of-the-state-replication.Replication_EEREV-D-25-01006
Replication materials for Culture and constitutional compliance
Replication materials for:
https://arxiv.org/abs/2608.23369
sshfighter-17-head-router-delay-replication-v1
SSH Fighter 17-Head Frozen-Expert Router v1 — Delay Stress
Seventeen character-specific heads route among nine immutable Agent Gym
policies. The package contains every three-seed checkpoint, averaged inference
weights, the grouped training tensor cache, and fresh-seed exact-engine
evaluation rows. This package is the predeclared six-scenario role-delay stress matrix.
Fresh exact-engine evaluation
condition
points rate
wins
losses
draws
switches/match… See the full description on the dataset page: https://huggingface.co/datasets/LisaMegaWatts/sshfighter-17-head-router-delay-replication-v1.lifemem-replication-2026-09-09e2emet_replication_packagespacehpc-feature-vjepa21l256-pusht-replication192-commonlneh-no-abstention-prompt-gated-replication
no-abstention-prompt-gated-replication -- aggregate exhaust
Aggregate-only: every file committed under this experiment's analysis-committed/ tree (dose-response tables, direction fits, gate AUROCs, manifests, and any other analysis artifact), copied byte-for-byte. No source question text, aliases, or per-row generation text -- analysis-committed/ never carries those.
HF repo: professorsynapse/eh-no-abstention-prompt-gated-replication
Provenance
Experiment:… See the full description on the dataset page: https://huggingface.co/datasets/professorsynapse/eh-no-abstention-prompt-gated-replication.factprobe-replication-b2-olmo7b-stage1-v1
factprobe-replication-b2-olmo7b-stage1-v1
Arm b2 of the FactProbe replication (He et al. arXiv:2503.22362): few-shot completion probing of the BASE OLMo-2-1124-7B final stage-1 checkpoint on symmetric-relation fact recognition (spouse P26, sibling P3373), forward and backward directions, all alias variations. One row per prompt; p_yes/p_no are next-token probability mass over Yes/No (question) or True/False (statement) single-token variants.
Dataset Info
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-b2-olmo7b-stage1-v1.factprobe-replication-asked-negatives-v1
Asking a model for a plausible wrong answer (100,776 questions)
Two prompts, put to gpt-5.1 for every subject surface form in the spouse
and sibling data, ten independent samples each:
(a) "Who is the spouse of X? Just the name, no explanation needed."
(b') "Name a person who could reasonably be mistaken for the spouse of X,
but is not the spouse of X. Just the name, no explanation needed."
50,388 surface forms across 10,592 entities, times two prompts, is
100,776 questions… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-asked-negatives-v1.spacehpc-feature-vjepa2l-pusht-replication192-commonlnfactprobe-replication-b2-olmo13b-stage1-v1
factprobe-replication-b2-olmo13b-stage1-v1
Arm b2 of the FactProbe replication (He et al. arXiv:2503.22362): few-shot completion probing of the BASE OLMo-2-1124-7B final stage-1 checkpoint on symmetric-relation fact recognition (spouse P26, sibling P3373), forward and backward directions, all alias variations. One row per prompt; p_yes/p_no are next-token probability mass over Yes/No (question) or True/False (statement) single-token variants.
Dataset Info
Rows:… See the full description on the dataset page: https://huggingface.co/datasets/latkes/factprobe-replication-b2-olmo13b-stage1-v1.
