datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
replicationbench
ReplicationBench Dataset
Dataset Description
A benchmark to evaluate AI agents in astrophysics research through replicating existing research papers.
Dataset Structure
Data Splits
ReplicationBench (source: epxert): Core expert-written benchmark
ReplicationBench-Plus (source: showyourwork): Extension dataset generated through hybrid LLM-expert system
Data Configurations
Each split contains two types of data:
metadata: Paper metadata and… See the full description on the dataset page: https://huggingface.co/datasets/replicationbench-submission/replicationbench.crimsonred-paper-replication
CrimsonRed — Cross-Architecture Emotion-Prime Steering Replication
Replication of the emotion-prime steering protocol from arXiv:2607.18691 (NSM
semantic primes as explanans for emotion in LLMs), extended across four
architectures. Generated by scripts/paper_faithful_steering.py in the
CrimsonRed project.
The finding
The paper's core claim — that semantic-prime recipe directions steer emotion more
strongly than Scherer appraisal directions — replicates… See the full description on the dataset page: https://huggingface.co/datasets/musicakamusic/crimsonred-paper-replication.
