datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eh-j-space-layer-contrast-replication-qwen3-4b
j-space-layer-contrast-replication-qwen3-4b -- aggregate exhaust
Aggregate-only: every file committed under this experiment's analysis-committed/ tree (dose-response tables, direction fits, gate AUROCs, manifests, and any other analysis artifact), copied byte-for-byte. No source question text, aliases, or per-row generation text -- analysis-committed/ never carries those.
HF repo: professorsynapse/eh-j-space-layer-contrast-replication-qwen3-4b
Provenance… See the full description on the dataset page: https://huggingface.co/datasets/professorsynapse/eh-j-space-layer-contrast-replication-qwen3-4b.math500-bon-prm-replication
Best-of-N Weighted Baseline with PRM — Replicating DeepMind's Test-Time Compute Scaling
Replication of the Best-of-N Weighted baseline from:
"Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters"
(Snell, Lee, Xu, Kumar — 2024) — arxiv:2408.03314
Paper Summary
The paper studies how to optimally scale inference-time computation in LLMs. The key finding: using a compute-optimal test-time strategy can improve efficiency by 4× compared… See the full description on the dataset page: https://huggingface.co/datasets/ramu3405/math500-bon-prm-replication.
