AlgorithmicResearchGroup/aria-repo-benchmark
ARIA Repo Benchmark The ARIA Repo Benchmark is part of the ARIA benchmark suite, a collection of closed-book benchmarks probing the ML knowledge that frontier models have internalized during training. This dataset contains 58 curated research paper implementations with metadata for evaluating whether ML experiments described in papers can be reproduced. Dataset Summary Size: 58 entries Coverage: Computer Vision, NLP, Time Series, Graph, Bioinformatics Purpose:… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/aria-repo-benchmark.
ARIA Repo Benchmark
The ARIA Repo Benchmark is part of the ARIA benchmark suite, a collection of closed-book benchmarks probing the ML knowledge that frontier models have internalized during training. This dataset contains 58 curated research paper implementations with metadata for evaluating whether ML experiments described in papers can be reproduced.
Dataset Summary
- Size: 58 entries
- Coverage: Computer Vision, NLP, Time Series, Graph, Bioinformatics
- Purpose: Evaluate AI agents on their ability to locate, understand, and reproduce ML research experiments
Each entry links a research paper to its code repository, dataset, metrics, and compute requirements, along with verification of whether the experiment is reproducible on constrained hardware.
Dataset Structure
Key Fields
Compute & Reproducibility Fields
Usage
from datasets import load_dataset
ds = load_dataset("AlgorithmicResearchGroup/aria-repo-benchmark", split="train")
for entry in ds:
print(f"{entry['paper_title']} - {entry['modality']} - Reproducible: {entry['run_possible']}")Related Resources
Citation
@misc{aria_repo_benchmark,
title={ARIA Repo Benchmark},
author={Algorithmic Research Group},
year={2024},
publisher={Hugging Face},
url={https://huggingface.co/datasets/AlgorithmicResearchGroup/aria-repo-benchmark}
}