Team Ai
Datasetpublic

AlgorithmicResearchGroup/aria-repo-benchmark

ARIA Repo Benchmark The ARIA Repo Benchmark is part of the ARIA benchmark suite, a collection of closed-book benchmarks probing the ML knowledge that frontier models have internalized during training. This dataset contains 58 curated research paper implementations with metadata for evaluating whether ML experiments described in papers can be reproduced. Dataset Summary Size: 58 entries Coverage: Computer Vision, NLP, Time Series, Graph, Bioinformatics Purpose:… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/aria-repo-benchmark.

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes13downloads
Dataset Card

ARIA Repo Benchmark

The ARIA Repo Benchmark is part of the ARIA benchmark suite, a collection of closed-book benchmarks probing the ML knowledge that frontier models have internalized during training. This dataset contains 58 curated research paper implementations with metadata for evaluating whether ML experiments described in papers can be reproduced.

Dataset Summary

  • —Size: 58 entries
  • —Coverage: Computer Vision, NLP, Time Series, Graph, Bioinformatics
  • —Purpose: Evaluate AI agents on their ability to locate, understand, and reproduce ML research experiments

Each entry links a research paper to its code repository, dataset, metrics, and compute requirements, along with verification of whether the experiment is reproducible on constrained hardware.

Dataset Structure

Key Fields

FieldTypeDescription
paper_titlestringTitle of the research paper
paper_urlstringArXiv URL
paper_datetimestampPublication date
paper_textstringFull paper text
datasetstringDataset used in the paper
dataset_linkstringLink to the dataset
model_namestringModel name
code_linkslist[string]GitHub repository links
metricsstringPerformance metrics reported
table_metricslist[string]Detailed metrics from tables
promptsstringEvaluation prompts
modalitystringData modality (CV, NLP, Time Series, Graph, etc.)

Compute & Reproducibility Fields

FieldTypeDescription
compute_hoursfloat64Estimated training compute hours
num_gpusint64Number of GPUs required
reasoningstringReasoning about compute estimates
trainable_single_gpu_8hstringTrainable on a single GPU in 8 hours
verifiedstringVerification status
time_and_compute_verificationstringCompute verification notes
link_to_colab_notebookstringGoogle Colab notebook link
run_possiblestringWhether the code runs successfully
notesstringAdditional notes

Usage

python
from datasets import load_dataset

ds = load_dataset("AlgorithmicResearchGroup/aria-repo-benchmark", split="train")

for entry in ds:
    print(f"{entry['paper_title']} - {entry['modality']} - Reproducible: {entry['run_possible']}")

Related Resources

Citation

bibtex
@misc{aria_repo_benchmark,
    title={ARIA Repo Benchmark},
    author={Algorithmic Research Group},
    year={2024},
    publisher={Hugging Face},
    url={https://huggingface.co/datasets/AlgorithmicResearchGroup/aria-repo-benchmark}
}