AlgorithmicResearchGroup/aria-search-benchmark_v2-public
ARIA Search Benchmark v2 The ARIA Search Benchmark is part of the ARIA benchmark suite, a collection of closed-book benchmarks probing the ML knowledge that frontier models have internalized during training. This dataset tests whether models can answer factual questions about ML research papers, models, datasets, and benchmark results without access to external retrieval. Dataset Summary Size: 3,517 question-answer pairs Split: benchmark Paper date range: May… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/aria-search-benchmark_v2-public.
ARIA Search Benchmark v2
The ARIA Search Benchmark is part of the ARIA benchmark suite, a collection of closed-book benchmarks probing the ML knowledge that frontier models have internalized during training. This dataset tests whether models can answer factual questions about ML research papers, models, datasets, and benchmark results without access to external retrieval.
Dataset Summary
- Size: 3,517 question-answer pairs
- Split:
benchmark - Paper date range: May 2023 to December 2024
- Coverage: Spans models, datasets, and metrics across CV, NLP, audio, video, and multimodal domains
Dataset Structure
Usage
from datasets import load_dataset
ds = load_dataset("AlgorithmicResearchGroup/aria-search-benchmark_v2-public", split="benchmark")
for example in ds.select(range(5)):
print(f"Q: {example['prompts']}")
print(f"A: {example['answer']}")
print(f"Paper: {example['paper_title']}")
print()Related Resources
Citation
@misc{aria_search_benchmark_v2,
title={ARIA Search Benchmark v2},
author={Algorithmic Research Group},
year={2024},
publisher={Hugging Face},
url={https://huggingface.co/datasets/AlgorithmicResearchGroup/aria-search-benchmark_v2-public}
}