Team Ai
Datasetpublic

facebook/airs-bench

AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents The AI Research Science Benchmark (AIRS-Bench) quantifies the autonomous research abilities of LLM agents in the area of machine learning. AIRS-Bench comprises 20 tasks from state-of-the-art machine learning papers spanning diverse domains: NLP, Code, Math, biochemical modelling, and time series forecasting. Each task is specified by a ⟨problem, dataset, metric⟩ triplet and a SOTA value. The agent receives… See the full description on the dataset page: https://huggingface.co/datasets/facebook/airs-bench.

sourceHugging Facecc-by-nc-4.0updated 7mo agoView on Hugging Face
6likes70downloads
Dataset Card

AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents

![License: CC BY-NC 4.0](https://creativecommons.org/licenses/by-nc/4.0/deed.en) ![arXiv](https://arxiv.org/abs/2602.06855)

The AI Research Science Benchmark (AIRS-Bench) quantifies the autonomous research abilities of LLM agents in the area of machine learning. AIRS-Bench comprises 20 tasks from state-of-the-art machine learning papers spanning diverse domains: NLP, Code, Math, biochemical modelling, and time series forecasting.

Each task is specified by a ⟨problem, dataset, metric⟩ triplet and a SOTA value. The agent receives the full task specification and is expected to develop a solution that generates predictions on a test set, which are then evaluated and compared against the state-of-the-art (SOTA) score from a published paper.

For full details see the paper and the GitHub repository.


Dataset Description

This dataset contains the task specification files for the 20 AIRS-Bench tasks, formatted for use with the aira-dojo agentic harness.

Categories

Category# Tasks
Text Classification2
Question Answering4
Text Extraction and Matching3
Molecules and Proteins ML5
Time Series3
Code2
Math1

Data Fields

ColumnTypeDescription
taskstringTask identifier (directory name, e.g. SentimentAnalysisYelpReviewFullAccuracy)
categorystringHigh-level domain category (e.g. Text Classification, Code)
research_problemstringThe specific research problem the task addresses
datasetstringHuggingFace dataset identifier used for the task
metricstringEvaluation metric (e.g. Accuracy, MeanAbsoluteError, Rouge1)
metadata.yamlstringFull content of the task metadata file (dataset config, SOTA info, requirements)
project_description.mdstringThe task prompt provided to the agent
prepare.pystringDataset preparation script (creates train/test splits, hides test labels)
evaluate_prepare.pystringEvaluation data preparation script (creates test labels for scoring)
evaluate.pystringEvaluation script used to score the agent's submission
custom_labels.pystringOptional custom label handler for non-standard label formats (empty if unused)
utils.pystringOptional shared utilities across task scripts (empty if unused)

Citation

bibtex
@article{lupidi2026airsbenchsuitetasksfrontier,
      title={AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents},
      author={Alisia Lupidi and Bhavul Gauri and Thomas Simon Foster and Bassel Al Omari and Despoina Magka and Alberto Pepe and Alexis Audran-Reiss and Muna Aghamelu and Nicolas Baldwin and Lucia Cipolina-Kun and Jean-Christophe Gagnon-Audet and Chee Hau Leow and Sandra Lefdal and Hossam Mossalam and Abhinav Moudgil and Saba Nazir and Emanuel Tewolde and Isabel Urrego and Jordi Armengol Estape and Amar Budhiraja and Gaurav Chaurasia and Abhishek Charnalia and Derek Dunfield and Karen Hambardzumyan and Daniel Izcovich and Martin Josifoski and Ishita Mediratta and Kelvin Niu and Parth Pathak and Michael Shvartsman and Edan Toledo and Anton Protopopov and Roberta Raileanu and Alexander Miller and Tatiana Shavrina and Jakob Foerster and Yoram Bachrach},
      year={2026},
      eprint={2602.06855},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2602.06855},
}

License

This dataset is released under the CC BY-NC 4.0 license.