datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenMLE-Tasks
OpenMLE Tasks
📄 Paper
•
🌐 Project
•
💻 Code
•
🤗 Models
•
📚 SFT Traces
OpenMLE Tasks provides machine-learning Task environments. The public SFT trajectories are released separately in OpenMLE-SFT-Traces. These resources accompany the paper Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering and the OpenRSI code release.
Release form
What is included… See the full description on the dataset page: https://huggingface.co/datasets/FrontisAI/OpenMLE-Tasks.proofwriter
Dataset Card for "proofwriter"
More Information needed
osworld_v2_tasks
OSWorld V2 Task Classes
This gated dataset contains the official root-level task_*.py Python task classes for OSWorld V2.
The public GitHub repository keeps the task loader, helper utilities, and documentation. The task implementations are gated to reduce benchmark leakage and to help prevent evaluated agents from finding task answers, setup logic, or evaluator details online while executing a task.
Download from the public repository root with:
uvx --from huggingface_hub hf… See the full description on the dataset page: https://huggingface.co/datasets/xlangai/osworld_v2_tasks.esci
Dataset Card for "esci"
ESCI product search dataset
https://github.com/amazon-science/esci-data/
Preprocessings:
-joined the two relevant files
-product_text aggregate all product text
-mapped esci_label to full name
@article{reddy2022shopping,
title={Shopping Queries Dataset: A Large-Scale {ESCI} Benchmark for Improving Product Search},
author={Chandan K. Reddy and Lluís Màrquez and Fran Valero and Nikhil Rao and Hugo Zaragoza and Sambaran Bandyopadhyay and Arnab Biswas and Anlu… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/esci.foliohttps://github.com/Yale-LILY/FOLIO
@article{han2022folio,
title={FOLIO: Natural Language Reasoning with First-Order Logic},
author = {Han, Simeng and Schoelkopf, Hailey and Zhao, Yilun and Qi, Zhenting and Riddell, Martin and Benson, Luke and Sun, Lucy and Zubova, Ekaterina and Qiao, Yujie and Burtell, Matthew and Peng, David and Fan, Jonathan and Liu, Yixin and Wong, Brian and Sailor, Malcolm and Ni, Ansong and Nan, Linyong and Kasai, Jungo and Yu, Tao and Zhang, Rui and Joty, Shafiq and… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/folio.finance-tasks
Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024)
This repo contains the evaluation datasets for our paper Adapting Large Language Models via Reading Comprehension.
We explore continued pre-training on domain-specific corpora for large language models. While this approach enriches LLMs with domain knowledge, it significantly hurts their prompting ability for question answering. Inspired by human learning via reading comprehension, we propose a simple method to… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/finance-tasks.procedural-typed-decisions
procedural-typed-decisions
Procedurally generated decision problems. Each row is one structured state
(JSON, or a table, CSV, key=value lines, or prose for the arithmetic,
retrieval, and aggregation configs) with several typed questions over that same state, following the
Jev / System One request shape: choice (pick one criterion), noul (a
number in [0, 1]; a probability or a yes/no), and score (an ordered rubric).
Every answer is computed exactly from the state by rules that… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/procedural-typed-decisions.rna-downstream-tasks
GB.RNA Benchmark Datasets
mRNA related tasks
Translation efficiency prediction from Chu et al.(2024) [1]
3 cell lines: Muscle, pc3, HEK
input sequence: 5'UTR
10-fold cross-validation split
mRNA expression level prediction from Chu et al.(2024) [1]
3 cell lines: Muscle, pc3, HEK
input sequence: 5'UTR
10-fold cross-validation split
Mean ribosome load prediction from Sample et al. (2019) [2]
input sequence: 5'UTR
ouput: mean ribosome load
the original data… See the full description on the dataset page: https://huggingface.co/datasets/genbio-ai/rna-downstream-tasks.RoboTwin-LeRobot-unseen-tasks-cross-embPlantCAD2_zero_shot_tasks
🌱 PlantCAD2 Zero-Shot Tasks
Zero-shot evaluation tasks for plant genomics using PlantCAD2.This dataset contains tasks designed to evaluate model performance without task-specific training.
📂 Available Tasks
🔬 Cross-species Evolutionary Conservation
Task Name
Description
Samples
Metric
conservation_within_andropogoneae
Predict conserved vs non-conserved sites using alignments within 35 Andropogoneae genomes
19,030 vs 19,030
AUROC… See the full description on the dataset page: https://huggingface.co/datasets/plantcad/PlantCAD2_zero_shot_tasks.agent-harbor-tasks
🐚 Agent Harbor Tasks
400 original shell-agent tasks, Italian and English, from a 4-call chore to a 20-call investigation.
Each task drops an agent into a container with a small, realistic file system (a home folder, a repository, a
server's logs and configs, an office archive) and a request written the way a person would type it. The agent works
in the shell, and hidden tests decide the reward deterministically: no LLM in the reward path. For every task the
reference… See the full description on the dataset page: https://huggingface.co/datasets/efederici/agent-harbor-tasks.Tmax-Tasks-Clean
Tmax-Tasks-Clean
New: longlongcheck (2026-09-14)
Use configuration longlongcheck, split longlongcheck, for the 431-task snapshot combining the selected Codex and Claude Code repairs with passing historical GPT-6 terminal solutions. The existing splits below retain their earlier data.
From the latest local 452-task repaired snapshot, this split holds out the requested 15 old-pass/current-fail tasks, four additional tasks without any passing current GPT-6 replay… See the full description on the dataset page: https://huggingface.co/datasets/Fzz1/Tmax-Tasks-Clean.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.pdbthink-coordinate-tasks
PDBThink Coordinate Tasks
100,000 new coordinate-interpretation tasks across all 19 active PDBThink families,
from 2,671 experimental PDB entries in 1,867 source groups.
Version 1.3.0; deterministic seed 2026100101.
The model receives sanitised, rotated, rounded protein coordinates and a question.
It must answer without tools. This release contains no sequence-to-structure
prediction tasks and no retired MECH tasks. It is intended for additional evaluation,
RL with deterministic… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/pdbthink-coordinate-tasks.Lightwheel-Tasks-G1-ControllerThis dataset was created using LeRobot.
1 Dataset Description
This dataset includes 62 Lightwheel-Robocasa-Tasks, collected using the g1-controller robot in environments provided by LW-BenchHub.The robot configuration used during data collection is G1-Controller, using UnitreeG1ControllerEnvCfg inherits from UnitreeG1EnvCfg.
1.1 Robot State
The robot state recorded in the environment is stored in: observation.state
Dimension: 31
Description: Joint positions of… See the full description on the dataset page: https://huggingface.co/datasets/LightwheelAI/Lightwheel-Tasks-G1-Controller.chaos-mnli-ambiguity
chaos-mnli-ambiguity
ChaosNLI, MNLI portion: 1,599 MNLI pairs relabeled by 100 annotators each (Nie et al., 2020).
label_dist and label_count follow the entailment/neutral/contradiction order, and gini is the Gini
coefficient of label_dist (0 = annotators evenly split, 1 = unanimous). Built from the jsonl first uploaded
here, which flattens the ChaosNLI release (https://github.com/easonnie/ChaosNLI) and adds gini; the
variable-key label_counter (a duplicate of label_count) is… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/chaos-mnli-ambiguity.Lightwheel-Tasks-X7SThis dataset was created using LeRobot.
1 Dataset Description
This dataset includes 117 Lightwheel-Libero-Tasks and 89 Lightwheel-Robocasa-Tasks, collected using the x7s robot in environments provided by LW-BenchHub.The robot configuration used during data collection is X7s-Abs, using X7SAbsEnvCfg inherits from X7SEnvCfg.
1.1 Robot State
The robot state recorded in the environment is stored in: observation.state
Dimension: 25
Description: Joint positions of… See the full description on the dataset page: https://huggingface.co/datasets/LightwheelAI/Lightwheel-Tasks-X7S.Lightwheel-Tasks-G1-WBCThis dataset was created using LeRobot.
1 Dataset Description
This dataset includes 32 Lightwheel-Robocasa-Tasks, collected using the g1-wbc robot in environments provided by LW-BenchHub.The robot configuration used during data collection is G1-Controller-DecoupledWBC, using UnitreeG1ControllerDecoupledWBCEnvCfg inherits from UnitreeG1ControllerEnvCfg.
1.1 Robot State
The robot state recorded in the environment is stored in: observation.state
Dimension: 43… See the full description on the dataset page: https://huggingface.co/datasets/LightwheelAI/Lightwheel-Tasks-G1-WBC.Math-RL-Tasks
Ulam AI Math RL Tasks
Forty original, verifier-backed mathematical reasoning tasks packaged as ten
independent RL environments. The collection spans advanced graduate exercises,
research-style exact computation and structural generalization problems in
algebraic geometry, arithmetic geometry, combinatorics, topology, probability
and spectral analysis.
Each suite pairs a runnable rl_env/ with a preserved blind_run/ by
GPT-5.6 Sol Pro. The model name describes the evaluation actor… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/Math-RL-Tasks.Lightwheel-Tasks-Double-PiperThis dataset was created using LeRobot.
1 Dataset Description
This dataset includes 130 Lightwheel-Libero-Tasks, collected using the double-piper robot in environments provided by LW-BenchHub.The robot configuration used during data collection is DoublePiper-Abs, using DoublePiperAbsEnvCfg inherits from DoublePiperEnvCfg.
1.1 Robot State
The robot state recorded in the environment is stored in: observation.state
Dimension: 16
Description: Joint positions of… See the full description on the dataset page: https://huggingface.co/datasets/LightwheelAI/Lightwheel-Tasks-Double-Piper.jigsaw_toxicitybioR_tasksAudio-Video-Engineering-Agentic-Tasks-1M
Audio/Video Engineering Agentic Tasks (1M)
Abstract
A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.medicine-tasks
Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024)
This repo contains the evaluation datasets for our paper Adapting Large Language Models via Reading Comprehension.
We explore continued pre-training on domain-specific corpora for large language models. While this approach enriches LLMs with domain knowledge, it significantly hurts their prompting ability for question answering. Inspired by human learning via reading comprehension, we propose a simple method to… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/medicine-tasks.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.planbench
Dataset Card for "planbench"
https://arxiv.org/abs/2206.10498
@article{valmeekam2024planbench,
title={Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change},
author={Valmeekam, Karthik and Marquez, Matthew and Olmo, Alberto and Sreedharan, Sarath and Kambhampati, Subbarao},
journal={Advances in Neural Information Processing Systems},
volume={36},
year={2024}
}
QuALITY
Dataset Card for "QuALITY"
@article{bowman2022quality,
title={QuALITY: Question Answering with Long Input Texts, Yes!},
author={Bowman, Samuel R and Chen, Angelica and He, He and Joshi, Nitish and Ma, Johnny and Nangia, Nikita and Padmakumar, Vishakh and Pang, Richard Yuanzhe and Parrish, Alicia and Phang, Jason and others},
journal={NAACL 2022},
year={2022}
}
Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.PlantCAD2_fine_tuning_tasksPlantCAD2_zero_shot_tasks
🌱 PlantCAD2 Zero-Shot Tasks
Zero-shot evaluation tasks for plant genomics using PlantCAD2.This dataset contains tasks designed to evaluate model performance without task-specific training.
📂 Available Tasks
🔬 Cross-species Evolutionary Conservation
Task Name
Description
Samples
Metric
conservation_within_andropogoneae
Predict conserved vs non-conserved sites using alignments within 35 Andropogoneae genomes
19,030 vs 19,030
AUROC… See the full description on the dataset page: https://huggingface.co/datasets/Yangximiao/PlantCAD2_zero_shot_tasks.
