Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01FrontisAI /OpenMLE-Tasks OpenMLE Tasks 📄 Paper &nbsp;•&nbsp; 🌐 Project &nbsp;•&nbsp; 💻 Code &nbsp;•&nbsp; 🤗 Models &nbsp;•&nbsp; 📚 SFT Traces OpenMLE Tasks provides machine-learning Task environments. The public SFT trajectories are released separately in OpenMLE-SFT-Traces. These resources accompany the paper Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering and the OpenRSI code release. Release form What is included… See the full description on the dataset page: https://huggingface.co/datasets/FrontisAI/OpenMLE-Tasks.tabular5 likes15k downloads1mo agoHugging Face02tasksource /proofwriter Dataset Card for "proofwriter" More Information needed tabular100K<n<1M12 likes9.5k downloads3y agoHugging Face03xlangai /osworld_v2_tasksgated OSWorld V2 Task Classes This gated dataset contains the official root-level task_*.py Python task classes for OSWorld V2. The public GitHub repository keeps the task loader, helper utilities, and documentation. The task implementations are gated to reduce benchmark leakage and to help prevent evaluated agents from finding task answers, setup logic, or evaluator details online while executing a task. Download from the public repository root with: uvx --from huggingface_hub hf… See the full description on the dataset page: https://huggingface.co/datasets/xlangai/osworld_v2_tasks.tabularn<1K37 likes5.4k downloads1mo agoHugging Face04tasksource /esci Dataset Card for "esci" ESCI product search dataset https://github.com/amazon-science/esci-data/ Preprocessings: -joined the two relevant files -product_text aggregate all product text -mapped esci_label to full name @article{reddy2022shopping, title={Shopping Queries Dataset: A Large-Scale {ESCI} Benchmark for Improving Product Search}, author={Chandan K. Reddy and Lluís Màrquez and Fran Valero and Nikhil Rao and Hugo Zaragoza and Sambaran Bandyopadhyay and Arnab Biswas and Anlu… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/esci.tabulartext-classification1M<n<10M9 likes3.6k downloads3y agoHugging Face05tasksource /foliohttps://github.com/Yale-LILY/FOLIO @article{han2022folio, title={FOLIO: Natural Language Reasoning with First-Order Logic}, author = {Han, Simeng and Schoelkopf, Hailey and Zhao, Yilun and Qi, Zhenting and Riddell, Martin and Benson, Luke and Sun, Lucy and Zubova, Ekaterina and Qiao, Yujie and Burtell, Matthew and Peng, David and Fan, Jonathan and Liu, Yixin and Wong, Brian and Sailor, Malcolm and Ni, Ansong and Nan, Linyong and Kasai, Jungo and Yu, Tao and Zhang, Rui and Joty, Shafiq and… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/folio.tabulartext-classification1K<n<10K19 likes3k downloads3y agoHugging Face06AdaptLLM /finance-tasks Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024) This repo contains the evaluation datasets for our paper Adapting Large Language Models via Reading Comprehension. We explore continued pre-training on domain-specific corpora for large language models. While this approach enriches LLMs with domain knowledge, it significantly hurts their prompting ability for question answering. Inspired by human learning via reading comprehension, we propose a simple method to… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/finance-tasks.tabulartext-classification10K<n<100K83 likes2.1k downloads2y agoHugging Face07tasksource /procedural-typed-decisions procedural-typed-decisions Procedurally generated decision problems. Each row is one structured state (JSON, or a table, CSV, key=value lines, or prose for the arithmetic, retrieval, and aggregation configs) with several typed questions over that same state, following the Jev / System One request shape: choice (pick one criterion), noul (a number in [0, 1]; a probability or a yes/no), and score (an ordered rubric). Every answer is computed exactly from the state by rules that… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/procedural-typed-decisions.tabulartext-classification100K<n<1M4 likes1.9k downloads11d agoHugging Face08genbio-ai /rna-downstream-tasks GB.RNA Benchmark Datasets mRNA related tasks Translation efficiency prediction from Chu et al.(2024) [1] 3 cell lines: Muscle, pc3, HEK input sequence: 5'UTR 10-fold cross-validation split mRNA expression level prediction from Chu et al.(2024) [1] 3 cell lines: Muscle, pc3, HEK input sequence: 5'UTR 10-fold cross-validation split Mean ribosome load prediction from Sample et al. (2019) [2] input sequence: 5'UTR ouput: mean ribosome load the original data… See the full description on the dataset page: https://huggingface.co/datasets/genbio-ai/rna-downstream-tasks.tabular1M<n<10M0 likes1.8k downloads1mo agoHugging Face09SJTU-DENG-Lab /RoboTwin-LeRobot-unseen-tasks-cross-embtabular100K<n<1M0 likes1.8k downloads4mo agoHugging Face10plantcad /PlantCAD2_zero_shot_tasks 🌱 PlantCAD2 Zero-Shot Tasks Zero-shot evaluation tasks for plant genomics using PlantCAD2.This dataset contains tasks designed to evaluate model performance without task-specific training. 📂 Available Tasks 🔬 Cross-species Evolutionary Conservation Task Name Description Samples Metric conservation_within_andropogoneae Predict conserved vs non-conserved sites using alignments within 35 Andropogoneae genomes 19,030 vs 19,030 AUROC… See the full description on the dataset page: https://huggingface.co/datasets/plantcad/PlantCAD2_zero_shot_tasks.tabulartext-classification1M<n<10M0 likes1.6k downloads1y agoHugging Face11efederici /agent-harbor-tasks 🐚 Agent Harbor Tasks 400 original shell-agent tasks, Italian and English, from a 4-call chore to a 20-call investigation. Each task drops an agent into a container with a small, realistic file system (a home folder, a repository, a server's logs and configs, an office archive) and a request written the way a person would type it. The agent works in the shell, and hidden tests decide the reward deterministically: no LLM in the reward path. For every task the reference… See the full description on the dataset page: https://huggingface.co/datasets/efederici/agent-harbor-tasks.tabularothern<1K0 likes1.3k downloads5d agoHugging Face12Fzz1 /Tmax-Tasks-Clean Tmax-Tasks-Clean New: longlongcheck (2026-09-14) Use configuration longlongcheck, split longlongcheck, for the 431-task snapshot combining the selected Codex and Claude Code repairs with passing historical GPT-6 terminal solutions. The existing splits below retain their earlier data. From the latest local 452-task repaired snapshot, this split holds out the requested 15 old-pass/current-fail tasks, four additional tasks without any passing current GPT-6 replay… See the full description on the dataset page: https://huggingface.co/datasets/Fzz1/Tmax-Tasks-Clean.tabular1K<n<10K0 likes1.3k downloads26d agoHugging Face13yatin-superintelligence /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M29 likes1.3k downloads7mo agoHugging Face14open-athena /pdbthink-coordinate-tasks PDBThink Coordinate Tasks 100,000 new coordinate-interpretation tasks across all 19 active PDBThink families, from 2,671 experimental PDB entries in 1,867 source groups. Version 1.3.0; deterministic seed 2026100101. The model receives sanitised, rotated, rounded protein coordinates and a question. It must answer without tools. This release contains no sequence-to-structure prediction tasks and no retired MECH tasks. It is intended for additional evaluation, RL with deterministic… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/pdbthink-coordinate-tasks.tabulartext-generation100K<n<1M0 likes1.2k downloads8d agoHugging Face15LightwheelAI /Lightwheel-Tasks-G1-ControllerThis dataset was created using LeRobot. 1 Dataset Description This dataset includes 62 Lightwheel-Robocasa-Tasks, collected using the g1-controller robot in environments provided by LW-BenchHub.The robot configuration used during data collection is G1-Controller, using UnitreeG1ControllerEnvCfg inherits from UnitreeG1EnvCfg. 1.1 Robot State The robot state recorded in the environment is stored in: observation.state Dimension: 31 Description: Joint positions of… See the full description on the dataset page: https://huggingface.co/datasets/LightwheelAI/Lightwheel-Tasks-G1-Controller.tabularrobotics1M<n<10M0 likes1.2k downloads9mo agoHugging Face16tasksource /chaos-mnli-ambiguity chaos-mnli-ambiguity ChaosNLI, MNLI portion: 1,599 MNLI pairs relabeled by 100 annotators each (Nie et al., 2020). label_dist and label_count follow the entailment/neutral/contradiction order, and gini is the Gini coefficient of label_dist (0 = annotators evenly split, 1 = unanimous). Built from the jsonl first uploaded here, which flattens the ChaosNLI release (https://github.com/easonnie/ChaosNLI) and adds gini; the variable-key label_counter (a duplicate of label_count) is… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/chaos-mnli-ambiguity.tabular1K<n<10K0 likes1.1k downloads16d agoHugging Face17LightwheelAI /Lightwheel-Tasks-X7SThis dataset was created using LeRobot. 1 Dataset Description This dataset includes 117 Lightwheel-Libero-Tasks and 89 Lightwheel-Robocasa-Tasks, collected using the x7s robot in environments provided by LW-BenchHub.The robot configuration used during data collection is X7s-Abs, using X7SAbsEnvCfg inherits from X7SEnvCfg. 1.1 Robot State The robot state recorded in the environment is stored in: observation.state Dimension: 25 Description: Joint positions of… See the full description on the dataset page: https://huggingface.co/datasets/LightwheelAI/Lightwheel-Tasks-X7S.tabularrobotics10M<n<100M1 likes869 downloads9mo agoHugging Face18LightwheelAI /Lightwheel-Tasks-G1-WBCThis dataset was created using LeRobot. 1 Dataset Description This dataset includes 32 Lightwheel-Robocasa-Tasks, collected using the g1-wbc robot in environments provided by LW-BenchHub.The robot configuration used during data collection is G1-Controller-DecoupledWBC, using UnitreeG1ControllerDecoupledWBCEnvCfg inherits from UnitreeG1ControllerEnvCfg. 1.1 Robot State The robot state recorded in the environment is stored in: observation.state Dimension: 43… See the full description on the dataset page: https://huggingface.co/datasets/LightwheelAI/Lightwheel-Tasks-G1-WBC.tabularrobotics1M<n<10M2 likes819 downloads9mo agoHugging Face19ulamai /Math-RL-Tasks Ulam AI Math RL Tasks Forty original, verifier-backed mathematical reasoning tasks packaged as ten independent RL environments. The collection spans advanced graduate exercises, research-style exact computation and structural generalization problems in algebraic geometry, arithmetic geometry, combinatorics, topology, probability and spectral analysis. Each suite pairs a runnable rl_env/ with a preserved blind_run/ by GPT-5.6 Sol Pro. The model name describes the evaluation actor… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/Math-RL-Tasks.tabularquestion-answering1K<n<10K1 likes813 downloads1mo agoHugging Face20LightwheelAI /Lightwheel-Tasks-Double-PiperThis dataset was created using LeRobot. 1 Dataset Description This dataset includes 130 Lightwheel-Libero-Tasks, collected using the double-piper robot in environments provided by LW-BenchHub.The robot configuration used during data collection is DoublePiper-Abs, using DoublePiperAbsEnvCfg inherits from DoublePiperEnvCfg. 1.1 Robot State The robot state recorded in the environment is stored in: observation.state Dimension: 16 Description: Joint positions of… See the full description on the dataset page: https://huggingface.co/datasets/LightwheelAI/Lightwheel-Tasks-Double-Piper.tabularrobotics1M<n<10M1 likes787 downloads9mo agoHugging Face21tasksource /jigsaw_toxicitytabular100K<n<1M2 likes778 downloads3y agoHugging Face22wanglab /bioR_taskstabular1M<n<10M1 likes674 downloads1y agoHugging Face23yatin-superintelligence /Audio-Video-Engineering-Agentic-Tasks-1M Audio/Video Engineering Agentic Tasks (1M) Abstract A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.tabulartext-generation1M<n<10M14 likes599 downloads7mo agoHugging Face24AdaptLLM /medicine-tasks Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024) This repo contains the evaluation datasets for our paper Adapting Large Language Models via Reading Comprehension. We explore continued pre-training on domain-specific corpora for large language models. While this approach enriches LLMs with domain knowledge, it significantly hurts their prompting ability for question answering. Inspired by human learning via reading comprehension, we propose a simple method to… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/medicine-tasks.tabulartext-classification1K<n<10K33 likes575 downloads2y agoHugging Face25rAVEUK /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M3 likes554 downloads7mo agoHugging Face26tasksource /planbench Dataset Card for "planbench" https://arxiv.org/abs/2206.10498 @article{valmeekam2024planbench, title={Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change}, author={Valmeekam, Karthik and Marquez, Matthew and Olmo, Alberto and Sreedharan, Sarath and Kambhampati, Subbarao}, journal={Advances in Neural Information Processing Systems}, volume={36}, year={2024} } tabular10K<n<100K12 likes505 downloads2y agoHugging Face27tasksource /QuALITY Dataset Card for "QuALITY" @article{bowman2022quality, title={QuALITY: Question Answering with Long Input Texts, Yes!}, author={Bowman, Samuel R and Chen, Angelica and He, He and Joshi, Nitish and Ma, Johnny and Nangia, Nikita and Padmakumar, Vishakh and Pang, Richard Yuanzhe and Parrish, Alicia and Phang, Jason and others}, journal={NAACL 2022}, year={2022} } tabular1K<n<10K1 likes477 downloads2y agoHugging Face28kryp1234 /Creative-Professionals-Agentic-Tasks-1M Creative Professionals Agentic Tasks (1M) Abstract A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.tabulartext-generation1M<n<10M1 likes474 downloads7mo agoHugging Face29plantcad /PlantCAD2_fine_tuning_taskstabular10M<n<100M0 likes454 downloads3mo agoHugging Face30Yangximiao /PlantCAD2_zero_shot_tasks 🌱 PlantCAD2 Zero-Shot Tasks Zero-shot evaluation tasks for plant genomics using PlantCAD2.This dataset contains tasks designed to evaluate model performance without task-specific training. 📂 Available Tasks 🔬 Cross-species Evolutionary Conservation Task Name Description Samples Metric conservation_within_andropogoneae Predict conserved vs non-conserved sites using alignments within 35 Andropogoneae genomes 19,030 vs 19,030 AUROC… See the full description on the dataset page: https://huggingface.co/datasets/Yangximiao/PlantCAD2_zero_shot_tasks.tabulartext-classification1M<n<10M0 likes442 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.