datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gaia2
Gaia2
Paper | Code | Project Page
Dataset Summary
Gaia2 is a benchmark dataset for evaluating AI agent capabilities in simulated environments. The dataset contains 800 scenarios that test agent performance in environments where time flows continuously and events occur dynamically.
The dataset evaluates seven core capabilities: Execution (multi-step planning and state changes), Search (information gathering and synthesis), Adaptability (dynamic response to environmental… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2.cad-environments
CAD Environments
CAD Environments is a multimodal dataset of complete, human-performed workflows in desktop CAD software. The current release contains 51 task workflows totaling 99.03 hours, covering eight software groups across mechanical design, architecture, MEP, structural design, and general 3D modeling.
Each workflow preserves the full task context—not just the final model—including the problem statement, reference and input files, a gold output, evaluation rubrics, a… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/cad-environments.gaia2_filesystem
GAIA2 Filesystem
This is a dataset containing files for the GAIA2 benchmark. You should not use this dataset on its own, but instead use the Meta Agents Research Environments framework to execute scenarios from that GAIA2 dataset.
Dataset Link
https://huggingface.co/datasets/meta-agents-research-environments/gaia2
Contact Details
Publishing POC: Meta AI Research Team
Affiliation: Meta Platforms, Inc.
Website:… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2_filesystem.environmental_datasetcorral-environment-tasks
Corral – Environment Tasks
Task definitions across the 8 Corral environments, including descriptions, allowed tools, scoring functions, and submission formats
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the task definitions for the 8 environments included in the Corral benchmark.
The dataset is organized into multiple configurations… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-environment-tasks.proc-gen-environments
geodesic-research/proc-gen-environments
Local-pipeline snapshot published via --push-from-local (GH #52). All configs below were built locally (Hub-independent) and uploaded in a single commit at one snapshot revision.
Pipeline run params hash: dd13a8843fce63fefa4e70c743f85c6cba30904523b6097a96bfe98100c298eb
Configs in this snapshot: case-elaboration-conversation, cases-conversation, cases-fields-conversation, cases-intermediate-conversation, contracts-conversation
Per-run… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/proc-gen-environments.africa-world-bank-climate-environment-time-series
Africa World Bank Climate and Environment Labeled Time Series Data
This repository is part of the Africa Temporal Intelligence Corpus (ATIC). It contains sector-specific temporal corpus packages for African countries.
ATIC sector repositories are designed for machine consumption first: Parquet tables, stable IDs, reproducible metadata, explicit provenance, review status, and separable semantic layers.
Sector Scope
Temporal climate and environmental indicators for… See the full description on the dataset page: https://huggingface.co/datasets/africatic/africa-world-bank-climate-environment-time-series.Environment-and-Natural-Resources-Indicators-For-African-Countries
Environment and Natural Resources Indicators For African Countries | Africa (World Health Organization)
Size category: 1K<n<10K - Formats: csv - Sector: climate_environment - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Environment-and-Natural-Resources-Indicators-For-African-Countries.gaia2-cli
GAIA2 CLI
Benchmark dataset for gaia2-cli, the CLI-based agent evaluation harness.
Schema
Each row has two columns:
Column
Type
Description
scenario_id
string
Unique scenario identifier (e.g. scenario_universe_21_1qgjj6)
scenario
string
Complete scenario as a JSON string
Usage
from datasets import load_dataset
import json
# Load a specific config (160 scenarios)
ds = load_dataset("meta-agents-research-environments/gaia2-cli", "adaptability"… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2-cli.digital-hospital-environment
Digital Hospital Environment
Digital Hospital is an open-source clinical AI benchmark environment for evaluating agents that must operate inside a structured hospital workflow. It combines role-specific medical knowledge checks, patient-facing clinical operations, cross-role communication, deterministic grading, dense process rewards, and rollout capture in one downloadable runtime. The benchmark is designed for model evaluation, process-supervision datasets, offline… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/digital-hospital-environment.environmental_claims
Dataset Card for environmental_claims
Dataset Summary
We introduce an expert-annotated dataset for detecting real-world environmental claims made by listed companies.
Supported Tasks and Leaderboards
The dataset supports a binary classification task of whether a given sentence is an environmental claim or not.
Languages
The text in the dataset is in English.
Dataset Structure
Data Instances
{
"text": "It will enable E.ON to… See the full description on the dataset page: https://huggingface.co/datasets/climatebert/environmental_claims.environment-contracts
geodesic-research/environment-contracts
Local-pipeline snapshot published via --push-from-local (GH #52). All configs below were built locally (Hub-independent) and uploaded in a single commit at one snapshot revision.
Pipeline run params hash: dd13a8843fce63fefa4e70c743f85c6cba30904523b6097a96bfe98100c298eb
Configs in this snapshot: conversation
Per-run provenance: _pipeline_state/dd13a8843fce63fefa4e70c743f85c6cba30904523b6097a96bfe98100c298eb.json (in this repo) and each… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/environment-contracts.termgrade-environments
TermGrade Environments: 1,004 executable Linux tasks
Part of TermGrade: graded environments and trajectories for terminal agents. Read the blog post.
1,004 executable environments · graded by 6 models · measured pass rates
Each task is a container, an instruction, and a pytest verifier, and each one ships a
measured pass rate per model.
Companion release: ai-and/termgrade-trajectories, 36k agent episodes with the raw terminal output and per-test results for every trial here.… See the full description on the dataset page: https://huggingface.co/datasets/ai-and/termgrade-environments.SPADE-Environments-Qwen3-30B-ToolUse
qwen3-30B-A3B-Instruct-0703-tooluse-glory-kl005 — generated environments
Environments generated by the SPARE proposer during training run
2hjdrbeh (qwen3-30B-A3B-Instruct-0703-tooluse-glory-kl005), recovered from the spare-viz durable cache.
The run's scratch directory no longer exists; this dataset is the surviving copy.
Games
260
Steps covered
7 (step 0–192)
With recovered skill
260
With hint
0
Actor / proposer model… See the full description on the dataset page: https://huggingface.co/datasets/msr-spare-1/SPADE-Environments-Qwen3-30B-ToolUse.TerminalHorizon-Environment
TerminalHorizon-Environment
Scaling Agentic Data for Long-Horizon Terminal Intelligence
TerminalHorizon is a fully automated data synthesis engine for long-horizon
terminal agents. Grounded in real-world professional work, it constructs
executable environments and generates training trajectories that connect agentic
behavioral patterns across stages toward a shared goal.
This repository releases the 1,535 task environments underlying the project.
Each task includes a public… See the full description on the dataset page: https://huggingface.co/datasets/shunzou05/TerminalHorizon-Environment.vn-provinces-society-environment-master
Vietnam health, living standards, culture, justice and environment master
Wide geo×year master joining NSO society, health, environment, justice and related locality packs (poverty, HDI/Gini, health workforce and facilities, waste, accidents, culture indicators, and others listed in the README). Overlapping column names are prefixed. Province names follow ar_core.vn_geo.
Figures
Hero
Comparison
Color key
Files
provinces (1450 rows)… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-society-environment-master.dream-network-environment-cards
Dream Network — Environment Location Cards (20)
20 fully-annotated environment reference cards — the locations of the Dream Network, each a self-contained 1:1 card: a cinematic establishing view plus alternate views and a full worldbuilding stat panel rendered into the image.
Companion to dream-network-player-cards (the characters) and unhinged-cast-20 (their turnaround sheets). Together they form a complete production bible: who the characters are, what they look like, and… See the full description on the dataset page: https://huggingface.co/datasets/TheMindExpansionNetwork/dream-network-environment-cards.processrl-terminal-environments
ProcessRL Terminal Environments
ProcessRL is a collection of behavior-conditioned terminal environments for training and evaluating agent process control. The tasks are designed around failures that appear in interactive terminal work: stopping after a misleading successful command, repeating an unproductive action, failing to pivot after a dead end, losing track of migrated state, and leaving partial progress unfinished.
This release contains the first public train/heldout… See the full description on the dataset page: https://huggingface.co/datasets/Jarrodbarnes/processrl-terminal-environments.africa-world-bank-environment-indicators-for-cameroon
Cameroon - Environment | Africa (original)
Size category: 1K<n<10K - Formats: parquet - Sector: climate_environment - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts inspect structured… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-world-bank-environment-indicators-for-cameroon.environmental-dialogueenvironmental-forecast-v1
Environmental Forecasting v1
Overview
Environmental Forecasting v1 is a multivariate environmental time-series dataset designed for studying short-term forecasting of indoor environmental conditions and early warning of abnormal future trajectories.
The dataset represents five environmental sensing channels corresponding to measurements commonly associated with BME680 and MQ-135 sensor systems:
Temperature
Relative humidity
Atmospheric pressure
Gas resistance… See the full description on the dataset page: https://huggingface.co/datasets/Ajeya95/environmental-forecast-v1.vibeworlding-environments-10
VibeWorlding 场景准备:10 个场景
独立场景源码与验证记录,方便协作下载。 不是作者 V2 原生训练集,也不是统一通过完整物理验证的数据集。场景未与新增资产库提前绑定。
下载和打开
下载完整环境包(67.5 MB)
解压后在包根目录运行 python3 -m http.server 8766 --bind 127.0.0.1,打开 http://127.0.0.1:8766/batch-001-ten/gallery.html 看场景,或 http://127.0.0.1:8766/v2-physics-pilot/index.html 看物理测试回放。无需部署 AI 模型;HF 数据集页面供下载,HTML 预览需上述本地静态服务。
现有验证状态
10/10 通过静态准备检查;8/10 通用 Three.js JSON 重载与单视角画面对照通过。10 个场景都实际运行了 Rapier 20 秒、240 Hz 刚体仿真,5/10 满足本轮全部保守判据。保存了全部失败和逐物体轨迹。… See the full description on the dataset page: https://huggingface.co/datasets/anon123312/vibeworlding-environments-10.environmental_2kSPADE-Environment-Pool-GPT5.5-ToolUse
SPARE GPT-5.5 Multi-Turn Tool-Use Games v1
A public static pool of 11,039 validated multi-turn tool-use environments generated by GPT-5.5 for SPARE actor training.
Training alignment
Source recipe: Qwen3-30B-A3B 0624 tool-use GAMES configuration
400 rollouts x 24 games/rollout = 9,600 no-reuse games required
11,039 validated games provide 1,439 games of headroom
Six balanced skills: API orchestration, data retrieval, state modification, error recovery, tool… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Environment-Pool-GPT5.5-ToolUse.EnvironmentalSoundClassification_ESC50-NaturalSoundscapesAndWaterSounds
Dataset Card for "environmental_sound_classification_natural_soundscapes_and_water_sounds_ESC50"
More Information needed
EnvironmentalSoundClassification_ESC50-HumanAndNonSpeechSounds
Dataset Card for "environmental_sound_classification_human_and_non_speech_sounds_ESC50"
More Information needed
EnvironmentalSoundClassification_UrbanSound8K-UrbanNoisesEnvironmentalSoundClassification_ESC50-Animals
Dataset Card for "environmental_sound_classification_animals_ESC50"
More Information needed
KOR-RE-natures-and-environments
Dataset Card for [KOR-RE-natures-and-environments]
You can find relation map, guidelines(written in Korean), short technical papers in this github repo. This work is done by as part of project for Boostcamp AI Tech supported by Naver Connect Foundation.
Main Data Fields
Sentences: sentences
Subject_entity: infos for subject entity in the sentence including words, start index, end index, type of entity
object_entity: infos for object entity in the sentence including words… See the full description on the dataset page: https://huggingface.co/datasets/kimcando/KOR-RE-natures-and-environments.environment5
