datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gaia2
Gaia2
Paper | Code | Project Page
Dataset Summary
Gaia2 is a benchmark dataset for evaluating AI agent capabilities in simulated environments. The dataset contains 800 scenarios that test agent performance in environments where time flows continuously and events occur dynamically.
The dataset evaluates seven core capabilities: Execution (multi-step planning and state changes), Search (information gathering and synthesis), Adaptability (dynamic response to environmental… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2.MIT_environmental_impulse_responsesMIT Environmental Impulse Response Dataset
The audio recordings in this dataset are originally created by the Computational Audition Lab at MIT. The source of the data can be found at: https://mcdermottlab.mit.edu/Reverb/IR_Survey.html.
The audio files in the dataset have been resampled to a sampling rate of 16 kHz. This resampling was done to reduce the size of the dataset while making it more suitable for various tasks, including data augmentation.
The dataset consists of 271 audio files… See the full description on the dataset page: https://huggingface.co/datasets/davidscripka/MIT_environmental_impulse_responses.cad-environments
CAD Environments
CAD Environments is a multimodal dataset of complete, human-performed workflows in desktop CAD software. The current release contains 51 task workflows totaling 99.03 hours, covering eight software groups across mechanical design, architecture, MEP, structural design, and general 3D modeling.
Each workflow preserves the full task context—not just the final model—including the problem statement, reference and input files, a gold output, evaluation rubrics, a… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/cad-environments.HydroGym-environmentsgaia2_filesystem
GAIA2 Filesystem
This is a dataset containing files for the GAIA2 benchmark. You should not use this dataset on its own, but instead use the Meta Agents Research Environments framework to execute scenarios from that GAIA2 dataset.
Dataset Link
https://huggingface.co/datasets/meta-agents-research-environments/gaia2
Contact Details
Publishing POC: Meta AI Research Team
Affiliation: Meta Platforms, Inc.
Website:… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2_filesystem.environmental_datasetdata_agent_rl_environment_train_multireward
AdithyaSK/data_agent_rl_environment_train_multireward
Multi-reward variant of AdithyaSK/data_agent_rl_environment_train (2238 tasks, identical
data/instructions). The only change: each task's verifier now emits a reward.json with
three named rewards instead of a single float:
reward
meaning
range
correctness
graded answer matches gold (exact / numeric / LLM-judge)
0 or 1
submission
a non-empty answer was written to /workdir/answer.txt
0 or 1
tool_efficiency
fewer… See the full description on the dataset page: https://huggingface.co/datasets/AdithyaSK/data_agent_rl_environment_train_multireward.data_agent_rl_environment_train_difficulty_ranked
AdithyaSK/data_agent_rl_environment_train_difficulty_ranked
A Harbor task suite of 2238 data-agent tasks, ordered easy -> hard by empirical
difficulty measured from a pass@4 rollout sweep (Qwen3.5-4B + 2B, bash harness).
Layout (standard Harbor spec)
tasks/<task_id>/{task.toml, instruction.md, environment/, tests/}
registry.json # tasks[] IN DIFFICULTY ORDER (rank 1 = easiest); each entry has rank/difficulty/solve_frac
manifest.json # full ranked table… See the full description on the dataset page: https://huggingface.co/datasets/AdithyaSK/data_agent_rl_environment_train_difficulty_ranked.data_agent_rl_environment_train
data_agent_rl_environment_train
The official verified training suite for the data-agent RL pipeline.
2238 Harbor-format data-analysis tasks, each with:
An LLM-assigned difficulty label (L1-L5)
A Kaggle dataset dependency (fetched at container start)
A tested reward function
This is the training-data counterpart to
AdithyaSK/data_agent_rl_environment_eval.
For your held-out eval split, use that one.
💡 Browse in your browser — click the badge above or open… See the full description on the dataset page: https://huggingface.co/datasets/AdithyaSK/data_agent_rl_environment_train.SPADE-Environments-Qwen3-30B-Games
SPADE generated environments: games
Paper | Code | All artifacts
Executable game environments written by the SPADE Environment Designer during the paper's 30B games self-play run. One Python file per environment; manifest.json records the generation checkpoint, training step, skill, and difficulty of each.
Environments
3310
Training steps covered
113 (step 0 to 396)
With skill label
3119
Designer / agent model
Qwen/Qwen3-30B-A3B-Instruct-2507… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Environments-Qwen3-30B-Games.MIT_environmental_impulse_responses
MIT Environmental Impulse Response Dataset
The audio recordings in this dataset are originally created by the Computational Audition Lab at MIT. The source of the data can be found at: https://mcdermottlab.mit.edu/Reverb/IR_Survey.html.
This mirror provides the 16 kHz WAV files used for wake-word training augmentation in the Tater Totterson trainer projects. The files were resampled to 16 kHz to keep the dataset small and convenient for machine-learning audio pipelines.… See the full description on the dataset page: https://huggingface.co/datasets/TaterTotterson/MIT_environmental_impulse_responses.corral-environment-tasks
Corral – Environment Tasks
Task definitions across the 8 Corral environments, including descriptions, allowed tools, scoring functions, and submission formats
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the task definitions for the 8 environments included in the Corral benchmark.
The dataset is organized into multiple configurations… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-environment-tasks.proc-gen-environments
geodesic-research/proc-gen-environments
Local-pipeline snapshot published via --push-from-local (GH #52). All configs below were built locally (Hub-independent) and uploaded in a single commit at one snapshot revision.
Pipeline run params hash: dd13a8843fce63fefa4e70c743f85c6cba30904523b6097a96bfe98100c298eb
Configs in this snapshot: case-elaboration-conversation, cases-conversation, cases-fields-conversation, cases-intermediate-conversation, contracts-conversation
Per-run… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/proc-gen-environments.africa-world-bank-climate-environment-time-series
Africa World Bank Climate and Environment Labeled Time Series Data
This repository is part of the Africa Temporal Intelligence Corpus (ATIC). It contains sector-specific temporal corpus packages for African countries.
ATIC sector repositories are designed for machine consumption first: Parquet tables, stable IDs, reproducible metadata, explicit provenance, review status, and separable semantic layers.
Sector Scope
Temporal climate and environmental indicators for… See the full description on the dataset page: https://huggingface.co/datasets/africatic/africa-world-bank-climate-environment-time-series.data_agent_rl_environment_train_subset_100
data_agent_rl_environment_train_subset_100
A 100-task quick-iteration subset of the data-agent RL training suite.
All tasks are L1 difficulty (the easiest tier) with a numeric reward function —
chosen so RL/eval loops converge fast and grade deterministically (no LLM-judge variance).
This is a strict subset of
AdithyaSK/data_agent_rl_environment_train; for the full
2238-task suite or the held-out eval split, use that one and
AdithyaSK/data_agent_rl_environment_eval.
💡 Browse… See the full description on the dataset page: https://huggingface.co/datasets/AdithyaSK/data_agent_rl_environment_train_subset_100.Environment-and-Natural-Resources-Indicators-For-African-Countries
Environment and Natural Resources Indicators For African Countries | Africa (World Health Organization)
Size category: 1K<n<10K - Formats: csv - Sector: climate_environment - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Environment-and-Natural-Resources-Indicators-For-African-Countries.SPADE-Environments-ToolUse
SPADE generated environments: tool use
Paper | Code | All artifacts
Multi-turn tool-use environments written by the SPADE designer during training, pooled
across every captured run. 2,231 environments across 7 runs and two model scales (30B-A3B and 4B).
Source run
Scale
Environments
qwen3-30b-0617-tooluse-regen32-mixed
30B-A3B
41
qwen3-30b-0624-tooluse-blend
30B-A3B
243
qwen3-30b-0703-tooluse-glory-kl005
30B-A3B
260
qwen3-4b-0630-tooluse-eval-aligned-r32
4B
456… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Environments-ToolUse.SPADE-Environment-Pool-GPT5.5-Games
SPARE GPT-5.5 Grounded Cognitive Multi-Turn Games
This public dataset contains 7,872 validated Python game environments for actor-only SPARE training.
Six cognitive skills, exactly 1,312 environments per skill
Generated with GPT-5.5 and grounded by spice_megascience_15k.jsonl
Grounding corpus SHA-256: a36a928b4940b5b5d9e3f4cb5804a94c69462360943adb3be14613c82f0f72c0
Maximum 25 turns and 32K generation context
Every environment passes load, reset, step, and replay validation with… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Environment-Pool-GPT5.5-Games.gaia2-cli
GAIA2 CLI
Benchmark dataset for gaia2-cli, the CLI-based agent evaluation harness.
Schema
Each row has two columns:
Column
Type
Description
scenario_id
string
Unique scenario identifier (e.g. scenario_universe_21_1qgjj6)
scenario
string
Complete scenario as a JSON string
Usage
from datasets import load_dataset
import json
# Load a specific config (160 scenarios)
ds = load_dataset("meta-agents-research-environments/gaia2-cli", "adaptability"… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2-cli.Environmental_geology-dem
Digital Elevation Model
Overview
This repository contains the Digital Elevation Model dataset, part of the Environmental Geology suite maintained by NORA Research Lab. The data is processed into Cloud-Optimized GeoTIFF (COG) format for efficient geospatial analysis.
Dataset Details
Variable: Digital Elevation Model
Units: metres
Format: Cloud-Optimized GeoTIFF (COG)
Compression: ZSTD
Coordinate System: EPSG:4326 (WGS84)
Spatial Resolution: ~30… See the full description on the dataset page: https://huggingface.co/datasets/NoraResearchLab/Environmental_geology-dem.data_agent_rl_environment_eval
data_agent_rl_environment_eval
The official verified eval suite for the data-agent RL pipeline. 366 Harbor-format
data-analysis tasks, each with an LLM-assigned difficulty label (L1–L5), a Kaggle
dataset dependency, and a tested reward function.
💡 Browse this dataset in your browser — click the badge above or open
AdithyaSK/harbor-visualiser
to inspect every task's spec, instruction, environment, tests, and difficulty.
Reproduce the eval — end to end
The… See the full description on the dataset page: https://huggingface.co/datasets/AdithyaSK/data_agent_rl_environment_eval.EnvShip-Bench_An_Environment-Enhanced_Benchmark_for_Short-Term_Vessel_Trajectory_Prediction
EnvShip-Bench
EnvShip-Bench is a benchmark for short-term vessel trajectory prediction built from raw AIS data released by the Danish Maritime Authority (DMA).
This Hugging Face release currently provides two source branches under a shared layout:
DMA/
NOAA/
The DMA branch is organized under:
DMA/benchmark/core/
DMA/benchmark/full/
DMA/mini_benchmark/ship_core_lite/
DMA/mini_benchmark/clean_ship_core_lite_v1/
The NOAA branch is organized under:
NOAA/benchmark/core/… See the full description on the dataset page: https://huggingface.co/datasets/mark000071/EnvShip-Bench_An_Environment-Enhanced_Benchmark_for_Short-Term_Vessel_Trajectory_Prediction.Environmental_geology-slope
Slope
Overview
This repository contains the Slope dataset, part of the Environmental Geology suite maintained by NORA Research Lab. The data is processed into Cloud-Optimized GeoTIFF (COG) format for efficient geospatial analysis.
Dataset Details
Variable: Slope
Units: degrees
Format: Cloud-Optimized GeoTIFF (COG)
Compression: ZSTD
Coordinate System: EPSG:4326 (WGS84)
Spatial Resolution: ~30 metres (1 arc-second)
Attribution… See the full description on the dataset page: https://huggingface.co/datasets/NoraResearchLab/Environmental_geology-slope.DROID-sim-environmentsWaterScopeAI-DataEnvironmental_geology-dem
Digital Elevation Model
Overview
This repository contains the Digital Elevation Model dataset, part of the Environmental Geology suite maintained by NORA Research Lab. The data is processed into Cloud-Optimized GeoTIFF (COG) format for efficient geospatial analysis.
Dataset Details
Variable: Digital Elevation Model
Units: metres
Format: Cloud-Optimized GeoTIFF (COG)
Compression: ZSTD
Coordinate System: EPSG:4326 (WGS84)
Spatial Resolution: ~30… See the full description on the dataset page: https://huggingface.co/datasets/o0Biggz0o/Environmental_geology-dem.digital-hospital-environment
Digital Hospital Environment
Digital Hospital is an open-source clinical AI benchmark environment for evaluating agents that must operate inside a structured hospital workflow. It combines role-specific medical knowledge checks, patient-facing clinical operations, cross-role communication, deterministic grading, dense process rewards, and rollout capture in one downloadable runtime. The benchmark is designed for model evaluation, process-supervision datasets, offline… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/digital-hospital-environment.environmental_claims
Dataset Card for environmental_claims
Dataset Summary
We introduce an expert-annotated dataset for detecting real-world environmental claims made by listed companies.
Supported Tasks and Leaderboards
The dataset supports a binary classification task of whether a given sentence is an environmental claim or not.
Languages
The text in the dataset is in English.
Dataset Structure
Data Instances
{
"text": "It will enable E.ON to… See the full description on the dataset page: https://huggingface.co/datasets/climatebert/environmental_claims.environment-contracts
geodesic-research/environment-contracts
Local-pipeline snapshot published via --push-from-local (GH #52). All configs below were built locally (Hub-independent) and uploaded in a single commit at one snapshot revision.
Pipeline run params hash: dd13a8843fce63fefa4e70c743f85c6cba30904523b6097a96bfe98100c298eb
Configs in this snapshot: conversation
Per-run provenance: _pipeline_state/dd13a8843fce63fefa4e70c743f85c6cba30904523b6097a96bfe98100c298eb.json (in this repo) and each… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/environment-contracts.Environmental_geology-roughness
Surface roughness (TRI)
Overview
This repository contains the Surface roughness (TRI) dataset, part of the Environmental Geology suite maintained by NORA Research Lab. The data is processed into Cloud-Optimized GeoTIFF (COG) format for efficient geospatial analysis.
Dataset Details
Variable: Surface roughness (TRI)
Units: metres
Format: Cloud-Optimized GeoTIFF (COG)
Compression: ZSTD
Coordinate System: EPSG:4326 (WGS84)
Spatial Resolution: ~30… See the full description on the dataset page: https://huggingface.co/datasets/NoraResearchLab/Environmental_geology-roughness.
