datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
3ambench
3amBench: can your agent write alerts that page the right human at 3 a.m., and only then?
3amBench (package alertforge) is an RL environment and benchmark for a job SRE teams do every week:
owning Prometheus alerting rules and Alertmanager routing as code. Each task drops the agent into a realistic
monitoring/ repo for a fictional company with a handful of change requests: onboard a service onto
multi-window burn-rate SLO alerts, fix the alert that paged 30 times last night… See the full description on the dataset page: https://huggingface.co/datasets/openenvforge/3ambench.openenv-pr-review-benchmark
OpenEnv Code Review Environment
This project is now an OpenEnv-style reinforcement learning environment for code review.
An external AI agent receives a static PR task (diff + context), submits a review as an action, and gets a reward based on planted ground-truth issues.
What Changed
Added GitHub Actions integration endpoint for live PR grading.
Removed Nova Act and Bedrock from the active model path.
Added OpenAI-backed analyzer utilities.
Added deterministic OpenEnv… See the full description on the dataset page: https://huggingface.co/datasets/adityam/openenv-pr-review-benchmark.openenv-smoke-fix-gitArena smoke test: the OpenEnv Terminal-Bench 2 fix-git task as source (environment/) and its built image as an OCI layout (images/fix-git/).
WorldAtlas
WorldAtlas: A Unified Large-Scale Benchmark forWorld Generation toward World Intelligence
OpenEnvision
Is your world model a consistent, controllable, and physically plausible scene renderer?
5,000 annotated images · 960 memory cases · 4,961 reference videos
WorldAtlas provides evaluation assets for generative world models: 5,000 annotated images, 960 initial images for closed-loop memory evaluation, and 4,961 reference… See the full description on the dataset page: https://huggingface.co/datasets/OpenEnvisionLab/WorldAtlas.openenv-sql-investigation
OpenEnv SQL Investigation
Ten original procedural SQL investigation families teach joins, duplicate-safe aggregation, missing-data handling, weighted rates, temporal conditions, cohorts, anti-joins, ranking, and streak analysis. All data are synthetic and generated locally; there are no personal data or downloaded task assets.
The container generates a fresh in-memory SQLite database on every unseeded reset. tasks.jsonl contains one reproducible seed-42 instance per family, with… See the full description on the dataset page: https://huggingface.co/datasets/burtenshaw/openenv-sql-investigation.openenv-examples-trl-2026-07-06openenv-scalingopenenv-pathway-analysis-env
OpenEnv Pathway Analysis Environment
This repository packages pathway_analysis_env for Hugging Face Hub publication.
Contents
envs/pathway_analysis_env/ environment code
reproducible GEO benchmark inputs for 3 tasks
task expansion scripts:
create_geo_task.py
append_task_to_manifest.py
Run locally
uv sync --all-extras
PYTHONPATH=src:envs uv run python envs/pathway_analysis_env/scripts/run_agent_eval_suite.py --manifest… See the full description on the dataset page: https://huggingface.co/datasets/aparulsarma/openenv-pathway-analysis-env.RealseeWMBench
Dataset Summary
Sunain Gameplay is a large-scale multimodal dataset containing synchronized gameplay videos and player input actions from 12 popular video games. The dataset comprises over 1000 curated clips (10-50 seconds each) with 2M+ frames and frame-level action annotations, designed specifically for training embodied AI agents, world models, and video generation systems. This release represents a curated sample of the full Sunain gameplay dataset. For access to the complete… See the full description on the dataset page: https://huggingface.co/datasets/OpenEnvisionLab/WMBench.openenv-job-assets
Chief of Staff Arena
OpenEnv environment for long-horizon executive coordination under pressure
Overview
Chief of Staff Arena evaluates whether an agent can make decisions that stay robust over time, not just optimize one step at a time.
It simulates real executive pressure:
contradictory requests,
hidden preference drift,
stakeholder trust tradeoffs,
deadline cascades,
family vs work conflict.
The core goal is to optimize outcomes + relationships + safety… See the full description on the dataset page: https://huggingface.co/datasets/ssanidhya0407/openenv-job-assets.sakthai-openenv-training
SakThai OpenEnv Training
Part of the SakThai model family.
Dataset Summary
SakThai OpenEnv Training is a pinned runtime-environment dataset for reproducing SakThai training workflows. It stores exact package versions, experimental interface requirements, and notes for openenv, trl, and GRPO integrations used during model training.
Purpose:
Ensure bit-for-bit reproducibility of training runs on GPU clusters
Document experimental dependencies (openenv==0.4.1… See the full description on the dataset page: https://huggingface.co/datasets/Nanthasit/sakthai-openenv-training.maas-openenv-rl-datasetarena-resultsblockgroupvoting
Problem and Opportunity
In the United States, voting is largely a private matter. A registered voter is given a randomized ballot form or machine to prevent linkage between their voting choices and their identity. This disconnect supports confidence in the election process, but it provides obstacles to an election's analysis. A common solution is to field exit polls, interviewing voters immediately after leaving their polling location. This method is rife with bias, however, and… See the full description on the dataset page: https://huggingface.co/datasets/openenvironments/blockgroupvoting.email-triage-openenv
Email Triage & Response — OpenEnv Environment
A real-world OpenEnv environment where AI agents learn to triage corporate email inboxes: categorize, prioritize, reply, forward, and flag emails.
🌟 Why Email Triage?
Email triage is a task performed by billions of knowledge workers daily. It requires:
Reading comprehension — understanding intent and urgency
Decision-making — choosing correct actions from a discrete set
Context reasoning — considering sender, deadlines… See the full description on the dataset page: https://huggingface.co/datasets/RuDubnium/email-triage-openenv.WorldEngineopenenv-training-assetsWorldArenasmartgrid-openenv-dataset
Smart Grid Energy Trader — Dataset
Complete training and evaluation dataset for the Smart Grid OpenEnv environment.
Files
File
Rows
Description
states.csv
216
All environment states across all tasks
episodes.csv
3,860
Full episode rollouts with agent decisions
price_patterns.csv
96
24-hour price curves for 4 market scenarios
llm_prompts.jsonl
500
LLM training examples in OpenAI chat format
states.csv columns
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/codechimanshu/smartgrid-openenv-dataset.energy_openenv
Smart Energy Management OpenEnv
Problem
Optimize energy usage while maintaining comfort.
How it works
Agent controls appliances to balance:
Energy consumption
Comfort level
Run
python agent.py
Evaluate
python grader.py
WM4Dgrid2op-openenv-datasetsArenaResultsmimo-openenv-software
MiMo software engineering for OpenEnv
An OpenEnv adapter for eight software engineering tasks from XiaomiMiMo/MiMo-V2.6-RL-oss. Each task retains its original problem statement, repository image, test patch, test command and binary reward: 1 if the original tests exit successfully, 0 otherwise. Infrastructure errors are reported as errors. No LLM judge or external service credentials are used.
The tasks cover Python and JavaScript library repairs. tasks.jsonl lists task… See the full description on the dataset page: https://huggingface.co/datasets/akseljoonas/mimo-openenv-software.openenv-python-repair
Python Repair Lab
An original OpenEnv curriculum of 1,200 deterministic Python function-repair episodes: 12 problem families, four distinct bug patterns per family, and 25 seeded case sets per pattern. There are 151,780 executable checks across the episodes. These are 48 repair patterns with data variants, not 1,200 unrelated algorithms. Tasks cover interval algorithms, rolling calculations, weighted statistics, stable deduplication, Unicode run-length encoding, Luhn checksums… See the full description on the dataset page: https://huggingface.co/datasets/Louistiti/openenv-python-repair.
