datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
3ambench
3amBench: can your agent write alerts that page the right human at 3 a.m., and only then?
3amBench (package alertforge) is an RL environment and benchmark for a job SRE teams do every week:
owning Prometheus alerting rules and Alertmanager routing as code. Each task drops the agent into a realistic
monitoring/ repo for a fictional company with a handful of change requests: onboard a service onto
multi-window burn-rate SLO alerts, fix the alert that paged 30 times last night… See the full description on the dataset page: https://huggingface.co/datasets/openenvforge/3ambench.openenv-pr-review-benchmark
OpenEnv Code Review Environment
This project is now an OpenEnv-style reinforcement learning environment for code review.
An external AI agent receives a static PR task (diff + context), submits a review as an action, and gets a reward based on planted ground-truth issues.
What Changed
Added GitHub Actions integration endpoint for live PR grading.
Removed Nova Act and Bedrock from the active model path.
Added OpenAI-backed analyzer utilities.
Added deterministic OpenEnv… See the full description on the dataset page: https://huggingface.co/datasets/adityam/openenv-pr-review-benchmark.openenv-python-repair
Python Repair Lab
An original OpenEnv curriculum of 1,200 deterministic Python function-repair episodes: 12 problem families, four distinct bug patterns per family, and 25 seeded case sets per pattern. There are 151,780 executable checks across the episodes. These are 48 repair patterns with data variants, not 1,200 unrelated algorithms. Tasks cover interval algorithms, rolling calculations, weighted statistics, stable deduplication, Unicode run-length encoding, Luhn checksums… See the full description on the dataset page: https://huggingface.co/datasets/Louistiti/openenv-python-repair.openenv-sql-investigation
OpenEnv SQL Investigation
Ten original procedural SQL investigation families teach joins, duplicate-safe aggregation, missing-data handling, weighted rates, temporal conditions, cohorts, anti-joins, ranking, and streak analysis. All data are synthetic and generated locally; there are no personal data or downloaded task assets.
The container generates a fresh in-memory SQLite database on every unseeded reset. tasks.jsonl contains one reproducible seed-42 instance per family, with… See the full description on the dataset page: https://huggingface.co/datasets/burtenshaw/openenv-sql-investigation.mimo-openenv-software
MiMo curriculum for OpenEnv
Twelve original tasks from XiaomiMiMo/MiMo-V2.6-RL-oss, served through one OpenEnv shell/finish interface. Nine are code tasks and three are terminal tasks. Problem statements, test patches, embedded test files and scoring rules are preserved. No LLM judge or external service credentials are required.
Source repository · Public dataset
Curriculum
tasks.jsonl records task identities, immutable upstream images, source positions and… See the full description on the dataset page: https://huggingface.co/datasets/akseljoonas/mimo-openenv-software.OpenEnv-Arena-MultiDomain-RL-Mix-v1
OpenEnv Arena Multi-Domain RL Mix v1
This release combines a bounded, source-attributed sample from public NVIDIA RL datasets and XiaomiMiMo's MiMo-V2.6-RL-oss with the locally authored Arena curriculum. Upstream records are stored as JSON in source_record_json; task prompts, verifier family, source revision, row locator, and source license are preserved on every row.
Dataset splits
train: Arena-authored training tasks plus upstream sample records assigned to… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/OpenEnv-Arena-MultiDomain-RL-Mix-v1.openenv-scalingopenenv-job-assets
Chief of Staff Arena
OpenEnv environment for long-horizon executive coordination under pressure
Overview
Chief of Staff Arena evaluates whether an agent can make decisions that stay robust over time, not just optimize one step at a time.
It simulates real executive pressure:
contradictory requests,
hidden preference drift,
stakeholder trust tradeoffs,
deadline cascades,
family vs work conflict.
The core goal is to optimize outcomes + relationships + safety… See the full description on the dataset page: https://huggingface.co/datasets/ssanidhya0407/openenv-job-assets.openenv-arena-compute-lab
Verified Compute Lab
An OpenEnv environment for
OpenEnv Arena: procedural,
multi-step tasks in math and formal reasoning, natural science and
finance that reward careful work with a tool rather than mental arithmetic.
Each episode is a fresh instance: a question, named answer fields and a few
small files (data, rules, sometimes a corrections memo). The agent works with
three tools and gets 12 calls:
{"tool": "read", "path": "rules.md"}
{"tool": "python", "code":… See the full description on the dataset page: https://huggingface.co/datasets/rycerzes/openenv-arena-compute-lab.openenv-office-workflow
Office Workflow Lab v1
Built by Leon AI (https://github.com/leon-ai/leon) with Louistiti.
Source: https://github.com/louistiti/openenv-office-workflow
Verification: https://github.com/louistiti/openenv-office-workflow/releases/tag/office-v1
36 synthetic seed previews: 12 task identities across six office/data workflow families, each shown at three reproducible seeds. The runtime image generates fresh input tables on reset. These are inspectable previews, not 36 independent… See the full description on the dataset page: https://huggingface.co/datasets/Louistiti/openenv-office-workflow.WorldArena
