ServiceNow-AI/cascade_bench
CascadeBench: Do Enterprise Systems Need Learned World Models? ๐ Accepted to NeurIPS 2026: The Fortieth Annual Conference on Neural Information Processing Systems (Main Track) A reasoning-focused benchmark for predicting enterprise business-rule cascades, built on synthetic schemas with rule-level attribution of every field change About In enterprise systems, the dynamics come from tenant-specific business logic that varies across deployments and changes over time.โฆ See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/cascade_bench.
<div align="center">
<h1>CascadeBench: Do Enterprise Systems Need Learned World Models?</h1>
<p><a href="https://arxiv.org/abs/2605.12178"><img src="https://img.shields.io/badge/Paper-red?logo=arxiv&logoColor=white" alt="Paper" /></a> <a href="https://neurips.cc/Conferences/2026"><img src="https://img.shields.io/badge/NeurIPS-2026%20Main%20Track-purple" alt="NeurIPS 2026" /></a> <a href="https://github.com/ServiceNow/SyGra/tree/scratch/ewm/tasks/examples/wowstatepredictor_da"><img src="https://img.shields.io/badge/Discovery%20Agent-GitHub-black?logo=github" alt="Discovery Agent" /></a></p>
<p>๐ <b>Accepted to NeurIPS 2026: The Fortieth Annual Conference on Neural Information Processing Systems (Main Track)</b></p>
<p><i>A reasoning-focused benchmark for predicting enterprise business-rule cascades, built on synthetic schemas with rule-level attribution of every field change</i></p>
</div>
<div align="center"><img src="assets/teaser.png" alt="CascadeBench overview" width="90%" /></div>
About
In enterprise systems, the dynamics come from tenant-specific business logic that varies across deployments and changes over time. Business rules, workflows and schema defaults decide what happens when a record changes. The same action can have different effects on different instances.
CascadeBench tests whether a model can predict those effects. Each example gives the current state $st$ and an action $at$. The model predicts the next state $s_{t+1}$: every field-level change across every table, including all changes made by the chain of business rules the action triggers.
The benchmark accompanies the paper *Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer Dynamics*. The paper finds that:
- Offline-trained world models perform well in-distribution but degrade as configurations change. Enterprise discovery agents stay more robust because they read the active rules at runtime.
- Having the rules is necessary but not sufficient. Accuracy still drops sharply as cascades compose, even when the active rules are in the prompt. Multi-step rule composition limits performance more than retrieval does.
Key Features
- ๐งช Real dynamics. A live ServiceNow rule engine produced every transition. Nothing is simulated.
- ๐ท๏ธ Rule-level attribution. A custom execution log traces each field change in the audit log to the business rule that caused it.
- ๐ Synthetic surface form. Table and field names are freshly generated with a
u_prefix, and reserved product namespaces are excluded. Models can't rely on recalling table structures they may have seen in pretraining. - ๐ฆ Full context per example. Each example includes table schemas, business rules, seed records and supporting records. You control how much of this the model sees.
- ๐งน Clean ground truth. Audits are restricted to content fields. System IDs, timestamps and bookkeeping fields are removed.
- โ Validated cascades. Every business rule passes 14 deterministic checks (schema correctness, cycle detection, filter validity, script safety) and is verified through execution.
Code
- Discovery agent: ServiceNow/SyGra โบ `wow_state_predictor_da`
- Dataset construction and evaluation code: coming soon! ๐ง
Dataset Summary
Complexity Tiers
Ground-truth changes are stratified into three tiers:
Field Descriptions
Table and field names in the four JSON-string columns (parameters,seed_data,supporting_data,schema) differ per domain. They are stored as serialized JSON to keep the Arrow schema stable. Calljson.loads()to use them.
Usage
from datasets import load_dataset
import json
ds = load_dataset("ServiceNow-AI/cascade_bench", split="train")
row = ds[0]
print(row["domain"], row["tool_name"])
schema = json.loads(row["schema"]) # table -> column schema
params = json.loads(row["parameters"]) # the action a_t
seed = json.loads(row["seed_data"]) # current state s_t (target record)
rules = row["business_rules"] # rules that may fire
gold = row["audits"] # ground-truth field changes s_{t+1}Evaluation Settings
The paper evaluates three settings, which differ in how much context the model gets:
Example Use Cases
- Benchmark world models and transition predictors on enterprise systems whose dynamics are specific to each deployment.
- Compare internalized and runtime-discovered dynamics by varying which context fields the model sees.
- Study multi-step rule composition by using rule-level attribution to measure how accuracy drops as rule hops increase.
- Evaluate discovery agents that query schemas, workflow definitions and business rules before acting.
Citation
@misc{nair2026enterprisesystemsneedlearned,
title={Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer Dynamics},
author={Jishnu Sethumadhavan Nair and Patrice Bechard and Rishabh Maheshwary and Surajit Dasgupta and Sravan Ramachandran and Aakash Bhagat and Shruthan Radhakrishna and Pulkit Pattnaik and Johan Obando-Ceron and Shiva Krishna Reddy Malay and Sagar Davasam and Seganrasan Subramanian and Vipul Mittal and Sridhar Krishna Nemala and Christopher Pal and Srinivas Sunkara and Sai Rajeswar},
year={2026},
eprint={2605.12178},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2605.12178},
}