Neurogica/forecast-workflow-bench
Forecast Workflow Bench (FWBench) 1,251 electricity and cycle-hire cases for evaluating LLM and SLM decisions with budgeted forecasting tools. Code · Leaderboard Evaluation protocol The Language Model evaluation covers 1,251 cases, two hosted and eight local configurations, and 22,518 conversations including paired local TSFM-removal runs. Agents minimize S = 0.5 * ((loss - F) / sigma + credits / B), where F is the minimum feasible loss with known target demand… See the full description on the dataset page: https://huggingface.co/datasets/Neurogica/forecast-workflow-bench.
Forecast Workflow Bench (FWBench)
1,251 electricity and cycle-hire cases for evaluating LLM and SLM decisions with budgeted forecasting tools.
Evaluation protocol
The Language Model evaluation covers 1,251 cases, two hosted and eight local configurations, and 22,518 conversations including paired local TSFM-removal runs.
Agents minimize S = 0.5 * ((loss - F) / sigma + credits / B), where F is the minimum feasible loss with known target demand, sigma is the calibrated domain scale, and B = 113726. Prompts omit the action-independent F. Invalid plans retain penalties. Cohort weights are 1/6 for each of three electricity cohorts and 1/2 for cycle hire; cases are averaged within series and lead before cohort weighting. Credits measure forecast-tool expenditure, excluding language-model inference.
The paper and leaderboard also report Sv, a diagnostic weighted mean over valid submissions only, with the original benchmark weights renormalized over that subset. Valid subsets differ across models, so Sv does not determine the ranking. Qwen3.6 35B-A3B without thinking has only three valid cases; its S_v must not be interpreted as representative full-set performance. All underlying case results are unchanged.
The dataset contains the exact full-tool instructions, histories, contracts, tariffs, separate reproduction targets, per-case results, and integrity manifest. These public targets support reproduction, not hidden-test evaluation. Give each agent only its isolated episode; never expose sibling episodes or targets. The full-tool main ranking has ten configurations. The eight local configurations also have no-TSFM controls in the results summary. Forecast-tool removal changes available actions and is not a pure test of forecast accuracy.
Data terms remain component-specific: EIA provenance and TfL Transport Data Service terms apply to their respective observations. Apache-2.0 covers contributed code and generated materials, not a relicensing of the source observations.
Download and verify
hf download Neurogica/forecast-workflow-bench --repo-type dataset --local-dir fwbench-data
python3 fwbench-data/examples/verify.py fwbench-dataRecord the Hugging Face commit hash to reproduce this snapshot. episodes/ contains agent-visible inputs; reproduction_targets/ is scorer-only. results/ contains per-case measurements and aggregate evaluations.
License and attribution
See LICENSE for component-specific terms. The code license does not relicense source observations. Powered by TfL Open Data. Contains OS data © Crown copyright and database rights 2016. Geomni UK Map data © and database rights [2019]. Electricity observations: U.S. Energy Information Administration.
