Team Ai
Datasetpublic

Neurogica/forecast-workflow-bench

Forecast Workflow Bench (FWBench) 1,251 electricity and cycle-hire cases for evaluating LLM and SLM decisions with budgeted forecasting tools. Code · Leaderboard Evaluation protocol The Language Model evaluation covers 1,251 cases, two hosted and eight local configurations, and 22,518 conversations including paired local TSFM-removal runs. Agents minimize S = 0.5 * ((loss - F) / sigma + credits / B), where F is the minimum feasible loss with known target demand… See the full description on the dataset page: https://huggingface.co/datasets/Neurogica/forecast-workflow-bench.

sourceHugging Faceotherupdated 18d agoView on Hugging Face
0likes411downloads
Dataset Card

Forecast Workflow Bench (FWBench)

1,251 electricity and cycle-hire cases for evaluating LLM and SLM decisions with budgeted forecasting tools.

Code · Leaderboard

Evaluation protocol

The Language Model evaluation covers 1,251 cases, two hosted and eight local configurations, and 22,518 conversations including paired local TSFM-removal runs.

Agents minimize S = 0.5 * ((loss - F) / sigma + credits / B), where F is the minimum feasible loss with known target demand, sigma is the calibrated domain scale, and B = 113726. Prompts omit the action-independent F. Invalid plans retain penalties. Cohort weights are 1/6 for each of three electricity cohorts and 1/2 for cycle hire; cases are averaged within series and lead before cohort weighting. Credits measure forecast-tool expenditure, excluding language-model inference.

The paper and leaderboard also report Sv, a diagnostic weighted mean over valid submissions only, with the original benchmark weights renormalized over that subset. Valid subsets differ across models, so Sv does not determine the ranking. Qwen3.6 35B-A3B without thinking has only three valid cases; its S_v must not be interpreted as representative full-set performance. All underlying case results are unchanged.

The dataset contains the exact full-tool instructions, histories, contracts, tariffs, separate reproduction targets, per-case results, and integrity manifest. These public targets support reproduction, not hidden-test evaluation. Give each agent only its isolated episode; never expose sibling episodes or targets. The full-tool main ranking has ten configurations. The eight local configurations also have no-TSFM controls in the results summary. Forecast-tool removal changes available actions and is not a pure test of forecast accuracy.

Data terms remain component-specific: EIA provenance and TfL Transport Data Service terms apply to their respective observations. Apache-2.0 covers contributed code and generated materials, not a relicensing of the source observations.

Download and verify

bash
hf download Neurogica/forecast-workflow-bench --repo-type dataset --local-dir fwbench-data
python3 fwbench-data/examples/verify.py fwbench-data

Record the Hugging Face commit hash to reproduce this snapshot. episodes/ contains agent-visible inputs; reproduction_targets/ is scorer-only. results/ contains per-case measurements and aggregate evaluations.

License and attribution

See LICENSE for component-specific terms. The code license does not relicense source observations. Powered by TfL Open Data. Contains OS data © Crown copyright and database rights 2016. Geomni UK Map data © and database rights [2019]. Electricity observations: U.S. Energy Information Administration.