AgentsSci/EMNLP_Cost-Aware-Protocol-Routing
Cost-Aware Protocol Routing: Matched Protocol Outcomes The short version. We ran the same 6,803 reasoning problems through four different LLM collaboration setups — from a single direct answer up to a four-agent deliberation — and recorded, for every problem, which ones got it right. Then we asked whether a model can look at a problem beforehand and predict which setup is worth paying for. It can predict whether it will fail. It cannot predict which collaboration protocol will… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/EMNLP_Cost-Aware-Protocol-Routing.
Cost-Aware Protocol Routing: Matched Protocol Outcomes
The short version. We ran the same 6,803 reasoning problems through four different LLM collaboration setups — from a single direct answer up to a four-agent deliberation — and recorded, for every problem, which ones got it right. Then we asked whether a model can look at a problem beforehand and predict which setup is worth paying for.
It can predict whether it will fail. It cannot predict which collaboration protocol will fix the failure. That gap is what this dataset is for.
This is the data release for the EMNLP 2026 paper *LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off*.
The four protocols
Every released problem was run under all four, with the same solver model.
The oracle label
For each problem, oracle_label is the first protocol that succeeded, scanned in the fixed cost order Baseline -> Single -> PER -> Broadcast. If all four failed, the label is none.
Two things to be clear about:
- `none` is not a fifth protocol. It is a router action: abstain, spend nothing, because nothing observed here worked. Four protocols were executed on every released problem.
- The oracle is retrospective. It is computed from one realized execution per protocol, not from repeated sampling. It is the best a router could have done on these specific runs, not a ground-truth best action. A protocol that failed once here might succeed on a resample.
What is in here
15,088 rows in the flagship table: 10 settings (5 benchmark conditions x 2 solver models), covering 6,803 distinct problems.
Alongside the outcomes: two documented split schemes, held-out router predictions and metrics for six settings, confidence-probe measurements, and the paper's eight aggregate result tables reproduced verbatim.
Full column-by-column documentation: `docs/schema.md`.
What is NOT in here, and why
No problem text. No gold answers. No answer options. No reference solutions.
The four upstream benchmarks carry four different licenses, and LAB-Bench — the most restrictive — is CC-BY-SA-4.0 with an upstream do-not-train request. Rather than ship a mixed-license text column that downstream users would have to untangle correctly, this release withholds text across the board and gives you stable identifiers plus reconstruction instructions instead.
That applies even to JEEBench and SciBench, whose MIT licenses would have permitted redistribution. The decision was made release-wide because the flagship table interleaves all four benchmarks in one file.
To attach the text yourself: `docs/reconstruction.md`. The licensing evidence: `docs/license_audit.md`.
Also not included: raw model generations, per-problem cost accounting for nine of the ten settings (one setting is covered, see below), MaScQA (excluded — NonCommercial, and not in the paper), and the Gemma-3-27B scope check. See `docs/provenance.md`.
Also not included: Baseline final-answer text. The post-answer probe consumed it, so data/probe_inputs.jsonl is an identifier and metadata manifest, not a runnable prompt set — rebuilding runnable probe inputs needs both the problem text rehydrated from upstream and that setting's Baseline final answer. Baseline answers are not released, and for three of the six probe settings they no longer exist in the project's own artifacts either, so the probe is not fully re-runnable by anyone. The probe's outputs are released in full, so the paper's numbers stay reproducible. Coverage is measured in `TODO.md`.
Cost: per-problem tokens, for one setting only
The paper's cost axis is tokens, and data/costs/omnimath_per_protocol_costs.csv gives per-problem token totals and model-call counts for all four protocols — but only for `omnimath2__competition_math_4181__gpt_oss_120b`, one of the ten settings. The other nine have no per-problem cost data here; their cost appears only through the aggregate tables.
These are protocol-level totals, summed over every model call the protocol made. That is the paper's accounting. They are not a single call's prompt+completion — Baseline averages 9.67 model calls per problem in this setting, so the two quantities differ by roughly that factor. If you compare against another dataset's total_tokens, check which quantity it measures first.
4,155 of 4,181 problems are covered. The 26 omitted problems form 13 duplicate-text pairs where a cost row cannot be attributed to a specific problem_id; they were dropped rather than guessed. As a result, the mean Baseline token count recomputed from this file is 18,432.0, while the paper's published figure over all 4,181 problems is 18,385.4. That gap is the arithmetic of the 26 dropped rows, not a disagreement — cite 18,385.4 for the paper, expect 18,432.0 from this file, and do not reconcile them by adjusting either.
Two confidence probes — do not confuse them
The dataset ships two different instruments. They ask different questions at different points in the pipeline, and they support different numbers in the paper. Every row of both files carries a probe_type column.
The title claim comes from the post-answer probe. Its released predictions reproduce data/aggregate/postanswer_confidence.csv for all six settings to four decimals — including the headline gpt-oss-120b OmniMath figures of 4,181 rows, 4,151 parseable, and 0.8847 failure AUROC.
Reproducibility limit. The post-answer probe cannot be fully re-run from released artifacts, because the Baseline final-answer string that the probe consumes was not persisted for the gpt-oss runs. The probe's outputs and metrics are fully released and independently reproducible. Recoverable coverage is 6,462 / 12,928 probe rows (50.0%): 100% for the three Gemma settings, 0% for the three gpt-oss settings. See `TODO.md`.
Neither probe ever sees the gold answer, the correctness label, the oracle label, or any protocol outcome. The post-answer prompt states that boundary explicitly and is injection-hardened — it marks the problem text and the baseline answer as untrusted data, tells the model not to follow instructions inside them, and enumerates what the model does not have. It ships verbatim at `docs/confidence_probe_prompt.txt`.
`confidence` is 0-100, not 0-1, in both probe files, and the pre-answer file leaves it empty for all 222 fallback rows rather than filling in a value. The documented gate treats an empty confidence as escalate. Get either convention wrong and you get a plausible, wrong number: on the 423 test problems the correct reading reproduces the published 78.0%, while misreading the scale gives 60.76% and treating empties as "stay" gives 73.76%. See `docs/schema.md`.
Join the probe predictions on `example_id`, never on `problem_uid_run`. The Gemma and gpt-oss runs use different native identifier schemes for the same problems. Joining on the native id gives zero overlap for all three Gemma settings while the three gpt-oss settings join perfectly — it silently drops half the data and still looks like it worked. registry/example_id_crosswalk.csv maps (setting_id, example_id) -> problem_id so you never hit this.
Dataset Structure
Data Files
All measured row counts, verified against the files in this repository.
Data Splits
Disjointness is checked by validate.py, which fails if any problem appears in more than one split.
Data Fields
The key fields of the flagship table, matched_labels.csv:
oracle_label takes exactly one of baseline_llm, single_agent, per, broadcast, none. Do not trust it blindly — recompute it from the four outcome columns; validate.py does exactly that and reports zero mismatches across all 15,088 rows.
Full field-level documentation, including the confidence scale and null handling, is in `docs/schema.md`.
Leakage warning — read this before training anything
Features and labels are in separate files, on purpose. Do not merge them into your model input.
Three specific hazards:
- The six-setting splits are drawn independently per setting. A problem can be in
trainfor gpt-oss-120b and intestfor Gemma-4-31B-it. Measured cross-solver split agreement is about 54%. If you pool both solvers' rows and train one model, you will leak. Split onproblem_idyourself if you pool. - The oracle label is not a router input. It is computed from the outcomes you are trying to predict. It is deliberately absent from both split files.
- `baseline_correct` is an outcome, not a feature. Several analyses condition on it, which is legitimate for measurement but not for a router that must decide before any execution.
Quick start
Download
hf download AgentsSci/EMNLP_Cost-Aware-Protocol-Routing --repo-type dataset --local-dir data/emnlp_protocol_routingThe Python equivalent:
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="AgentsSci/EMNLP_Cost-Aware-Protocol-Routing",
repo_type="dataset",
local_dir="data/emnlp_protocol_routing",
)Validate what you downloaded
python data/emnlp_protocol_routing/validate.pyIt checks the release and schema versions, every expected file, all row counts, benchmark/model/protocol coverage, the oracle recomputation, split disjointness, the leakage guard, and every sha256 in registry/checksums.sha256, then reports which optional artifacts are absent by design. It exits nonzero on any failure.
Load
from datasets import load_dataset
# flagship table: one row per (setting, problem)
d = load_dataset("AgentsSci/EMNLP_Cost-Aware-Protocol-Routing", "matched_labels")["all"]
# feature side, no labels
p = load_dataset("AgentsSci/EMNLP_Cost-Aware-Protocol-Routing", "problems")["all"]Or just read the CSVs — everything under 50 MB is plain CSV on purpose:
import pandas as pd
m = pd.read_csv("data/matched_labels.csv") # 15,088 rows
m.groupby("setting_id").oracle_label.value_counts(normalize=True)Recompute the oracle yourself in one line, and confirm it matches:
import numpy as np
rec = np.select(
[m.baseline_correct == 1, m.single_correct == 1,
m.per_correct == 1, m.broadcast_correct == 1],
["baseline_llm", "single_agent", "per", "broadcast"],
default="none")
assert (rec == m.oracle_label).all() # holds for all 15,088 rowsPer-benchmark licensing
Project-authored content (documentation, registry tables, all derived measurements) is CC-BY-4.0. The reasoning, stated plainly: the released tables are our own measurements of model behaviour, and no upstream problem text or gold answers are redistributed here, so each upstream benchmark's own license continues to govern its own content and LAB-Bench's CC-BY-SA-4.0 share-alike obligation is not triggered — there is no LAB-Bench content in this release to share alike. Model terms: gpt-oss-120b weights are Apache-2.0; Gemma-4-31B-it is governed by the Gemma Terms of Use. Full notice in `LICENSE` and `NOTICE`.
Intended and out-of-scope use
Intended. Training and evaluating cost-aware routers; studying calibration and failure prediction; measuring how much collaboration is actually worth on a given problem; benchmarking against a retrospective oracle without paying for four protocol executions.
Out of scope. Treating the oracle labels as repeated-sampling expected optima. Reading per-protocol solve rates as general claims about those protocols outside these benchmarks and these two solvers. Using the identifiers here to assemble a training corpus over LAB-Bench, against its upstream do-not-train request. Any commercial use of MaScQA-derived work — which is moot here, since MaScQA is not included.
Verification
Every claim below was recomputed at staging time, not copied:
- Recomputing
oracle_labelfrom the four correctness flags in fixed order reproduces the stored label for all 15,088 rows, zero mismatches. - The per-problem labels reproduce the paper's
matched_protocol_coverageandoracle_label_distributiontables for all ten settings, to two decimals. - Both solvers cover an identical canonical problem set in all five paired settings.
- Splits are pairwise disjoint and exhaustive in both schemes.
Known discrepancies are stated, not hidden, in `docs/provenance.md` — including a 3-row disagreement between two source tables and a 73% clean-parse rate on the primary confidence probe.
Links
- Paper: <https://arxiv.org/abs/2608.14927>
- Code: <https://github.com/ChihHsuan-Yang/EMNLP_Cost-Aware-Protocol-Routing>
- Project site: <https://chihhsuan-yang.github.io/EMNLP_Cost-Aware-Protocol-Routing/>
- Model card: <https://huggingface.co/AgentsSci/EMNLP_Cost-Aware-Protocol-Routing>
Citation
@inproceedings{yang2026costaware,
title = {LLMs Can Predict Failure Risk, But Struggle to Predict Which
Collaboration Protocol Pays Off: Cost-Aware Protocol Routing
Across Reasoning Tasks},
author = {Yang, Chih-Hsuan and Jiang, Jingyan and Yang, Cheng-Hau and
Vasudevan, Vikram and Zheng, Huihuo and Vishwanath, Venkatram and
Thakur, Rajeev},
booktitle = {Proceedings of the 2026 Conference on Empirical Methods in
Natural Language Processing (EMNLP)},
year = {2026},
eprint = {2608.14927},
archivePrefix = {arXiv}
}Authors. Chih-Hsuan Yang<sup>1</sup>, Jingyan Jiang<sup>1</sup>, Cheng-Hau Yang<sup>1</sup>, Vikram Vasudevan<sup>2</sup>, Huihuo Zheng<sup>1</sup>, Venkatram Vishwanath<sup>1</sup>, Rajeev Thakur<sup>1</sup>
<sup>1</sup> Argonne National Laboratory, Lemont, IL, USA <sup>2</sup> Oregon State University, Corvallis, OR, USA
Contact: bellayang@anl.gov
Acknowledgment
This research used resources of the Argonne Leadership Computing Facility, a U.S. Department of Energy (DOE) Office of Science user facility at Argonne National Laboratory (ANL) operated under Contract No. DE-AC02-06CH11357.
