Team Ai
Datasetpublic

AgentsSci/EMNLP_Cost-Aware-Protocol-Routing

Cost-Aware Protocol Routing: Matched Protocol Outcomes The short version. We ran the same 6,803 reasoning problems through four different LLM collaboration setups — from a single direct answer up to a four-agent deliberation — and recorded, for every problem, which ones got it right. Then we asked whether a model can look at a problem beforehand and predict which setup is worth paying for. It can predict whether it will fail. It cannot predict which collaboration protocol will… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/EMNLP_Cost-Aware-Protocol-Routing.

sourceHugging Facecc-by-4.0updated 17d agoView on Hugging Face
0likes206downloads
Dataset Card

Cost-Aware Protocol Routing: Matched Protocol Outcomes

The short version. We ran the same 6,803 reasoning problems through four different LLM collaboration setups — from a single direct answer up to a four-agent deliberation — and recorded, for every problem, which ones got it right. Then we asked whether a model can look at a problem beforehand and predict which setup is worth paying for.

It can predict whether it will fail. It cannot predict which collaboration protocol will fix the failure. That gap is what this dataset is for.

This is the data release for the EMNLP 2026 paper *LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off*.

The four protocols

Every released problem was run under all four, with the same solver model.

ProtocolWhat it doesRelative cost
BaselineOne direct answer. No revision.cheapest
SingleOne solver iterates on its own answer with evaluator feedback.low
PERPlanner decomposes, executor solves, reviewer checks.medium
BroadcastSeveral agents deliberate over a shared channel.most expensive

The oracle label

For each problem, oracle_label is the first protocol that succeeded, scanned in the fixed cost order Baseline -> Single -> PER -> Broadcast. If all four failed, the label is none.

Two things to be clear about:

  • —`none` is not a fifth protocol. It is a router action: abstain, spend nothing, because nothing observed here worked. Four protocols were executed on every released problem.
  • —The oracle is retrospective. It is computed from one realized execution per protocol, not from repeated sampling. It is the best a router could have done on these specific runs, not a ground-truth best action. A protocol that failed once here might succeed on a resample.

What is in here

15,088 rows in the flagship table: 10 settings (5 benchmark conditions x 2 solver models), covering 6,803 distinct problems.

SettingBenchmarkConditionSolvern
omnimath2__competition_math_4181__gpt_oss_120bOmniMathcompetition mathgpt-oss-120b4,181
omnimath2__competition_math_4181__gemma_4_31bOmniMathcompetition mathGemma-4-31B-it4,181
labbench__text_no_tool__gpt_oss_120b_text_no_toolLAB-Benchtext, no toolsgpt-oss-120b1,542
labbench__text_no_tool__gemma_4_31b_text_no_toolLAB-Benchtext, no toolsGemma-4-31B-it1,542
labbench__llm_strict__gpt_oss_120bLAB-Benchstrict in-prompt evidencegpt-oss-120b741
labbench__llm_strict__gemma_4_31bLAB-Benchstrict in-prompt evidenceGemma-4-31B-it741
scibench__text_only__gpt_oss_120bSciBenchtext onlygpt-oss-120b565
scibench__text_only__gemma_4_31bSciBenchtext onlyGemma-4-31B-it565
jeebench__text_only__gpt_oss_120bJEEBenchtext onlygpt-oss-120b515
jeebench__text_only__gemma_4_31bJEEBenchtext onlyGemma-4-31B-it515

Alongside the outcomes: two documented split schemes, held-out router predictions and metrics for six settings, confidence-probe measurements, and the paper's eight aggregate result tables reproduced verbatim.

Full column-by-column documentation: `docs/schema.md`.

What is NOT in here, and why

No problem text. No gold answers. No answer options. No reference solutions.

The four upstream benchmarks carry four different licenses, and LAB-Bench — the most restrictive — is CC-BY-SA-4.0 with an upstream do-not-train request. Rather than ship a mixed-license text column that downstream users would have to untangle correctly, this release withholds text across the board and gives you stable identifiers plus reconstruction instructions instead.

That applies even to JEEBench and SciBench, whose MIT licenses would have permitted redistribution. The decision was made release-wide because the flagship table interleaves all four benchmarks in one file.

To attach the text yourself: `docs/reconstruction.md`. The licensing evidence: `docs/license_audit.md`.

Also not included: raw model generations, per-problem cost accounting for nine of the ten settings (one setting is covered, see below), MaScQA (excluded — NonCommercial, and not in the paper), and the Gemma-3-27B scope check. See `docs/provenance.md`.

Also not included: Baseline final-answer text. The post-answer probe consumed it, so data/probe_inputs.jsonl is an identifier and metadata manifest, not a runnable prompt set — rebuilding runnable probe inputs needs both the problem text rehydrated from upstream and that setting's Baseline final answer. Baseline answers are not released, and for three of the six probe settings they no longer exist in the project's own artifacts either, so the probe is not fully re-runnable by anyone. The probe's outputs are released in full, so the paper's numbers stay reproducible. Coverage is measured in `TODO.md`.

Cost: per-problem tokens, for one setting only

The paper's cost axis is tokens, and data/costs/omnimath_per_protocol_costs.csv gives per-problem token totals and model-call counts for all four protocols — but only for `omnimath2__competition_math_4181__gpt_oss_120b`, one of the ten settings. The other nine have no per-problem cost data here; their cost appears only through the aggregate tables.

These are protocol-level totals, summed over every model call the protocol made. That is the paper's accounting. They are not a single call's prompt+completion — Baseline averages 9.67 model calls per problem in this setting, so the two quantities differ by roughly that factor. If you compare against another dataset's total_tokens, check which quantity it measures first.

4,155 of 4,181 problems are covered. The 26 omitted problems form 13 duplicate-text pairs where a cost row cannot be attributed to a specific problem_id; they were dropped rather than guessed. As a result, the mean Baseline token count recomputed from this file is 18,432.0, while the paper's published figure over all 4,181 problems is 18,385.4. That gap is the arithmetic of the 26 dropped rows, not a disagreement — cite 18,385.4 for the paper, expect 18,432.0 from this file, and do not reconcile them by adjusting either.

Two confidence probes — do not confuse them

The dataset ships two different instruments. They ask different questions at different points in the pipeline, and they support different numbers in the paper. Every row of both files carries a probe_type column.

**post-answer probe****pre-answer probe**
probe_typepost_answer_pre_collaborationpre_answer_q1
filedata/confidence/postanswer_confidence_predictions.csvdata/confidence/primary_omnimath_confidence_predictions.csv
rows12,928 (all six analysed settings)839 (OmniMath primary split)
runsafter Baseline answers, before any collaborationbefore any solving
seesthe problem, allowed metadata, and the model's own Baseline final answerthe problem only
asks"is this Baseline answer correct?""how likely am I to solve this in one pass?"
supportsthe paper's headline failure-risk resultthe confidence-gate policy row

The title claim comes from the post-answer probe. Its released predictions reproduce data/aggregate/postanswer_confidence.csv for all six settings to four decimals — including the headline gpt-oss-120b OmniMath figures of 4,181 rows, 4,151 parseable, and 0.8847 failure AUROC.

Reproducibility limit. The post-answer probe cannot be fully re-run from released artifacts, because the Baseline final-answer string that the probe consumes was not persisted for the gpt-oss runs. The probe's outputs and metrics are fully released and independently reproducible. Recoverable coverage is 6,462 / 12,928 probe rows (50.0%): 100% for the three Gemma settings, 0% for the three gpt-oss settings. See `TODO.md`.

Neither probe ever sees the gold answer, the correctness label, the oracle label, or any protocol outcome. The post-answer prompt states that boundary explicitly and is injection-hardened — it marks the problem text and the baseline answer as untrusted data, tells the model not to follow instructions inside them, and enumerates what the model does not have. It ships verbatim at `docs/confidence_probe_prompt.txt`.

`confidence` is 0-100, not 0-1, in both probe files, and the pre-answer file leaves it empty for all 222 fallback rows rather than filling in a value. The documented gate treats an empty confidence as escalate. Get either convention wrong and you get a plausible, wrong number: on the 423 test problems the correct reading reproduces the published 78.0%, while misreading the scale gives 60.76% and treating empties as "stay" gives 73.76%. See `docs/schema.md`.

Join the probe predictions on `example_id`, never on `problem_uid_run`. The Gemma and gpt-oss runs use different native identifier schemes for the same problems. Joining on the native id gives zero overlap for all three Gemma settings while the three gpt-oss settings join perfectly — it silently drops half the data and still looks like it worked. registry/example_id_crosswalk.csv maps (setting_id, example_id) -> problem_id so you never hit this.

Dataset Structure

Data Files

All measured row counts, verified against the files in this repository.

FileRowsColsWhat it is
data/matched_labels.csv15,08815Start here. One row per (setting, problem): the four protocol outcomes and the fixed-order oracle label, for all 10 settings.
data/problems.csv6,80310Router-visible problem metadata. No labels — this is the feature side.
data/labels_for_scoring.csv12,92815The label side, kept in a separate file from the features on purpose.
data/probe_inputs.jsonl12,9289Identifier manifest for the confidence probe (not a runnable prompt set — see below).
data/confidence/postanswer_confidence_predictions.csv12,92812Post-answer probe outputs, 6 settings. Backs the headline result.
data/confidence/primary_omnimath_confidence_predictions.csv83915Pre-answer q1 probe on the primary split. A different instrument.
data/router/six_setting_test_predictions.csv5,83210Held-out predictions for 3 routers × 6 settings on identical test ids.
data/costs/omnimath_per_protocol_costs.csv4,15510Per-protocol token totals and model calls. One setting only.
data/splits/six_setting_splits.csv12,928670/15/15 stratified by oracle label, seed 20260712.
data/splits/primary_omnimath_splits.csv4,181580/10/10 stratified, seed 42, test n=423.
data/aggregate/*.csv6–24 each—The eight camera-ready aggregate tables, as published.
registry/*——Experiment, benchmark, model and protocol registries; schema; manifest; checksums; the example_id crosswalk.

Data Splits

Split setTrainDevTestSeedStratified by
Six-setting router9,0461,9381,94420260712oracle label, within setting
Primary (OmniMath)3,34241642342oracle label

Disjointness is checked by validate.py, which fails if any problem appears in more than one split.

Data Fields

The key fields of the flagship table, matched_labels.csv:

ColumnTypeMeaning
setting_idstringbenchmark__condition__solver, one of 10
problem_idstringStable problem identifier, unique within a setting
baseline_correct, single_correct, per_correct, broadcast_correct0/1Did that protocol solve this problem
oracle_labelenumFirst success in the fixed order; none if all four failed
any_protocol_solved0/1Did any of the four succeed
model, model_endpoint, domain, benchmark_id, slice_id, run_idstringProvenance

oracle_label takes exactly one of baseline_llm, single_agent, per, broadcast, none. Do not trust it blindly — recompute it from the four outcome columns; validate.py does exactly that and reports zero mismatches across all 15,088 rows.

Full field-level documentation, including the confidence scale and null handling, is in `docs/schema.md`.

Leakage warning — read this before training anything

Features and labels are in separate files, on purpose. Do not merge them into your model input.

FileRoleContains outcomes?
data/problems.csvfeature sideno
data/probe_inputs.jsonlfeature sideno
data/costs/omnimath_per_protocol_costs.csvfeature sideno
data/confidence/postanswer_confidence_predictions.csvfeature sideno
data/confidence/primary_omnimath_confidence_predictions.csvfeature sideno
data/matched_labels.csvlabel sideyes
data/labels_for_scoring.csvlabel sideyes
data/router/six_setting_test_predictions.csvlabel sideyes

Three specific hazards:

  1. 1.The six-setting splits are drawn independently per setting. A problem can be in train for gpt-oss-120b and in test for Gemma-4-31B-it. Measured cross-solver split agreement is about 54%. If you pool both solvers' rows and train one model, you will leak. Split on problem_id yourself if you pool.
  2. 2.The oracle label is not a router input. It is computed from the outcomes you are trying to predict. It is deliberately absent from both split files.
  3. 3.`baseline_correct` is an outcome, not a feature. Several analyses condition on it, which is legitimate for measurement but not for a router that must decide before any execution.

Quick start

Download

bash
hf download AgentsSci/EMNLP_Cost-Aware-Protocol-Routing --repo-type dataset --local-dir data/emnlp_protocol_routing

The Python equivalent:

python
from huggingface_hub import snapshot_download

snapshot_download(
    repo_id="AgentsSci/EMNLP_Cost-Aware-Protocol-Routing",
    repo_type="dataset",
    local_dir="data/emnlp_protocol_routing",
)

Validate what you downloaded

bash
python data/emnlp_protocol_routing/validate.py

It checks the release and schema versions, every expected file, all row counts, benchmark/model/protocol coverage, the oracle recomputation, split disjointness, the leakage guard, and every sha256 in registry/checksums.sha256, then reports which optional artifacts are absent by design. It exits nonzero on any failure.

Load

python
from datasets import load_dataset

# flagship table: one row per (setting, problem)
d = load_dataset("AgentsSci/EMNLP_Cost-Aware-Protocol-Routing", "matched_labels")["all"]

# feature side, no labels
p = load_dataset("AgentsSci/EMNLP_Cost-Aware-Protocol-Routing", "problems")["all"]

Or just read the CSVs — everything under 50 MB is plain CSV on purpose:

python
import pandas as pd
m = pd.read_csv("data/matched_labels.csv")          # 15,088 rows
m.groupby("setting_id").oracle_label.value_counts(normalize=True)

Recompute the oracle yourself in one line, and confirm it matches:

python
import numpy as np
rec = np.select(
    [m.baseline_correct == 1, m.single_correct == 1,
     m.per_correct == 1, m.broadcast_correct == 1],
    ["baseline_llm", "single_agent", "per", "broadcast"],
    default="none")
assert (rec == m.oracle_label).all()   # holds for all 15,088 rows

Per-benchmark licensing

BenchmarkUpstream licenseText redistributed here?Upstream source
Omni-MATH-2Apache-2.0nomartheballon/Omni-MATH-2
JEEBenchMITnodair-iitd/jeebench
SciBenchMITnomandyyyyii/scibench
LAB-BenchCC-BY-SA-4.0, do-not-train requestnofuturehouse/lab-bench
MaScQACC-BY-NC-SA-4.0benchmark excluded entirelyM3RG-IITD/MaScQA

Project-authored content (documentation, registry tables, all derived measurements) is CC-BY-4.0. The reasoning, stated plainly: the released tables are our own measurements of model behaviour, and no upstream problem text or gold answers are redistributed here, so each upstream benchmark's own license continues to govern its own content and LAB-Bench's CC-BY-SA-4.0 share-alike obligation is not triggered — there is no LAB-Bench content in this release to share alike. Model terms: gpt-oss-120b weights are Apache-2.0; Gemma-4-31B-it is governed by the Gemma Terms of Use. Full notice in `LICENSE` and `NOTICE`.

Intended and out-of-scope use

Intended. Training and evaluating cost-aware routers; studying calibration and failure prediction; measuring how much collaboration is actually worth on a given problem; benchmarking against a retrospective oracle without paying for four protocol executions.

Out of scope. Treating the oracle labels as repeated-sampling expected optima. Reading per-protocol solve rates as general claims about those protocols outside these benchmarks and these two solvers. Using the identifiers here to assemble a training corpus over LAB-Bench, against its upstream do-not-train request. Any commercial use of MaScQA-derived work — which is moot here, since MaScQA is not included.

Verification

Every claim below was recomputed at staging time, not copied:

  • —Recomputing oracle_label from the four correctness flags in fixed order reproduces the stored label for all 15,088 rows, zero mismatches.
  • —The per-problem labels reproduce the paper's matched_protocol_coverage and oracle_label_distribution tables for all ten settings, to two decimals.
  • —Both solvers cover an identical canonical problem set in all five paired settings.
  • —Splits are pairwise disjoint and exhaustive in both schemes.

Known discrepancies are stated, not hidden, in `docs/provenance.md` — including a 3-row disagreement between two source tables and a 73% clean-parse rate on the primary confidence probe.

Links

  • —Paper: <https://arxiv.org/abs/2608.14927>
  • —Code: <https://github.com/ChihHsuan-Yang/EMNLP_Cost-Aware-Protocol-Routing>
  • —Project site: <https://chihhsuan-yang.github.io/EMNLP_Cost-Aware-Protocol-Routing/>
  • —Model card: <https://huggingface.co/AgentsSci/EMNLP_Cost-Aware-Protocol-Routing>

Citation

bibtex
@inproceedings{yang2026costaware,
  title     = {LLMs Can Predict Failure Risk, But Struggle to Predict Which
               Collaboration Protocol Pays Off: Cost-Aware Protocol Routing
               Across Reasoning Tasks},
  author    = {Yang, Chih-Hsuan and Jiang, Jingyan and Yang, Cheng-Hau and
               Vasudevan, Vikram and Zheng, Huihuo and Vishwanath, Venkatram and
               Thakur, Rajeev},
  booktitle = {Proceedings of the 2026 Conference on Empirical Methods in
               Natural Language Processing (EMNLP)},
  year      = {2026},
  eprint    = {2608.14927},
  archivePrefix = {arXiv}
}

Authors. Chih-Hsuan Yang<sup>1</sup>, Jingyan Jiang<sup>1</sup>, Cheng-Hau Yang<sup>1</sup>, Vikram Vasudevan<sup>2</sup>, Huihuo Zheng<sup>1</sup>, Venkatram Vishwanath<sup>1</sup>, Rajeev Thakur<sup>1</sup>

<sup>1</sup> Argonne National Laboratory, Lemont, IL, USA <sup>2</sup> Oregon State University, Corvallis, OR, USA

Contact: bellayang@anl.gov

Acknowledgment

This research used resources of the Argonne Leadership Computing Facility, a U.S. Department of Energy (DOE) Office of Science user facility at Argonne National Laboratory (ANL) operated under Contract No. DE-AC02-06CH11357.