Team Ai
Datasetpublic

CinderD/wildtrace

WildTrace strict481 WildTrace is a source-internal long-context multi-hop reasoning benchmark built from natural evidence trails. Unlike reverse-synthetic QA, its tasks are mined in situ from long source documents before questions are written. The strict481 release contains 481 locked tasks over 214 public long-form sources, with full-document, evidence-withheld evaluation. The model under test receives only the source document and the public question; evidence spans, clue… See the full description on the dataset page: https://huggingface.co/datasets/CinderD/wildtrace.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
1likes298downloads
Dataset Card

WildTrace strict481

WildTrace is a source-internal long-context multi-hop reasoning benchmark built from natural evidence trails. Unlike reverse-synthetic QA, its tasks are mined in situ from long source documents before questions are written. The strict481 release contains 481 locked tasks over 214 public long-form sources, with full-document, evidence-withheld evaluation. The model under test receives only the source document and the public question; evidence spans, clue counts, reference answers, rubrics, and validation notes are hidden.

Freeze id: wildtrace_strict481_public_20260710. Release version: wildtrace-strict481-public-2026-07-10. Public task and source identifiers are opaque and stable within this release.

Contents

  • —data/wildtrace_strict481.with_answers.json: full nested benchmark artifact with questions, ground truth, rubrics, and evidence clues.
  • —data/wildtrace_strict481.with_answers.jsonl: Hugging Face / Arrow-friendly view of the same rows, with ground_truth and clues serialized as JSON strings.
  • —data/wildtrace_strict481.questions_only.json and .jsonl: model-input rows without answers/rubrics/clues.
  • —corpus/*.txt: the 214 source documents referenced by the tasks.
  • —eval/run_eval.py, eval/run_judge.py, eval/config.json: minimal evidence-withheld evaluation and rubric-judge harness.
  • —methodology/EVAL_PROTOCOL.md: prompt templates, generation settings, context policy, judge prompt, scoring normalization, and aggregation rules.
  • —metadata/: schema, corpus manifest, release manifest, checksums, and dataset statistics.
  • —croissant.json: MLCommons Croissant metadata with Responsible AI fields.

Snapshot Statistics

ItemValue
Tasks481
Source documents214
Languagesen: 441, zh: 40
Min estimated document tokens19,106
Median estimated document tokens302,084
Max estimated document tokens2,549,520

Evidence Geometries

GeometryCount
abductive_inference65
causal_attribution67
comparative67
counterfactual_reasoning67
forward_chain73
intersection_query71
temporal_reconstruction71

Context Tiers

TierCount
L0_<=128K90
L1_128K_181K59
L2_181K_256K61
L3_256K_362K58
L4_362K_512K59
L5_512K_724K58
L6_724K_1M93
L7_>1M3

Source Families

Source familyTask count
cjk_literature40
en_literature318
technical_report123

Usage

Local package:

python
from datasets import load_dataset

ds = load_dataset(
    "json",
    data_files="data/wildtrace_strict481.with_answers.jsonl",
    split="train",
)
print(ds[0]["question_id"])
print(ds[0]["question_text"])

After upload to Hugging Face:

python
from datasets import load_dataset

ds = load_dataset("CinderD/wildtrace", "with_answers", split="test")

To evaluate a model, clone the complete dataset repository, configure the model and judge endpoints in eval/config.json, and run:

bash
cd eval
python run_eval.py --config config.json \
  --data ../data/wildtrace_strict481.with_answers.json \
  --corpus ../corpus --out ../results/mymodel.responses.json
python run_judge.py --config config.json \
  --data ../data/wildtrace_strict481.with_answers.json \
  --responses ../results/mymodel.responses.json \
  --out ../results/mymodel.scores.json

The scripts create the output directory automatically. Both the nested .json and Hugging Face .jsonl representations are accepted. See methodology/EVAL_PROTOCOL.md before reporting new results.

Scoring Denominators

The score file reports both scored_overall (valid-response quality) and all_tasks_overall/overall (coverage-sensitive All481 quality). Missing, failed, and out-of-context model responses receive zero in All481. A judge API or parse failure leaves overall null and exits nonzero instead of silently averaging a partial judge panel; rerun the same command to resume.

Maintenance Notes

  • —2026-08-13: normalized three legacy rubric encodings, added HF JSONL support, automatic output-directory creation, complete-panel judging, and explicit Scored/All481 aggregation. Refreshed release checksums.

Quality Assurance

Each retained item is tied to source-grounded clues, a criterion-level rubric, and validation records used during construction. The public package strips internal iteration metadata while retaining the artifacts needed to inspect the released tasks. The final release passed checks for paradigm validity, leave-one-out multi-hop necessity, grounding, uniqueness, contamination resistance, and coherent-source constraints.

Responsible AI

The source corpus consists of public-domain literary works, Chinese literary sources, and publicly available technical incident reports. Some sources mention injury, death, violence, or distressing events because they are faithful to the original public documents. WildTrace is intended for diagnostic evaluation of source-grounded long-context reasoning, not for deployment certification in legal, medical, safety-critical, financial, or investigative settings.

Public release can increase benchmark contamination risk. New leaderboard claims should disclose whether a model may have seen the benchmark items or source documents during training.

License

WildTrace is a mixed-license release. Benchmark annotations, questions, rubrics, metadata, and scripts are released under CC BY 4.0. Source documents are not relicensed by WildTrace: they retain their original public-domain, public-agency, Project Gutenberg, or source-specific terms, which can vary by jurisdiction and by source. Users are responsible for checking source-specific terms before redistributing or adapting source texts.

Citation

bibtex
@article{chen2026wildtrace,
  title={WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning},
  author={Chen, Zixin and Liu, Peng and Li, Haobo and Sheng, Rui and Tu, Jianhong and Deng, Xiaodong and Huang, Fei and Shum, Kashun and Liu, Dayiheng and Qu, Huamin},
  year={2026},
  journal={arXiv preprint arXiv:2607.09328},
  url={https://arxiv.org/abs/2607.09328}
}