CinderD/wildtrace
WildTrace strict481 WildTrace is a source-internal long-context multi-hop reasoning benchmark built from natural evidence trails. Unlike reverse-synthetic QA, its tasks are mined in situ from long source documents before questions are written. The strict481 release contains 481 locked tasks over 214 public long-form sources, with full-document, evidence-withheld evaluation. The model under test receives only the source document and the public question; evidence spans, clue… See the full description on the dataset page: https://huggingface.co/datasets/CinderD/wildtrace.
WildTrace strict481
WildTrace is a source-internal long-context multi-hop reasoning benchmark built from natural evidence trails. Unlike reverse-synthetic QA, its tasks are mined in situ from long source documents before questions are written. The strict481 release contains 481 locked tasks over 214 public long-form sources, with full-document, evidence-withheld evaluation. The model under test receives only the source document and the public question; evidence spans, clue counts, reference answers, rubrics, and validation notes are hidden.
Freeze id: wildtrace_strict481_public_20260710. Release version: wildtrace-strict481-public-2026-07-10. Public task and source identifiers are opaque and stable within this release.
Contents
data/wildtrace_strict481.with_answers.json: full nested benchmark artifact with questions, ground truth, rubrics, and evidence clues.data/wildtrace_strict481.with_answers.jsonl: Hugging Face / Arrow-friendly view of the same rows, withground_truthandcluesserialized as JSON strings.data/wildtrace_strict481.questions_only.jsonand.jsonl: model-input rows without answers/rubrics/clues.corpus/*.txt: the 214 source documents referenced by the tasks.eval/run_eval.py,eval/run_judge.py,eval/config.json: minimal evidence-withheld evaluation and rubric-judge harness.methodology/EVAL_PROTOCOL.md: prompt templates, generation settings, context policy, judge prompt, scoring normalization, and aggregation rules.metadata/: schema, corpus manifest, release manifest, checksums, and dataset statistics.croissant.json: MLCommons Croissant metadata with Responsible AI fields.
Snapshot Statistics
Evidence Geometries
Context Tiers
Source Families
Usage
Local package:
from datasets import load_dataset
ds = load_dataset(
"json",
data_files="data/wildtrace_strict481.with_answers.jsonl",
split="train",
)
print(ds[0]["question_id"])
print(ds[0]["question_text"])After upload to Hugging Face:
from datasets import load_dataset
ds = load_dataset("CinderD/wildtrace", "with_answers", split="test")To evaluate a model, clone the complete dataset repository, configure the model and judge endpoints in eval/config.json, and run:
cd eval
python run_eval.py --config config.json \
--data ../data/wildtrace_strict481.with_answers.json \
--corpus ../corpus --out ../results/mymodel.responses.json
python run_judge.py --config config.json \
--data ../data/wildtrace_strict481.with_answers.json \
--responses ../results/mymodel.responses.json \
--out ../results/mymodel.scores.jsonThe scripts create the output directory automatically. Both the nested .json and Hugging Face .jsonl representations are accepted. See methodology/EVAL_PROTOCOL.md before reporting new results.
Scoring Denominators
The score file reports both scored_overall (valid-response quality) and all_tasks_overall/overall (coverage-sensitive All481 quality). Missing, failed, and out-of-context model responses receive zero in All481. A judge API or parse failure leaves overall null and exits nonzero instead of silently averaging a partial judge panel; rerun the same command to resume.
Maintenance Notes
- 2026-08-13: normalized three legacy rubric encodings, added HF JSONL support, automatic output-directory creation, complete-panel judging, and explicit Scored/All481 aggregation. Refreshed release checksums.
Quality Assurance
Each retained item is tied to source-grounded clues, a criterion-level rubric, and validation records used during construction. The public package strips internal iteration metadata while retaining the artifacts needed to inspect the released tasks. The final release passed checks for paradigm validity, leave-one-out multi-hop necessity, grounding, uniqueness, contamination resistance, and coherent-source constraints.
Responsible AI
The source corpus consists of public-domain literary works, Chinese literary sources, and publicly available technical incident reports. Some sources mention injury, death, violence, or distressing events because they are faithful to the original public documents. WildTrace is intended for diagnostic evaluation of source-grounded long-context reasoning, not for deployment certification in legal, medical, safety-critical, financial, or investigative settings.
Public release can increase benchmark contamination risk. New leaderboard claims should disclose whether a model may have seen the benchmark items or source documents during training.
License
WildTrace is a mixed-license release. Benchmark annotations, questions, rubrics, metadata, and scripts are released under CC BY 4.0. Source documents are not relicensed by WildTrace: they retain their original public-domain, public-agency, Project Gutenberg, or source-specific terms, which can vary by jurisdiction and by source. Users are responsible for checking source-specific terms before redistributing or adapting source texts.
Citation
@article{chen2026wildtrace,
title={WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning},
author={Chen, Zixin and Liu, Peng and Li, Haobo and Sheng, Rui and Tu, Jianhong and Deng, Xiaodong and Huang, Fei and Shum, Kashun and Liu, Dayiheng and Qu, Huamin},
year={2026},
journal={arXiv preprint arXiv:2607.09328},
url={https://arxiv.org/abs/2607.09328}
}