skinny-cloud/dx3-recall-generalization-benchmark
Dx3 Recall Generalization Benchmark — held-out v0 Author: Asif Waliuddin · NXTG.AI · CC BY 4.0 A held-out recall goldenset (20 query→target pairs), disjoint from the tuning set, built to test whether a retrieval fix generalizes rather than memorizes the set it was tuned against. Every target was full-text-confirmed present in the live store before inclusion — so a recall@5 miss means weak retrieval/ranking, not absence. Honest methodology artifact: a generalization gate, not a… See the full description on the dataset page: https://huggingface.co/datasets/skinny-cloud/dx3-recall-generalization-benchmark.
Dx3 Recall Generalization Benchmark — held-out v0
Author: Asif Waliuddin · NXTG.AI · CC BY 4.0
A held-out recall goldenset (20 query→target pairs), disjoint from the tuning set, built to test whether a retrieval fix generalizes rather than memorizes the set it was tuned against. Every target was full-text-confirmed present in the live store before inclusion — so a recall@5 miss means weak retrieval/ranking, not absence. Honest methodology artifact: a generalization gate, not a leaderboard.
heldout-goldenset-v0.json— the 20 pairs (id, question, target).
Companion research: CRUCIBLE — Measurement Integrity (DOI 10.57967/hf/9867).
