Team Ai
Datasetpublic

skinny-cloud/dx3-recall-generalization-benchmark

Dx3 Recall Generalization Benchmark — held-out v0 Author: Asif Waliuddin · NXTG.AI · CC BY 4.0 A held-out recall goldenset (20 query→target pairs), disjoint from the tuning set, built to test whether a retrieval fix generalizes rather than memorizes the set it was tuned against. Every target was full-text-confirmed present in the live store before inclusion — so a recall@5 miss means weak retrieval/ranking, not absence. Honest methodology artifact: a generalization gate, not a… See the full description on the dataset page: https://huggingface.co/datasets/skinny-cloud/dx3-recall-generalization-benchmark.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes5downloads
Dataset Card

Dx3 Recall Generalization Benchmark — held-out v0

Author: Asif Waliuddin · NXTG.AI · CC BY 4.0

A held-out recall goldenset (20 query→target pairs), disjoint from the tuning set, built to test whether a retrieval fix generalizes rather than memorizes the set it was tuned against. Every target was full-text-confirmed present in the live store before inclusion — so a recall@5 miss means weak retrieval/ranking, not absence. Honest methodology artifact: a generalization gate, not a leaderboard.

  • —heldout-goldenset-v0.json — the 20 pairs (id, question, target).

Companion research: CRUCIBLE — Measurement Integrity (DOI 10.57967/hf/9867).