muradil211/AetherSearch_Eval_1400
🔭 AetherSearch Eval-1400 One frozen benchmark for training-time evaluation and final checkpoint assessment 🏠 Project · 🎓 SFT Data · 🤖 SFT Model · ⚖️ DPO Data · 🧠 DPO Model Dataset overview AetherSearch Eval-1400 is a frozen, 1,400-question evaluation suite for agentic search. It combines seven official held-out QA sources and isolates their questions from the audited AetherSearch SFT, DPO, and RL training inputs. This… See the full description on the dataset page: https://huggingface.co/datasets/muradil211/AetherSearch_Eval_1400.
<div align="center">
<img src="assets/aethersearch-mark.svg" alt="AetherSearch monogram" width="140">
🔭 AetherSearch Eval-1400
One frozen benchmark for training-time evaluation and final checkpoint assessment
<p> <img src="https://img.shields.io/badge/Questions-1%2C400-F59E0B?style=flat-square" alt="1,400 questions"> <img src="https://img.shields.io/badge/Sources-7-2563EB?style=flat-square" alt="7 source datasets"> <img src="https://img.shields.io/badge/Release-Frozen%20v1-0F766E?style=flat-square" alt="Frozen v1 release"> <img src="https://img.shields.io/badge/Seed-42-7C3AED?style=flat-square" alt="Sampling seed 42"> <img src="https://img.shields.io/badge/Training%20Overlap-0-0F766E?style=flat-square" alt="Zero normalized training-question overlap"> </p>
🏠 Project · 🎓 SFT Data · 🤖 SFT Model · ⚖️ DPO Data · 🧠 DPO Model
</div>
Dataset overview
AetherSearch Eval-1400 is a frozen, 1,400-question evaluation suite for agentic search. It combines seven official held-out QA sources and isolates their questions from the audited AetherSearch SFT, DPO, and RL training inputs.
This is the dataset for `SFT_eval`: the same 1,400 questions are used both during SFT training and after training finishes. The frozen release also provides a common benchmark for subsequent Base, SFT, DPO, and RL checkpoint comparisons.
The Hub exposes one split, `validation`, backed by eval_1400.jsonl. Training-time and final evaluation reuse that exact split, question order, and gold aliases. A final score on this suite is therefore a score on the shared evaluation benchmark; this release does not supply a separate final holdout.
Evaluation contract
All 1,400 questions remain evaluation data and are excluded from gradient training, preference-pair construction, and RL optimization. Gold aliases and official supporting evidence must stay outside normal agent prompts.
The release fixes questions and reference answers. For checkpoint comparisons, also keep the scoring implementation, retrieval corpus, tool budget, prompt, and decoding settings consistent. No model performance results are included.
Release at a glance
Source composition
Bamboogle supplies 125 questions. Its 75-question shortfall from the initial 200-per-source proposal is distributed evenly across the other six sources: 12 extra questions each, with the remaining three assigned to NQ, TriviaQA, and PopQA in the fixed source order. The resulting composition is frozen.
Each record preserves its original source_split; the combined Hub split name validation describes how AetherSearch consumes this benchmark. Source revisions, download locations, and source-file SHA-256 values are recorded in eval_1400_manifest.json.
Public schema
NQ-Open and Bamboogle do not provide original example IDs. Their source_id is null; row numbers, filenames, and the source hashes in the manifest provide their original identity.
The main evaluation file contains questions, gold aliases, and provenance. eval_1400_rollout_questions.jsonl contains exactly id and question for the same 1,400 rows in the same order. The diagnostic file contains supporting metadata only for sources that provide it, joined through the benchmark ID.
Load with Datasets
from datasets import load_dataset
benchmark = load_dataset(
"muradil211/AetherSearch_Eval_1400",
revision="frozen-v1",
split="validation",
)
assert len(benchmark) == 1400
# Use this same benchmark during SFT training and after training completes.
agent_inputs = benchmark.select_columns(["id", "question"])
gold_by_id = {row["id"]: row["answers"] for row in benchmark}
# Pass agent_inputs to the rollout runner; keep gold_by_id in the scorer.
print(agent_inputs[0])For a runner that consumes JSONL directly, download the question-only file:
from huggingface_hub import hf_hub_download
rollout_path = hf_hub_download(
repo_id="muradil211/AetherSearch_Eval_1400",
repo_type="dataset",
revision="frozen-v1",
filename="eval_1400_rollout_questions.jsonl",
)Curation and integrity
The frozen build follows this sequence:
- Register the current canonical SFT, DPO, and RL training releases and historical versions supported by actual training logs or release evidence.
- Normalize every training and candidate question with Unicode NFKC, casefold, whitespace collapse, strip, and removal of trailing ASCII or full-width question marks.
- Remove exact training overlaps, conservative punctuation/template equivalents, and internal or cross-source normalized duplicates.
- Sample deterministically with seed 42. Candidates are sorted by source ID, normalized question, and original row number before a separate seeded shuffle for each source. Reviewed exclusions trigger the same deterministic selection again; a shortage fails the build.
- Audit the selected questions against the complete training-question union with exhaustive character-similarity retrieval and rare-token retrieval. Resolve each flagged pair through a conservative question-specific review.
- Independently reconstruct the exclusion sets from original training files, verify selected questions and aliases against original held-out rows, replay sampling, and verify that training-file hashes stayed unchanged.
Stage exclusion sets overlap with one another; their union contains 186,326 questions. The counts include evidenced historical training inputs alongside the canonical releases of 2,600 SFT rows, 2,126 DPO rows, and 169,615 RL rows.
Across the 51,713 official candidate rows, isolation removed 877 exact training overlaps, 12 automatic near-duplicates, and 1,216 internal duplicates. Review removed 32 additional near-duplicate candidates. The final 1,400-question scan flagged 33 pairs involving 16 questions; all were resolved with explicit reasons for retaining distinct factual questions. No known high-confidence training near-duplicate remained.
Similarity retrieves pairs for review; it is not an automatic semantic deletion rule. Changed entities, years, answer slots, or reasoning targets are retained when they ask different questions. This audit does not claim detection of every possible arbitrary semantic paraphrase. Every exclusion and reviewed decision is recorded in eval_1400_exclusion_audit.json.
Files
The benchmark and its five accompanying frozen JSON/JSONL artifacts are published byte-for-byte from the verified local release. The exclusion audit preserves the build-time training-source registry. Raw training files and historical training logs remain outside this repository.
Checksums and versioning
The canonical data SHA-256 is:
ecd634aac24013213ceca3523dc024d82544c5479546fe205da24d8e8ffea678After downloading the repository, verify it without loading a model:
sha256sum -c checksums.sha256
python verify_release.pyPin revision="frozen-v1" when comparing checkpoints. Changes to question membership, gold aliases, composition, or sampling require a new benchmark version. This repository publishes the newly curated Eval-1400 suite; the existing AetherSearch_Eval repository documents the earlier full 51,713-row evaluation release.
Attribution and rights
The upstream QA datasets retain their applicable terms and attribution requirements. This combined release does not assert a blanket license over all source questions, aliases, or supporting metadata; the card therefore uses license: unknown, consistent with the AetherSearch SFT/DPO dataset cards. See ATTRIBUTION.md and the pinned source references in the manifest.
