Team Ai
Datasetpublic

muradil211/AetherSearch_Eval_1400

🔭 AetherSearch Eval-1400 One frozen benchmark for training-time evaluation and final checkpoint assessment 🏠 Project · 🎓 SFT Data · 🤖 SFT Model · ⚖️ DPO Data · 🧠 DPO Model Dataset overview AetherSearch Eval-1400 is a frozen, 1,400-question evaluation suite for agentic search. It combines seven official held-out QA sources and isolates their questions from the audited AetherSearch SFT, DPO, and RL training inputs. This… See the full description on the dataset page: https://huggingface.co/datasets/muradil211/AetherSearch_Eval_1400.

sourceHugging Faceunknownupdated 9d agoView on Hugging Face
1likes767downloads
Dataset Card

<div align="center">

<img src="assets/aethersearch-mark.svg" alt="AetherSearch monogram" width="140">

🔭 AetherSearch Eval-1400

One frozen benchmark for training-time evaluation and final checkpoint assessment

<p> <img src="https://img.shields.io/badge/Questions-1%2C400-F59E0B?style=flat-square" alt="1,400 questions"> <img src="https://img.shields.io/badge/Sources-7-2563EB?style=flat-square" alt="7 source datasets"> <img src="https://img.shields.io/badge/Release-Frozen%20v1-0F766E?style=flat-square" alt="Frozen v1 release"> <img src="https://img.shields.io/badge/Seed-42-7C3AED?style=flat-square" alt="Sampling seed 42"> <img src="https://img.shields.io/badge/Training%20Overlap-0-0F766E?style=flat-square" alt="Zero normalized training-question overlap"> </p>

🏠 Project · 🎓 SFT Data · 🤖 SFT Model · ⚖️ DPO Data · 🧠 DPO Model

</div>


Dataset overview

AetherSearch Eval-1400 is a frozen, 1,400-question evaluation suite for agentic search. It combines seven official held-out QA sources and isolates their questions from the audited AetherSearch SFT, DPO, and RL training inputs.

This is the dataset for `SFT_eval`: the same 1,400 questions are used both during SFT training and after training finishes. The frozen release also provides a common benchmark for subsequent Base, SFT, DPO, and RL checkpoint comparisons.

The Hub exposes one split, `validation`, backed by eval_1400.jsonl. Training-time and final evaluation reuse that exact split, question order, and gold aliases. A final score on this suite is therefore a score on the shared evaluation benchmark; this release does not supply a separate final holdout.

Evaluation contract

WhenUse
During SFT trainingPeriodic SFT_eval of intermediate checkpoints on the frozen validation split
After SFT trainingEvaluate the completed SFT checkpoint on the same full split
Base / SFT / DPO / RL comparisonReuse the same revision, 1,400 question IDs, gold aliases, and evaluation settings
Agent rolloutSupply only id and question from the question-only file or projection
ScoringKeep answers in the evaluator and join predictions by id
Evidence diagnosticsRead supporting facts from the separate diagnostic sidecar when auditing

All 1,400 questions remain evaluation data and are excluded from gradient training, preference-pair construction, and RL optimization. Gold aliases and official supporting evidence must stay outside normal agent prompts.

The release fixes questions and reference answers. For checkpoint comparisons, also keep the scoring implementation, retrieval corpus, tool budget, prompt, and decoding settings consistent. No model performance results are included.

Release at a glance

ItemValue
Evaluation questions1,400
Source datasets7
Unique raw / normalized questions1,400 / 1,400
Hub splitvalidation
Sampling seed42
Training-question exclusion union186,326 unique normalized questions
Exact SFT / DPO / RL question overlap0 / 0 / 0
Internal / cross-dataset duplicates0 / 0
Records with non-empty gold aliases1,400
Records with verified provenance1,400
Known high-confidence training near-duplicates after audit0
Release revisionfrozen-v1

Source composition

Source datasetQuestionsOfficial source splitRelease variant
NQ213devNQ-Open
TriviaQA213devunfiltered.nocontext; author release calls this validation
PopQA213testOfficial test.tsv
HotpotQA212devDistractor v1; author release calls this validation
2WikiMultiHopQA212devOfficial April 2021 release and entity aliases
MuSiQue212devMuSiQue-Ans v1.0, answerable questions
Bamboogle125evalComplete official 125-question evaluation release
Total1,400Official held-out partitions only

Bamboogle supplies 125 questions. Its 75-question shortfall from the initial 200-per-source proposal is distributed evenly across the other six sources: 12 extra questions each, with the remaining three assigned to NQ, TriviaQA, and PopQA in the fixed source order. The resulting composition is frozen.

Each record preserves its original source_split; the combined Hub split name validation describes how AetherSearch consumes this benchmark. Source revisions, download locations, and source-file SHA-256 values are recorded in eval_1400_manifest.json.

Public schema

FieldTypeMeaning
idstringFrozen benchmark ID, eval1400_0001 through eval1400_1400
questionstringOriginal held-out question text
answerslist[string]Non-empty official answer aliases for the scorer
source_datasetstringOne of the seven source names above
source_splitstringOriginal official partition: dev, test, or eval
source_idstring or nullOriginal example ID when supplied by the source
source_row_numberintegerOne-based row location in the pinned source file
source_filestringSource-file identity within the provenance snapshot

NQ-Open and Bamboogle do not provide original example IDs. Their source_id is null; row numbers, filenames, and the source hashes in the manifest provide their original identity.

The main evaluation file contains questions, gold aliases, and provenance. eval_1400_rollout_questions.jsonl contains exactly id and question for the same 1,400 rows in the same order. The diagnostic file contains supporting metadata only for sources that provide it, joined through the benchmark ID.

Load with Datasets

python
from datasets import load_dataset

benchmark = load_dataset(
    "muradil211/AetherSearch_Eval_1400",
    revision="frozen-v1",
    split="validation",
)

assert len(benchmark) == 1400

# Use this same benchmark during SFT training and after training completes.
agent_inputs = benchmark.select_columns(["id", "question"])
gold_by_id = {row["id"]: row["answers"] for row in benchmark}

# Pass agent_inputs to the rollout runner; keep gold_by_id in the scorer.
print(agent_inputs[0])

For a runner that consumes JSONL directly, download the question-only file:

python
from huggingface_hub import hf_hub_download

rollout_path = hf_hub_download(
    repo_id="muradil211/AetherSearch_Eval_1400",
    repo_type="dataset",
    revision="frozen-v1",
    filename="eval_1400_rollout_questions.jsonl",
)

Curation and integrity

The frozen build follows this sequence:

  1. 1.Register the current canonical SFT, DPO, and RL training releases and historical versions supported by actual training logs or release evidence.
  2. 2.Normalize every training and candidate question with Unicode NFKC, casefold, whitespace collapse, strip, and removal of trailing ASCII or full-width question marks.
  3. 3.Remove exact training overlaps, conservative punctuation/template equivalents, and internal or cross-source normalized duplicates.
  4. 4.Sample deterministically with seed 42. Candidates are sorted by source ID, normalized question, and original row number before a separate seeded shuffle for each source. Reviewed exclusions trigger the same deterministic selection again; a shortage fails the build.
  5. 5.Audit the selected questions against the complete training-question union with exhaustive character-similarity retrieval and rare-token retrieval. Resolve each flagged pair through a conservative question-specific review.
  6. 6.Independently reconstruct the exclusion sets from original training files, verify selected questions and aliases against original held-out rows, replay sampling, and verify that training-file hashes stayed unchanged.
Training stageAudited input rowsUnique normalized exclusion questionsFinal exact overlap
SFT60,20630,0930
DPO18,58415,2690
RL169,615169,6040

Stage exclusion sets overlap with one another; their union contains 186,326 questions. The counts include evidenced historical training inputs alongside the canonical releases of 2,600 SFT rows, 2,126 DPO rows, and 169,615 RL rows.

Across the 51,713 official candidate rows, isolation removed 877 exact training overlaps, 12 automatic near-duplicates, and 1,216 internal duplicates. Review removed 32 additional near-duplicate candidates. The final 1,400-question scan flagged 33 pairs involving 16 questions; all were resolved with explicit reasons for retaining distinct factual questions. No known high-confidence training near-duplicate remained.

Similarity retrieves pairs for review; it is not an automatic semantic deletion rule. Changed entities, years, answer slots, or reasoning targets are retained when they ask different questions. This audit does not claim detection of every possible arbitrary semantic paraphrase. Every exclusion and reviewed decision is recorded in eval_1400_exclusion_audit.json.

Files

FilePurpose
eval_1400.jsonlCanonical 1,400-row evaluation split, with evaluator gold aliases
eval_1400_rollout_questions.jsonlSame IDs and questions, with exactly two fields for agent input
eval_1400_manifest.jsonFrozen composition, pinned official sources, isolation checks, and hashes
eval_1400_exclusion_audit.jsonTraining-source registry, exclusions, and near-duplicate review decisions
eval_1400_diagnostics.jsonlOfficial supporting facts and diagnostic metadata; audit use only
independent_validation.jsonIndependent source, gold, training-isolation, and replay verification report
verify_release.pyStandalone release-file, row-integrity, and rollout-projection verifier
checksums.sha256SHA-256 values for all published content files
ATTRIBUTION.mdUpstream attribution and rights-status record
assets/aethersearch-mark.svgAetherSearch release mark

The benchmark and its five accompanying frozen JSON/JSONL artifacts are published byte-for-byte from the verified local release. The exclusion audit preserves the build-time training-source registry. Raw training files and historical training logs remain outside this repository.

Checksums and versioning

The canonical data SHA-256 is:

text
ecd634aac24013213ceca3523dc024d82544c5479546fe205da24d8e8ffea678

After downloading the repository, verify it without loading a model:

bash
sha256sum -c checksums.sha256
python verify_release.py

Pin revision="frozen-v1" when comparing checkpoints. Changes to question membership, gold aliases, composition, or sampling require a new benchmark version. This repository publishes the newly curated Eval-1400 suite; the existing AetherSearch_Eval repository documents the earlier full 51,713-row evaluation release.

Attribution and rights

The upstream QA datasets retain their applicable terms and attribution requirements. This combined release does not assert a blanket license over all source questions, aliases, or supporting metadata; the card therefore uses license: unknown, consistent with the AetherSearch SFT/DPO dataset cards. See ATTRIBUTION.md and the pinned source references in the manifest.