Team Ai
Datasetpublic

muradil211/AetherSearch_DPO

🔭 AetherSearch DPO Preference pairs for reasoning, retrieval, and evidence-grounded answers 🏠 Project · 🧠 DPO Model · 🧪 Training Code · 🎓 SFT Data · 🤖 SFT Model Dataset overview AetherSearch DPO contains 2,126 preference pairs for training an agentic-search policy after supervised fine-tuning. Every row provides one shared prompt, a preferred assistant continuation, and a non-preferred continuation. This repository is… See the full description on the dataset page: https://huggingface.co/datasets/muradil211/AetherSearch_DPO.

sourceHugging Faceunknownupdated 1mo agoView on Hugging Face
1likes67downloads
Dataset Card

<div align="center">

<img src="assets/aethersearch-mark.svg" alt="AetherSearch monogram" width="140">

🔭 AetherSearch DPO

Preference pairs for reasoning, retrieval, and evidence-grounded answers

<p> <img src="https://img.shields.io/badge/Pairs-2%2C126-F59E0B?style=flat-square" alt="2,126 preference pairs"> <img src="https://img.shields.io/badge/Split-Train%20Only-0F766E?style=flat-square" alt="Train split only"> <img src="https://img.shields.io/badge/Questions-Unique-2563EB?style=flat-square" alt="Unique questions"> <img src="https://img.shields.io/badge/Format-Chosen%20%2F%20Rejected-7C3AED?style=flat-square" alt="Chosen and rejected format"> </p>

🏠 Project · 🧠 DPO Model · 🧪 Training Code · 🎓 SFT Data · 🤖 SFT Model

</div>

Dataset overview

AetherSearch DPO contains 2,126 preference pairs for training an agentic-search policy after supervised fine-tuning. Every row provides one shared prompt, a preferred assistant continuation, and a non-preferred continuation.

This repository is deliberately train-only. It contains no benchmark results, run logs, model checkpoints, or held-out result files.

The released `muradil211/AetherSearch_DPO` checkpoint was trained in one DPO stage over all 2,126 pairs in this exact release, using the public `AetherSearch/dpo` training code on a separate server.

Release at a glance

ItemValue
Preference pairs2,126
Unique normalized questions2,126
Identical chosen/rejected pairs0
Splittrain only
OrderingDeterministic seeded hash shuffle
Shuffle seed42
Training filetrain.jsonl

Source composition

Source datasetPairsShare
TriviaQA1,44567.97%
MuSiQue41019.28%
Natural Questions1306.11%
WebQuestions874.09%
2WikiMultiHopQA542.54%
Total2,126100.00%

Preference composition

Pair typePairs
answer_hard_negative1,166
query_hard_negative289
true_full_trajectory_preference235
insufficient_information_continue_search173
evidence_misread_negative94
premature_answer_negative93
multi_hop_decomposition_negative30
regression_protection_pair28
query_refinement_negative18
Total2,126

Public schema

Every JSONL row contains exactly these eight fields, in this order:

FieldTypeMeaning
idstringCanonical six-digit public identifier
questionstringOriginal question
source_datasetstringQuestion-source dataset
answerslist[string]Accepted answer aliases
pair_typestringPreference/error category
prompt_textstringShared system, user, and trajectory context
chosenstringPreferred assistant continuation
rejectedstringNon-preferred assistant continuation

The DPO training unit is:

text
(prompt_text, chosen, rejected)

chosen and rejected are continuations and do not duplicate prompt_text. Some full-trajectory continuations contain environment-provided <information>...</information> spans. A trainer reproducing the AetherSearch loss contract should exclude those environment tokens from preference loss.

Curation and integrity

Before publication:

  1. 1.Questions overlapping the referenced Search-R1 train/test sets or the AetherSearch SFT release were removed using normalized-question matching.
  2. 2.Only one preference pair was retained per normalized question.
  3. 3.Internal paths, stage labels, rollout filenames, and private source IDs were removed from the public schema.
  4. 4.Rows were globally reordered with a deterministic seeded hash shuffle, then assigned IDs 000001 through 002126.
  5. 5.Three spacing/punctuation-only normalizations standardized factual entity designations; preference meaning and labels were unchanged.

The release audit verifies 2,126 unique IDs, 2,126 unique normalized questions, 2,126 non-empty chosen continuations, 2,126 non-empty rejected continuations, and zero identical pairs. Full distributions and character-length summaries are recorded in dataset_manifest.json.

Load with Datasets

python
from datasets import load_dataset

dataset = load_dataset(
    "muradil211/AetherSearch_DPO",
    split="train",
)

print(dataset.column_names)
print(len(dataset))  # 2126

Files

FilePurpose
train.jsonlCanonical preference-training split
dataset_manifest.jsonSchema, distributions, curation, and integrity audit
checksums.sha256Release-file integrity checksums
ATTRIBUTION.mdSource attribution and rights-status record
assets/aethersearch-mark.svgAetherSearch release mark

Limitations

  • —Preference labels reflect curated negative construction and trajectory selection; they are not human preference votes for every row.
  • —Retrieved passages can contain incomplete, outdated, or incorrect evidence.
  • —The dataset is intended for preference training and is not a benchmark.
  • —No downstream quality or safety result is claimed by this data release.

Rights and attribution

No blanket license is asserted for the combined release. Review ATTRIBUTION.md and the terms of every upstream source before redistribution or downstream publication.

Checksums

After downloading the repository, verify the release with:

bash
sha256sum -c checksums.sha256