muradil211/AetherSearch_DPO
🔭 AetherSearch DPO Preference pairs for reasoning, retrieval, and evidence-grounded answers 🏠 Project · 🧠 DPO Model · 🧪 Training Code · 🎓 SFT Data · 🤖 SFT Model Dataset overview AetherSearch DPO contains 2,126 preference pairs for training an agentic-search policy after supervised fine-tuning. Every row provides one shared prompt, a preferred assistant continuation, and a non-preferred continuation. This repository is… See the full description on the dataset page: https://huggingface.co/datasets/muradil211/AetherSearch_DPO.
<div align="center">
<img src="assets/aethersearch-mark.svg" alt="AetherSearch monogram" width="140">
🔭 AetherSearch DPO
Preference pairs for reasoning, retrieval, and evidence-grounded answers
<p> <img src="https://img.shields.io/badge/Pairs-2%2C126-F59E0B?style=flat-square" alt="2,126 preference pairs"> <img src="https://img.shields.io/badge/Split-Train%20Only-0F766E?style=flat-square" alt="Train split only"> <img src="https://img.shields.io/badge/Questions-Unique-2563EB?style=flat-square" alt="Unique questions"> <img src="https://img.shields.io/badge/Format-Chosen%20%2F%20Rejected-7C3AED?style=flat-square" alt="Chosen and rejected format"> </p>
🏠 Project · 🧠 DPO Model · 🧪 Training Code · 🎓 SFT Data · 🤖 SFT Model
</div>
Dataset overview
AetherSearch DPO contains 2,126 preference pairs for training an agentic-search policy after supervised fine-tuning. Every row provides one shared prompt, a preferred assistant continuation, and a non-preferred continuation.
This repository is deliberately train-only. It contains no benchmark results, run logs, model checkpoints, or held-out result files.
The released `muradil211/AetherSearch_DPO` checkpoint was trained in one DPO stage over all 2,126 pairs in this exact release, using the public `AetherSearch/dpo` training code on a separate server.
Release at a glance
Source composition
Preference composition
Public schema
Every JSONL row contains exactly these eight fields, in this order:
The DPO training unit is:
(prompt_text, chosen, rejected)chosen and rejected are continuations and do not duplicate prompt_text. Some full-trajectory continuations contain environment-provided <information>...</information> spans. A trainer reproducing the AetherSearch loss contract should exclude those environment tokens from preference loss.
Curation and integrity
Before publication:
- Questions overlapping the referenced Search-R1 train/test sets or the AetherSearch SFT release were removed using normalized-question matching.
- Only one preference pair was retained per normalized question.
- Internal paths, stage labels, rollout filenames, and private source IDs were removed from the public schema.
- Rows were globally reordered with a deterministic seeded hash shuffle, then assigned IDs
000001through002126. - Three spacing/punctuation-only normalizations standardized factual entity designations; preference meaning and labels were unchanged.
The release audit verifies 2,126 unique IDs, 2,126 unique normalized questions, 2,126 non-empty chosen continuations, 2,126 non-empty rejected continuations, and zero identical pairs. Full distributions and character-length summaries are recorded in dataset_manifest.json.
Load with Datasets
from datasets import load_dataset
dataset = load_dataset(
"muradil211/AetherSearch_DPO",
split="train",
)
print(dataset.column_names)
print(len(dataset)) # 2126Files
Limitations
- Preference labels reflect curated negative construction and trajectory selection; they are not human preference votes for every row.
- Retrieved passages can contain incomplete, outdated, or incorrect evidence.
- The dataset is intended for preference training and is not a benchmark.
- No downstream quality or safety result is claimed by this data release.
Rights and attribution
No blanket license is asserted for the combined release. Review ATTRIBUTION.md and the terms of every upstream source before redistribution or downstream publication.
Checksums
After downloading the repository, verify the release with:
sha256sum -c checksums.sha256