AdarshSingh7647/Eklav-Reranker-AnswerOnly-Data
Eklav-Reranker-AnswerOnly-Data Training data for the Eklav paper. Task: passage reranking (BRIGHT / NevIR benchmarks) Method: Answer-only (no reasoning of any kind -- the no-CoT floor) Examples: 381,934 train / 3,857 held-out val Format: ShareGPT (system + conversations: [{from, value}]), used for LoRA SFT via LLaMA-Factory. Single-turn ShareGPT conversations. Each row: a query+passage relevance-judgment prompt (human turn) and a bare true/false judgment (gpt turn) -- no hint… See the full description on the dataset page: https://huggingface.co/datasets/AdarshSingh7647/Eklav-Reranker-AnswerOnly-Data.
Eklav-Reranker-AnswerOnly-Data
Training data for the Eklav paper.
- Task: passage reranking (BRIGHT / NevIR benchmarks)
- Method: Answer-only (no reasoning of any kind -- the no-CoT floor)
- Examples: 381,934 train / 3,857 held-out val
- Format: ShareGPT (
system+conversations: [{from, value}]), used for LoRA SFT via LLaMA-Factory.
Single-turn ShareGPT conversations. Each row: a query+passage relevance-judgment prompt (human turn) and a bare true/false judgment (gpt turn) -- no hint, no chain-of-thought, no <think> block anywhere. Loss is computed over the entire (one-token) gpt turn.
Built from the same source query/passage/judgment triples as Eklav-Reranker-Data (Eklav method) and Eklav-Reranker-CotGen-Data (std-SFT method), with the teacher reasoning removed entirely rather than hinted at or reproduced -- the train/val split (381,934 / 3,857, seed 42) matches those two repos exactly so all three methods are directly comparable example-for-example.
Files:
train.json-- training splitval.json-- held-out validation split
See the Eklav paper for full dataset construction methodology, and the corresponding Eklav-*-PassageReranking-AnswerOnly model repos for checkpoints trained on this data.
