Team Ai
Datasetpublic

lightonai/embeddings-fine-tuning

Overview This dataset is composed of high quality data sources with mined hard negatives. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using this dataset or its curated version. This dataset has originally been created to follow the nv-retrieve setup, that mines the closest negatives to the query in a dataset and filter false negatives if their bi-encoder similarity is higher than a… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning.

sourceHugging Faceupdated 3mo agoView on Hugging Face
26likes4.4kdownloads
Dataset Card

Overview

This dataset is composed of high quality data sources with mined hard negatives. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using this dataset or its curated version. This dataset has originally been created to follow the nv-retrieve setup, that mines the closest negatives to the query in a dataset and filter false negatives if their bi-encoder similarity is higher than a percentage of the query-positive similarity score. To allow the exploration of various threshold and sampling methods, we decided, as for our pre-training datasets, to be the least destructive possible. Thus, instead of giving the final filtered samples given a method/threshold, we share all of the data, including all the (2048) mined negatives alongside their scores so anyone can apply their own strategy before training easily. The mined datasets are FiQa, NaturalQuestion, HotpotQA, MSMARCO, FEVER, SquadV2 and TriviaQA, for a total of 1.88M queries with 2048 mined negatives and their scores, alongside the positive. The model used for mining is gte-modernbert-base

For more information, please refer to our blogpost.

How to use

If you want to directly use the data as contrastive data with nv-retrieve filtering in either sentence-transformers or PyLate, you can simply map it to the (query, positive, negative_1, negative_2, ..., negative_n) like so:

<details> <summary> Python code to cast to contrastive format </summary>

python
import datasets
import os 
class KDToContrastive:
    """Dataset processing class for converting a KD dataset into a contrastive one.

    Parameters
    ----------
    queries
        Queries dataset.
    documents
        Documents dataset.
    split
        Split to use for the queries and documents datasets. Used only if the queries and documents are of type `datasets.DatasetDict`.
    num_negatives
        Number of negatives to keep.
    nv_threshold
        Threshold for the nv-embed filtering
    """

    def __init__(
        self,
        queries: datasets.Dataset | datasets.DatasetDict,
        documents: datasets.Dataset | datasets.DatasetDict,
        split: str = "train",
        num_negatives: int = 32,
        nv_threshold: float = 0.95,
    ) -> None:
        if isinstance(queries, datasets.DatasetDict):
            self.queries = queries[split]
        else:
            self.queries = queries

        if isinstance(documents, datasets.DatasetDict):
            self.documents = documents[split]
        else:
            self.documents = documents

        self.num_negatives = num_negatives
        self.nv_threshold = nv_threshold

        self.queries_index = {
            query_id: i for i, query_id in enumerate(iterable=self.queries["query_id"])
        }

        self.documents_index = {
            document_id: i
            for i, document_id in enumerate(iterable=self.documents["document_id"])
        }

    def has_enough_negatives(self, example):
        """Check if example has at least 50 valid negatives"""
        scores = example["scores"]
        positive_score = scores[0]

        count = sum(
            1 for score in scores[1:] if score < self.nv_threshold * positive_score
        )
        return count >= self.num_negatives

    def map_to_query_positive_negatives(self, example):
        """
        Maps a scores example to the desired format:
        query, positive, negative_0, negative_1, ..., negative_49
        """
        query_id = example["query_id"]
        document_ids = example["document_ids"]
        scores = example["scores"]

        # Get query text
        query_text = self.queries[self.queries_index[query_id]]

        # First document_id is the positive
        positive_id = document_ids[0]

        positive_text = self.documents[self.documents_index[positive_id]]
        positive_score = scores[0]

        # Create the row
        row = {"query": query_text, "positive": positive_text}

        # Add negatives (starting from index 1)
        total_negatives = 0
        for i in range(1, len(document_ids)):
            if scores[i] < self.nv_threshold * positive_score:
                negative_id = document_ids[i]
                row[f"negative_{total_negatives}"] = self.documents[
                    self.documents_index[negative_id]
                ]
                total_negatives += 1
                if total_negatives >= self.num_negatives:
                    break

        return row


def load_train_datasets():
    """Load all available splits from raphael data, with caching"""
    cache_dir = "nv_retrieve_99_50_cached"
    os.makedirs(cache_dir, exist_ok=True)
    train_dataset = datasets.DatasetDict()
    splits = ["trivia", "hotpotqa", "nq", "msmarco", "fever", "squadv2", "fiqa"]

    for split in splits:
        try:
            dataset = datasets.Dataset.load_from_disk(f"{cache_dir}/{split}")
            print("Loaded dataset from disk")
        except FileNotFoundError:
            print("Creating dataset")
            dataset = datasets.load_dataset(
                "lightonai/nv-embed-supervised-distill-dedup",
                name="scores",
                num_proc=144,
                split=split,
            )
            queries = datasets.load_dataset(
                "lightonai/nv-embed-supervised-distill-dedup",
                name="queries",
                num_proc=144,
                split=split,
            )
            documents = datasets.load_dataset(
                "lightonai/nv-embed-supervised-distill-dedup",
                name="documents",
                num_proc=144,
                split=split,
            )
            processor = KDToContrastive(
                queries, documents, num_negatives=50, nv_threshold=0.99
            )
            dataset = dataset.filter(
                processor.has_enough_negatives,
                desc="Filtering examples with <50 negatives",
            ).map(
                processor.map_to_query_positive_negatives,
                remove_columns=dataset.column_names,
                desc="Creating query-positive-negatives dataset",
            )
            dataset.save_to_disk(f"{cache_dir}/{split}")

        train_dataset[split] = dataset
    return train_dataset

</details>

Dataset structure

The dataset is composed of 7 high quality datasets, defined by the splits parameters. Each split contains 3 subsets, one containing the queries, one containing the documents and one joining tables also containing the corresponding pairwise query-documents scores.

Documents

ColumnTypeDescription
document_idint64Unique identifier of the document within the split.
documentstringRaw text of the document/passage.
SplitRows
fiqa57.6k
nq10.1M
hotpotqa5.22M
msmarco8.84M
fever5.38M
squadv219k
trivia21M
Total50.64M

Queries

ColumnTypeDescription
query_idint64Unique identifier of the query within the split.
querystringRaw text of the query.
SplitRows
fiqa5.5k
nq307k
hotpotqa85k
msmarco503k
fever110k
squadv2130k
trivia78.8k
Total1.22M

Scores

ColumnTypeDescription
query_idint64Identifier joining back to the corresponding row in queries.
document_idslist[int64]List of document IDs (joining back to documents). The first element is the positive document, followed by the top-2048 mined for the query.
scoreslist[float]Relevance scores for each document w.r.t the query. The first element is the positive document, followed by the top-2048 mined for the query. Can be used for nv-retrieve filtering or knowledge distillation.
SplitRows
fiqa14.2k
hotpotqa170k
nq152k
msmarco533k
fever140k
squadv2130k
trivia741k
Total1.88M

Citation

If you are using this dataset, please consider citing our work

bibtex
@misc{sourty2025denseonlateon,
  title={DenseOn with LateOn: Open State-of-the-Art Single and Multi-Vector Models},
  author={Sourty, Raphael and Chaffin, Antoine and Weller, Orion and Demoura, Paulo and Chatelain, Amelie},
  year={2026},
  howpublished={\url{https://huggingface.co/blog/lightonai/denseon-lateon}},
}```