Team Ai
Datasetpublic

allenai/multinews_sparse_oracle

This is a copy of the Multi-News dataset, except the input source documents of its test split have been replaced by a sparse retriever. The retrieval pipeline used: query: The summary field of each example corpus: The union of all documents in the train, validation and test splits retriever: BM25 via PyTerrier with default settings top-k strategy: "oracle", i.e. the number of documents retrieved, k, is set as the original number of input documents for each example Retrieval results on the… See the full description on the dataset page: https://huggingface.co/datasets/allenai/multinews_sparse_oracle.

sourceHugging Faceotherupdated 4y agoView on Hugging Face
1likes133downloads
Dataset Card

This is a copy of the Multi-News dataset, except the input source documents of its test split have been replaced by a _sparse_ retriever. The retrieval pipeline used:

  • —_query_: The summary field of each example
  • —_corpus_: The union of all documents in the train, validation and test splits
  • —_retriever_: BM25 via PyTerrier with default settings
  • —_top-k strategy_: "oracle", i.e. the number of documents retrieved, k, is set as the original number of input documents for each example

Retrieval results on the test set:

Recall@100RprecPrecision@kRecall@k
0.87750.74800.74800.7480