Team Ai
Datasetpublic

allenai/multixscience_sparse_max

This is a copy of the Multi-XScience dataset, except the input source documents of its test split have been replaced by a sparse retriever. The retrieval pipeline used: query: The related_work field of each example corpus: The union of all documents in the train, validation and test splits retriever: BM25 via PyTerrier with default settings top-k strategy: "max", i.e. the number of documents retrieved, k, is set as the maximum number of documents seen across examples in this dataset, in this… See the full description on the dataset page: https://huggingface.co/datasets/allenai/multixscience_sparse_max.

sourceHugging Faceunknownupdated 4y agoView on Hugging Face
0likes36downloads
settings

This repository belongs to allenai on Hugging Face.

Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namemultixscience_sparse_max
visibilitypublic
licenceunknown
gatedno
ownerallenai
Account settings
allenai/multixscience_sparse_max · Team Ai