Team Ai
Datasetpublic

allenai/multixscience_sparse_max

This is a copy of the Multi-XScience dataset, except the input source documents of its test split have been replaced by a sparse retriever. The retrieval pipeline used: query: The related_work field of each example corpus: The union of all documents in the train, validation and test splits retriever: BM25 via PyTerrier with default settings top-k strategy: "max", i.e. the number of documents retrieved, k, is set as the maximum number of documents seen across examples in this dataset, in this… See the full description on the dataset page: https://huggingface.co/datasets/allenai/multixscience_sparse_max.

sourceHugging Faceunknownupdated 4y agoView on Hugging Face
0likes36downloads
9 commits on main
59efc384y ago

Update README.md

johngiorgi
6e2f2b44y ago

Update README.md

johngiorgi
7f3fadb4y ago

Update README.md

johngiorgi
c828b0f4y ago

Create README.md

johngiorgi
708a7de4y ago

Upload dataset_infos.json with huggingface_hub

johngiorgi
3bf81e24y ago

Upload data/validation-00000-of-00001-666f8f42dff99a14.parquet with huggingface_hub

johngiorgi
22725264y ago

Upload data/test-00000-of-00001-f586e99c0013eedb.parquet with huggingface_hub

johngiorgi
67e27774y ago

Upload data/train-00000-of-00001-e71278f311282096.parquet with huggingface_hub

johngiorgi
7350b414y ago

initial commit

johngiorgi