datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
japanese-splade-v1-hard-negatives日本語 SPLADE v2 の学習に用いたデータセットです。
SPLADE モデルである hotchpotch/japanese-splade-base-v1-mmarco-only, japanese-splade-base-v1_5 を用いてハードネガティブマイニングを行なっています。また BAAI/bge-reranker-v2-m3 を用いたリランカースコアを付与しています。
mqa, mmarco はhpprc/emb のデータを用いています。
mqa の query 作成には MinHash を用い約40万件になるようフィルタしました。
msmarco-ja は hpprc/msmarco-jaのデータを用いています。
ライセンスは、各データセットのライセンスを継承します。
dbpedia-entities-efficient-splade-100K
DBPedia SPLADE + OpenAI: 100,000 SPLADE Sparse Vectors + OpenAI Embedding
This dataset has both OpenAI and SPLADE vectors for 100,000 DBPedia entries. This adds SPLADE Vectors to KShivendu/dbpedia-entities-openai-1M/
Model id used to make these vectors:
model_id = "naver/efficient-splade-VI-BT-large-doc"
For processing the query, use this:
model_id = "naver/efficient-splade-VI-BT-large-query"
If you'd like to extract the indices and weights/values from the vectors, you can do so… See the full description on the dataset page: https://huggingface.co/datasets/nirantk/dbpedia-entities-efficient-splade-100K.msmarco_v2.1_segmented-spladev3-anserinibrowsecomp-plus-spladev3-anserinimsmarco-passage-v2.splade-lg.pisa
msmarco-passage-v2.splade-lg.pisa
Description
TODO: What is the artifact?
Usage
# Load the artifact
import pyterrier as pt
artifact = pt.Artifact.from_hf('pyterrier/msmarco-passage-v2.splade-lg.pisa')
# TODO: Show how you use the artifact
Benchmarks
TODO: Provide benchmarks for the artifact.
Reproduction
# TODO: Show how you constructed the artifact.
Metadata
{
"type": "sparse_index",
"format":… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/msmarco-passage-v2.splade-lg.pisa.wikipedia-en-splade-bge
Wikipedia English with SPLADE and BGE-M3
Pre-computed SPLADE sparse and BGE-M3 dense embeddings for 6.4M English Wikipedia articles.
Direct Usage
HuggingFace automatically discovers parquet files. You can load this dataset directly:
from datasets import load_dataset
# Stream the entire dataset (recommended for large dataset)
dataset = load_dataset("Sicheng-Chroma/wikipedia-en-splade-bge", streaming=True)
# Load specific splits
train_dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Sicheng-Chroma/wikipedia-en-splade-bge.ragwiki-splade-cocondenser-ensembledistil.pisa
ragwiki-splade-cocondenser-ensembledistil.pisa
Description
TODO: What is the artifact?
Usage
# Load the artifact
import pyterrier as pt
artifact = pt.Artifact.from_hf('DandyTian/ragwiki-splade-cocondenser-ensembledistil.pisa')
# TODO: Show how you use the artifact
Benchmarks
TODO: Provide benchmarks for the artifact.
Reproduction
# TODO: Show how you constructed the artifact.
Metadata
{
"type":… See the full description on the dataset page: https://huggingface.co/datasets/DandyTian/ragwiki-splade-cocondenser-ensembledistil.pisa.seismic-msmarco-spladekannolo-msmarco-spladeqrecc-distillation-mistral-spladedbpedia-entity.splade-v3.cache
dbpedia-entity.splade-v3.cache
Description
TODO: What is the artifact?
Usage
# Load the artifact
import pyterrier_alpha as pta
artifact = pta.Artifact.from_hf('pyterrier/dbpedia-entity.splade-v3.cache')
# TODO: Show how you use the artifact
Benchmarks
TODO: Provide benchmarks for the artifact.
Reproduction
# TODO: Show how you constructed the artifact.
Metadata
{
"type": "indexer_cache",
"format": "lz4pickle"… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/dbpedia-entity.splade-v3.cache.topiocqa-distillation-llama-mistral-splade
DiSCo: LLM Knowledge Distillation for Efficient Sparse Retrieval in Conversational Search
This distillation dataset can be used to distill knowledge from Llama and Mistral LLMs into a conversational search dataset (TopiOCQA).
Please refer to the DiSCo github for complete usage [github].
Citation
If you use our distillation file, please cite our work:
@article{lupart2024disco,
title={DiSCo Meets LLMs: A Unified Approach for Sparse Retrieval and Contextual Distillation in… See the full description on the dataset page: https://huggingface.co/datasets/slupart/topiocqa-distillation-llama-mistral-splade.msmarco-train-distil-splade-v3-similaritybright.sustainable.splade.cache
bright.sustainable.splade.cache
Description
TODO: What is the artifact?
Usage
# Load the artifact
import pyterrier as pt
artifact = pt.Artifact.from_hf('pyterrier-tutorial/bright.sustainable.splade.cache')
# TODO: Show how you use the artifact
Benchmarks
TODO: Provide benchmarks for the artifact.
Reproduction
# TODO: Show how you constructed the artifact.
Metadata
{
"type": "indexer_cache"… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier-tutorial/bright.sustainable.splade.cache.hotpotqa.splade-v3.cache
hotpotqa.splade-v3.cache
Description
TODO: What is the artifact?
Usage
# Load the artifact
import pyterrier_alpha as pta
artifact = pta.Artifact.from_hf('pyterrier/hotpotqa.splade-v3.cache')
# TODO: Show how you use the artifact
Benchmarks
TODO: Provide benchmarks for the artifact.
Reproduction
# TODO: Show how you constructed the artifact.
Metadata
{
"type": "indexer_cache",
"format": "lz4pickle",
"package_hint":… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/hotpotqa.splade-v3.cache.msmarco-passage.splade-lg.cache
msmarco-passage.splade-lg.cache
Description
TODO: What is the artifact?
Usage
# Load the artifact
import pyterrier as pt
artifact = pt.Artifact.from_hf('macavaney/msmarco-passage.splade-lg.cache')
# TODO: Show how you use the artifact
Benchmarks
TODO: Provide benchmarks for the artifact.
Reproduction
# TODO: Show how you constructed the artifact.
Metadata
{
"type": "indexer_cache",
"format": "lz4pickle",
"record_count":… See the full description on the dataset page: https://huggingface.co/datasets/macavaney/msmarco-passage.splade-lg.cache.climate-fever.splade-v1.cache
climate-fever.splade-v1.cache
Description
TODO: What is the artifact?
Usage
# Load the artifact
import pyterrier_alpha as pta
artifact = pta.Artifact.from_hf('pyterrier/climate-fever.splade-v1.cache')
# TODO: Show how you use the artifact
Benchmarks
TODO: Provide benchmarks for the artifact.
Reproduction
# TODO: Show how you constructed the artifact.
Metadata
{
"type": "indexer_cache",
"format": "lz4pickle"… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/climate-fever.splade-v1.cache.msmarco-spladev2-hard-negatives-scoresmsmarco-train-distil-splade-v3enwiki_20251001_spladepp_indexsplade-v3-ciffbright.sustainable.splade.pisa
bright.sustainable.splade.pisa
Description
TODO: What is the artifact?
Usage
# Load the artifact
import pyterrier as pt
artifact = pt.Artifact.from_hf('pyterrier-tutorial/bright.sustainable.splade.pisa')
# TODO: Show how you use the artifact
Benchmarks
TODO: Provide benchmarks for the artifact.
Reproduction
# TODO: Show how you constructed the artifact.
Metadata
{
"type": "sparse_index",
"format":… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier-tutorial/bright.sustainable.splade.pisa.nq.splade-v3.cache
nq.splade-v3.cache
Description
TODO: What is the artifact?
Usage
# Load the artifact
import pyterrier_alpha as pta
artifact = pta.Artifact.from_hf('pyterrier/nq.splade-v3.cache')
# TODO: Show how you use the artifact
Benchmarks
TODO: Provide benchmarks for the artifact.
Reproduction
# TODO: Show how you constructed the artifact.
Metadata
{
"type": "indexer_cache",
"format": "lz4pickle",
"package_hint":… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/nq.splade-v3.cache.seismic-msmarco-splade-bineti-embedding-splade-v7-distill-v1qrecc-distillation-human-mistral-splademsmarco_passage_trec_dl_2019_judged_splade_naver_splade_v2_distil.terrier
msmarco_passage_trec_dl_2019_judged_splade_naver_splade_v2_distil.terrier
Description
TODO: What is the artifact?
Usage
# Load the artifact
import pyterrier as pt
artifact = pt.Artifact.from_hf('JackMcKechnie/msmarco_passage_trec_dl_2019_judged_splade_naver_splade_v2_distil.terrier')
# TODO: Show how you use the artifact
Benchmarks
TODO: Provide benchmarks for the artifact.
Reproduction
# TODO: Show how you constructed the… See the full description on the dataset page: https://huggingface.co/datasets/JackMcKechnie/msmarco_passage_trec_dl_2019_judged_splade_naver_splade_v2_distil.terrier.msmarco_passage_trec_dl_2019_judged_splade_naver_splade_cocondenser_selfdistil.terrier
msmarco_passage_trec_dl_2019_judged_splade_naver_splade_cocondenser_selfdistil.terrier
Description
TODO: What is the artifact?
Usage
# Load the artifact
import pyterrier as pt
artifact = pt.Artifact.from_hf('JackMcKechnie/msmarco_passage_trec_dl_2019_judged_splade_naver_splade_cocondenser_selfdistil.terrier')
# TODO: Show how you use the artifact
Benchmarks
TODO: Provide benchmarks for the artifact.
Reproduction
# TODO: Show how… See the full description on the dataset page: https://huggingface.co/datasets/JackMcKechnie/msmarco_passage_trec_dl_2019_judged_splade_naver_splade_cocondenser_selfdistil.terrier.climate-fever.splade-v3.cache
climate-fever.splade-v3.cache
Description
TODO: What is the artifact?
Usage
# Load the artifact
import pyterrier_alpha as pta
artifact = pta.Artifact.from_hf('pyterrier/climate-fever.splade-v3.cache')
# TODO: Show how you use the artifact
Benchmarks
TODO: Provide benchmarks for the artifact.
Reproduction
# TODO: Show how you constructed the artifact.
Metadata
{
"type": "indexer_cache",
"format": "lz4pickle"… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/climate-fever.splade-v3.cache.hotpotqa.splade-v1.cache
hotpotqa.splade-v1.cache
Description
TODO: What is the artifact?
Usage
# Load the artifact
import pyterrier_alpha as pta
artifact = pta.Artifact.from_hf('pyterrier/hotpotqa.splade-v1.cache')
# TODO: Show how you use the artifact
Benchmarks
TODO: Provide benchmarks for the artifact.
Reproduction
# TODO: Show how you constructed the artifact.
Metadata
{
"type": "indexer_cache",
"format": "lz4pickle",
"package_hint":… See the full description on the dataset page: https://huggingface.co/datasets/pyterrier/hotpotqa.splade-v1.cache.
