Team Ai
Datasetpublic

nirantk/dbpedia-entities-efficient-splade-100K

DBPedia SPLADE + OpenAI: 100,000 SPLADE Sparse Vectors + OpenAI Embedding This dataset has both OpenAI and SPLADE vectors for 100,000 DBPedia entries. This adds SPLADE Vectors to KShivendu/dbpedia-entities-openai-1M/ Model id used to make these vectors: model_id = "naver/efficient-splade-VI-BT-large-doc" For processing the query, use this: model_id = "naver/efficient-splade-VI-BT-large-query" If you'd like to extract the indices and weights/values from the vectors, you can… See the full description on the dataset page: https://huggingface.co/datasets/nirantk/dbpedia-entities-efficient-splade-100K.

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
3likes153downloads
Dataset Card

DBPedia SPLADE + OpenAI: 100,000 SPLADE Sparse Vectors + OpenAI Embedding

This dataset has both OpenAI and SPLADE vectors for 100,000 DBPedia entries. This adds SPLADE Vectors to KShivendu/dbpedia-entities-openai-1M/

Model id used to make these vectors:

python
model_id = "naver/efficient-splade-VI-BT-large-doc"

For processing the query, use this:

python
model_id = "naver/efficient-splade-VI-BT-large-query"

If you'd like to extract the indices and weights/values from the vectors, you can do so using the following snippet:

python
import numpy as np
vec = np.array(ds[0]['vec']) # where ds is the dataset

def get_indices_values(vec):
  sparse_indices = vec.nonzero()
  sparse_values = vec[sparse_indices]
  return sparse_indices, sparse_values