Team Ai
Datasetpublic

habedi/multi-vector-hnsw-datasets

Multi-Vector HNSW Benchmark Datasets This repository contains benchmark datasets used by the Multi-Vector HNSW project. The datasets are from habedi/multi-vector-search-datasets. Each record includes a question ID and three distinct 768-dimensional vectors representing the title, body, and tags of the question. The text embeddings were generated using the all-mpnet-base-v2 text embedding model. There are three datasets; each includes questions from a separate Q&A community… See the full description on the dataset page: https://huggingface.co/datasets/habedi/multi-vector-hnsw-datasets.

sourceHugging Facecc-by-nc-4.0updated 1y agoView on Hugging Face
0likes18downloads
Dataset Card

Multi-Vector HNSW Benchmark Datasets

This repository contains benchmark datasets used by the Multi-Vector HNSW project. The datasets are from habedi/multi-vector-search-datasets. Each record includes a question ID and three distinct 768-dimensional vectors representing the title, body, and tags of the question. The text embeddings were generated using the all-mpnet-base-v2 text embedding model.

There are three datasets; each includes questions from a separate Q\&A community hosted on Stack Exchange. Each dataset consists of three files:

  • —train.json: The data used to build the index.
  • —test.json: The query data used for searching the index.
  • —neighbours.json: The ground truth, containing the actual k=100 nearest neighbors for each item in test.json.
#DatasetQ\&A CommunityNum VectorsDimensionsTrain SizeTest Size
1se_cs_768Computer Science376836,7124,080
2se_ds_768Data Science376826,0552,895
3se_p_768Political Science376811,1741,242