Team Ai
Datasetpublic

habedi/multi-vector-search-datasets

Multi-Vector Search Datasets The datasets listed below are used in the Multi-Vector HNSW project for testing and benchmarking multi-vector approximate nearest neighbor search algorithms and their implementations. Stack Exchange Datasets Source: habedi/stack-exchange-dataset Each row contains: id: unique post ID title: the post title body: the main body content (with HTML tags removed) tags: associated tags embedding: a list of three 768-dimensional vectors for… See the full description on the dataset page: https://huggingface.co/datasets/habedi/multi-vector-search-datasets.

sourceHugging Facecc-by-nc-4.0updated 1y agoView on Hugging Face
0likes25downloads
Dataset Card

Multi-Vector Search Datasets

The datasets listed below are used in the Multi-Vector HNSW project for testing and benchmarking multi-vector approximate nearest neighbor search algorithms and their implementations.


Stack Exchange Datasets

Source: habedi/stack-exchange-dataset

Each row contains:

  • —id: unique post ID
  • —title: the post title
  • —body: the main body content (with HTML tags removed)
  • —tags: associated tags
  • —embedding: a list of three 768-dimensional vectors for [title, body, tags]

Text embeddings were generated using `all-mpnet-base-v2` from Sentence Transformers.

IndexDatasetSize
1Computer Science40,792
2Data Science28,950
3Political Science12,416

Flickr8k Dataset

Source: habedi/flickr-8k-dataset-clean

Each row contains:

  • —id: image filename
  • —captions: a list of five human-written captions
  • —image: raw image data (JPEG format)
  • —embedding: a list of six 768-dimensional vectors: five for captions, one for the image

Caption embeddings were generated using `all-mpnet-base-v2`.

Image embeddings were generated using `openai/clip-vit-base-patch32` via the Hugging Face Transformers library.

IndexDatasetSize
1Flickr8k (captions + image)8,091