habedi/multi-vector-search-datasets
Multi-Vector Search Datasets The datasets listed below are used in the Multi-Vector HNSW project for testing and benchmarking multi-vector approximate nearest neighbor search algorithms and their implementations. Stack Exchange Datasets Source: habedi/stack-exchange-dataset Each row contains: id: unique post ID title: the post title body: the main body content (with HTML tags removed) tags: associated tags embedding: a list of three 768-dimensional vectors for… See the full description on the dataset page: https://huggingface.co/datasets/habedi/multi-vector-search-datasets.
Multi-Vector Search Datasets
The datasets listed below are used in the Multi-Vector HNSW project for testing and benchmarking multi-vector approximate nearest neighbor search algorithms and their implementations.
Stack Exchange Datasets
Source: habedi/stack-exchange-dataset
Each row contains:
id: unique post IDtitle: the post titlebody: the main body content (with HTML tags removed)tags: associated tagsembedding: a list of three 768-dimensional vectors for[title, body, tags]
Text embeddings were generated using `all-mpnet-base-v2` from Sentence Transformers.
Flickr8k Dataset
Source: habedi/flickr-8k-dataset-clean
Each row contains:
id: image filenamecaptions: a list of five human-written captionsimage: raw image data (JPEG format)embedding: a list of six 768-dimensional vectors: five for captions, one for the image
Caption embeddings were generated using `all-mpnet-base-v2`.
Image embeddings were generated using `openai/clip-vit-base-patch32` via the Hugging Face Transformers library.
