vector-dataset
hackernews-vector-search-datasetThe Hacker News dataset contains 28.74 million postings and their vector embeddings. The embeddings were generated using SentenceTransformers model all-MiniLM-L6-v2. The dimension of each embedding vector is 384.
Created by clickhouse more info: https://clickhouse.com/docs/getting-started/example-datasets/hackernews-vector-search-dataset
vector_lin3art_style_wn-datasetunbias-plus-dataset
Unbias Dataset
This dataset contains configurations used for the Unbias project at the Vector Institute:
train_4 (config, default): Our newest and highest quality training split.
other_splits (config): Contains the earlier splits below.
train_1: Training split sourced from VLDBench (regenerated version).
train_2: Another training split.
train_3: Another training split.
test_set: Test split sourced from BABE Golden 500.
⭐ train_4 is our newest, highest quality… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/unbias-plus-dataset.multi-vector-search-datasets
Multi-Vector Search Datasets
The datasets listed below are used in the Multi-Vector HNSW project for testing and benchmarking multi-vector approximate nearest neighbor search algorithms and their implementations.
Stack Exchange Datasets
Source: habedi/stack-exchange-dataset
Each row contains:
id: unique post ID
title: the post title
body: the main body content (with HTML tags removed)
tags: associated tags
embedding: a list of three 768-dimensional vectors for [title… See the full description on the dataset page: https://huggingface.co/datasets/habedi/multi-vector-search-datasets.resiplus-vector-dataset
ResiPlus Vector Dataset
Semantic search optimization dataset for medical documents. Expands queries with medical synonyms and generates Qdrant filters for document retrieval.
Dataset Details
Examples: 400
Language: Spanish (es)
Format: Chat messages (system, user, assistant)
Use Case: Fine-tuning LLMs for nursing home management system
Usage
from datasets import load_dataset
dataset = load_dataset("Alejandro284/resiplus-vector-dataset")
Training… See the full description on the dataset page: https://huggingface.co/datasets/Alejandro284/resiplus-vector-dataset.multi-vector-hnsw-datasets
Multi-Vector HNSW Benchmark Datasets
This repository contains benchmark datasets used by the Multi-Vector HNSW project.
The datasets are from habedi/multi-vector-search-datasets.
Each record includes a question ID and three distinct 768-dimensional vectors representing the title, body, and tags of the question.
The text embeddings were generated using the all-mpnet-base-v2 text embedding model.
There are three datasets; each includes questions from a separate Q&A community hosted on… See the full description on the dataset page: https://huggingface.co/datasets/habedi/multi-vector-hnsw-datasets.
