vector
creative-writing-control-vectors-v3.0Llama-3-70B-japanese-suzume-vector-v0.1CantoneseLLM-v2.0-8B-Thinking-Chat-Vector-Merged-i1-GGUFCantoneseLLM-v2.0-30B-A3B-Thinking-Chat-Vector-Merged-i1-GGUFfasttext-en-vectorsDelta-Vector_Austral-24B-Winton-GGUFflux-lora-gliff-tosti-vector-1flux-lora-tosti-vector-full-captions
Datasets
All datasets matching “vector”whisper_transcriptions.reazon_speech_all.wer_10.0.vectorizedwhisper_transcriptions.mls.wer_10.0.vectorizedemotion-vectors-gemma-4-31b-it-postfix
Emotion vectors, google/gemma-4-31b-it (corrected extraction)
Residual-stream activations for google/gemma-4-31b-it, pooled per story and averaged per
emotion. Each emotion ends up as one direction in the model's activation space.
Read LINEAGE.md before using this. This set supersedes
abotresol/emotion-vectors-gemma-4-31b-it. The earlier extraction ran
while the tokenizer padded on the left, so the step that skips a story's first
50 tokens skipped padding instead. This set… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-it-postfix.Wikidata_Vectors_0.2
Wikidata Entity Embeddings 0.2
Dataset Summary
Wikidata Entity Embeddings is a dataset of embedding vectors for Wikidata entities. Each vector represents a Wikidata item (Q...) or property (P...) based on textual information extracted from Wikidata.
The dataset is part of the Wikidata Embedding Project, an initiative led by Wikimedia Deutschland in collaboration with Jina AI and IBM DataStax. The project provides a publicly accessible Wikidata Vector Database to… See the full description on the dataset page: https://huggingface.co/datasets/philippesaade/Wikidata_Vectors_0.2.vibeThis repository contains the datasets presented in VIBE: Vector Index Benchmark for Embeddings:
https://github.com/vector-index-bench/vibe
The datasets can be downloaded manually from this repository, but the benchmark framework also downloads them automatically.
Datasets
In-distribution datasets
Name
Type
n
d
Distance
agnews-mxbai-1024-euclidean
Text
769,382
1024
euclidean
arxiv-nomic-768-normalized
Text
1,344,643
768
any
dpr-jina-768-normalized… See the full description on the dataset page: https://huggingface.co/datasets/vector-index-bench/vibe.open-pmc-18m
OPEN-PMC
Arxiv: Arxiv
|
Code: Open-PMC Github
|
Model Checkpoint: Hugging Face
Dataset Summary
This dataset consists of image-text pairs extracted from medical papers available on PubMed Central. It has been curated to support research in medical image understanding, particularly in natural language processing (NLP) and computer vision tasks related to medical imagery. The dataset includes:
Extracted images from research articles.… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/open-pmc-18m.
