datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
whisper_transcriptions.reazon_speech_all.wer_10.0.vectorizedwhisper_transcriptions.mls.wer_10.0.vectorizedemotion-vectors-gemma-4-31b-it-postfix
Emotion vectors, google/gemma-4-31b-it (corrected extraction)
Residual-stream activations for google/gemma-4-31b-it, pooled per story and averaged per
emotion. Each emotion ends up as one direction in the model's activation space.
Read LINEAGE.md before using this. This set supersedes
abotresol/emotion-vectors-gemma-4-31b-it. The earlier extraction ran
while the tokenizer padded on the left, so the step that skips a story's first
50 tokens skipped padding instead. This set… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-it-postfix.Wikidata_Vectors_0.2
Wikidata Entity Embeddings 0.2
Dataset Summary
Wikidata Entity Embeddings is a dataset of embedding vectors for Wikidata entities. Each vector represents a Wikidata item (Q...) or property (P...) based on textual information extracted from Wikidata.
The dataset is part of the Wikidata Embedding Project, an initiative led by Wikimedia Deutschland in collaboration with Jina AI and IBM DataStax. The project provides a publicly accessible Wikidata Vector Database to… See the full description on the dataset page: https://huggingface.co/datasets/philippesaade/Wikidata_Vectors_0.2.vibeThis repository contains the datasets presented in VIBE: Vector Index Benchmark for Embeddings:
https://github.com/vector-index-bench/vibe
The datasets can be downloaded manually from this repository, but the benchmark framework also downloads them automatically.
Datasets
In-distribution datasets
Name
Type
n
d
Distance
agnews-mxbai-1024-euclidean
Text
769,382
1024
euclidean
arxiv-nomic-768-normalized
Text
1,344,643
768
any
dpr-jina-768-normalized… See the full description on the dataset page: https://huggingface.co/datasets/vector-index-bench/vibe.open-pmc-18m
OPEN-PMC
Arxiv: Arxiv
|
Code: Open-PMC Github
|
Model Checkpoint: Hugging Face
Dataset Summary
This dataset consists of image-text pairs extracted from medical papers available on PubMed Central. It has been curated to support research in medical image understanding, particularly in natural language processing (NLP) and computer vision tasks related to medical imagery. The dataset includes:
Extracted images from research articles.… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/open-pmc-18m.emotion-vectors-gemma-4-31b
Emotion vectors — gemma-4-31b (base) probed on the external gemma-4-4B story corpus
Data provenance (what made these activations)
Probed model (whose activations these are): google/gemma-4-31b (base)
Input corpus: snae/emotion_stories_gemma_4_4B — third-person emotion stories written by gemma-4-4B, a smaller EXTERNAL model (the open replication's published corpus; generator is NOT the probed model)
Per-story pooled residual-stream activations and per-emotion… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b.v32-vectorsVectorSQLBenchassistant-axis-vectors
The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models
This repository contains pre-computed axes and persona vectors for Gemma 2 27B, Qwen 3 32B, and Llama 3.3 70B, as described in the paper The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models.
Paper | Code | Demo
The Assistant Axis is a direction in activation space that captures how "Assistant-like" a model's current persona is. It can be used to:
Monitor persona… See the full description on the dataset page: https://huggingface.co/datasets/lu-christina/assistant-axis-vectors.emotion-vectors-gemma-4-31b-it
Emotion vectors — gemma-4-31b-it (instruct) probed on the external gemma-4-4B story corpus
Data provenance (what made these activations)
Probed model: google/gemma-4-31b-it (instruct)
Input corpus: snae/emotion_stories_gemma_4_4B — stories written by gemma-4-4B, a smaller EXTERNAL model (generator is NOT the probed model)
Per-story pooled residual-stream activations and per-emotion mean vectors,
extracted with gemma4-emotion-vectors… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-it.emotion-vectors-experiment-artifacts
Emotion-vectors replication on Gemma-4-31B: experiment artifacts
Every activation tensor, prompt set, and scored output behind the report
notebooks of gemma4-emotion-vectors
(commit a8e2352), published so replication does NOT require re-running
inference. The research record (hypotheses, pre-registered predictions,
verdicts) is the repo's TREE.md; the daily log is RESEARCH_LOG.md.
Models: google/gemma-4-31b (base) and google/gemma-4-31b-it (instruct),
bf16. Layers: range(0, 60… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-experiment-artifacts.emotion-vectors-gemma-4-31b-postfix
Emotion vectors, google/gemma-4-31b (corrected extraction)
Residual-stream activations for google/gemma-4-31b, pooled per story and averaged per
emotion. Each emotion ends up as one direction in the model's activation space.
Read LINEAGE.md before using this. This set supersedes
abotresol/emotion-vectors-gemma-4-31b. The earlier extraction ran
while the tokenizer padded on the left, so the step that skips a story's first
50 tokens skipped padding instead. This set re-extracts… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-postfix.whisper_transcriptions.reazonspeech.all.wer_10.0.vectorizedemotion-dialogue-vectors-gemma-4-31b
Emotion vectors — gemma-4-31b (base) probed on base-generated dialogues
Data provenance (what made these activations)
Probed model: google/gemma-4-31b (base)
Input corpus: abotresol/emotion-dialogues-gemma-4-31b — two-person dialogues written by the base model (generator = probed model; 44% emotion-word leakage, documented)
Per-story pooled residual-stream activations and per-emotion mean vectors,
extracted with gemma4-emotion-vectors… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-dialogue-vectors-gemma-4-31b.sonic-o1
SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding
🎯 What is SONIC-O1?
The first open-source benchmark for evaluating omnimodal video understanding with systematic fairness analysis. SONIC-O1 requires models to jointly process audio, video, and social context from real-world interactions—not just transcripts.
Key… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/sonic-o1.openiti-vectors
OpenITI Vector Database — maktabati.ai
🇬🇧 English
This dataset contains the fully vectorized OpenITI RELEASE 2025-1-9 collection of classical Islamic texts, prepared for semantic search (RAG).
Each entry represents a text chunk from one of 8,943 works in the OpenITI corpus, together with its embedding vector and complete metadata.
Statistics:
4,696,703 chunks
8,943 works (primary editions only, status=pri from OpenITI TSV)
approx. 470 Parquet files (approx.… See the full description on the dataset page: https://huggingface.co/datasets/Maktabati/openiti-vectors.open-pmc
OPEN-PMC
Arxiv: Arxiv
|
Code: Open-PMC Github
|
Model Checkpoint: Hugging Face
Dataset Summary
This dataset consists of image-text pairs extracted from medical papers available on PubMed Central. It has been curated to support research in medical image understanding, particularly in natural language processing (NLP) and computer vision tasks related to medical imagery. The dataset includes:
Extracted images from research articles.… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/open-pmc.cwicr-vector-db-bgem3-v3
CWICR Vector Database — BGE-M3 V3 Snapshots
Production Qdrant snapshots for CWICR (Construction Works Items, Costs & Resources) — a multilingual catalogue of construction rate databases covering 30 countries / language locales. Each snapshot encodes one country's rate book using the BAAI/bge-m3 embedder and is ready to restore directly into a Qdrant server for hybrid semantic search.
These snapshots are the V3 production artifacts produced by the OpenConstructionEstimate / CWICR… See the full description on the dataset page: https://huggingface.co/datasets/DataDrivenConstruction/cwicr-vector-db-bgem3-v3.voice_medical_cut_medium_vectoronevision1.5shamela-vectors
🇬🇧 English
The largest open-source vector database of the complete al-Maktaba al-Shamela (المكتبة الشاملة) Islamic text corpus, prepared for semantic search and Retrieval-Augmented Generation (RAG). Includes the full Quran text (6,236 verses, Hafs 'an 'Asim).
Statistics:
11,482,164 chunks from 8,589 classical Islamic books
6,236 Quran verses (one verse = one chunk, included in total)
40 categories covering the full breadth of Islamic scholarship
Period: 1st century AH to 15th… See the full description on the dataset page: https://huggingface.co/datasets/Maktabati/shamela-vectors.vector-100k
VectorOS Vector 100k SimSat VLM Dataset
VectorOS Vector 100k is a high-fidelity multimodal instruction dataset for fine-tuning vision-language models on geospatial epidemiology tasks. It was built for the VectorOS hackathon project and targets LiquidAI/LFM2.5-VL-450M.
The dataset contains 100,000 chat-style examples derived from 10,000 geospatial chips across 30 AOIs. Every accepted chip has a real SimSat Sentinel-2 true-color view, a real SimSat Sentinel-2 NIR-red-green false-color… See the full description on the dataset page: https://huggingface.co/datasets/Alfaxad/vector-100k.MAPF-GPT-vectorsneutral-vectors-gemma-4-31b-it-postfix
Emotion vectors, google/gemma-4-31b-it (corrected extraction)
Residual-stream activations for google/gemma-4-31b-it, pooled per story and averaged per
emotion. Each emotion ends up as one direction in the model's activation space.
Read LINEAGE.md before using this. This set supersedes
abotresol/neutral-vectors-gemma-4-31b-it. The earlier extraction ran
while the tokenizer padded on the left, so the step that skips a story's first
50 tokens skipped padding instead. This set… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/neutral-vectors-gemma-4-31b-it-postfix.Orion-Creative_Writing-ComplexityGeoFidelity-Bench
GeoFidelity-Bench
GeoFidelity-Bench: Evaluating Block-Conditioned Geographic Fidelity in Text-to-Image Street-View Generation
Kaizhen Tan, NeurIPS 2026. Contact: kt3275@nyu.edu.
Version 3.1.0 aligns the public metadata with the original experiments:
112 named street blocks, 25 cities, 23 countries, 7,563 reference assignments
covering 7,433 distinct Mapillary image IDs, and 16,128 generated images.
The generations cover six models, six prompt conditions, and four samples
per… See the full description on the dataset page: https://huggingface.co/datasets/moss-vector-714/GeoFidelity-Bench.steering-vectors-llama70bAIME_2024_DeepSeek_R1_0528_Temp_1.0_L_16384Responses of deepseek-ai/DeepSeek-R1-0528 for AIME 2024 (original dataset: Maxwell-Jia/AIME_2024).
Generation temperature is set to 1.0 and maximum token is set to 16384.
Dutch-Judiciary-Court-Cases-Netherlands-Rechtspraak-Vector-V3
