datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sonic-o1
SONIC-O1: A Real-World Benchmark for Evaluating Multimodal Large Language Models on Audio-Video Understanding
🎯 What is SONIC-O1?
The first open-source benchmark for evaluating omnimodal video understanding with systematic fairness analysis. SONIC-O1 requires models to jointly process audio, video, and social context from real-world interactions—not just transcripts.
Key… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/sonic-o1.openiti-vectors
OpenITI Vector Database — maktabati.ai
🇬🇧 English
This dataset contains the fully vectorized OpenITI RELEASE 2025-1-9 collection of classical Islamic texts, prepared for semantic search (RAG).
Each entry represents a text chunk from one of 8,943 works in the OpenITI corpus, together with its embedding vector and complete metadata.
Statistics:
4,696,703 chunks
8,943 works (primary editions only, status=pri from OpenITI TSV)
approx. 470 Parquet files (approx.… See the full description on the dataset page: https://huggingface.co/datasets/Maktabati/openiti-vectors.shamela-vectors
🇬🇧 English
The largest open-source vector database of the complete al-Maktaba al-Shamela (المكتبة الشاملة) Islamic text corpus, prepared for semantic search and Retrieval-Augmented Generation (RAG). Includes the full Quran text (6,236 verses, Hafs 'an 'Asim).
Statistics:
11,482,164 chunks from 8,589 classical Islamic books
6,236 Quran verses (one verse = one chunk, included in total)
40 categories covering the full breadth of Islamic scholarship
Period: 1st century AH to 15th… See the full description on the dataset page: https://huggingface.co/datasets/Maktabati/shamela-vectors.VectorBenchmarkGeoFidelity-Bench
GeoFidelity-Bench
GeoFidelity-Bench: Evaluating Block-Conditioned Geographic Fidelity in Text-to-Image Street-View Generation
Kaizhen Tan, NeurIPS 2026. Contact: kt3275@nyu.edu.
Version 3.1.0 aligns the public metadata with the original experiments:
112 named street blocks, 25 cities, 23 countries, 7,563 reference assignments
covering 7,433 distinct Mapillary image IDs, and 16,128 generated images.
The generations cover six models, six prompt conditions, and four samples
per… See the full description on the dataset page: https://huggingface.co/datasets/moss-vector-714/GeoFidelity-Bench.s2ef-15m
Dataset Description
This dataset contains a collection of 3D atomistic datasets with force and energy labels gathered from a series of sources:
Open Catalyst Project
OC20, OC22, ODAC23
Materials Project Trajectory Dataset (MPtrj)
SPICE 1.1.4
Dataset Structure
Data Instances
For each instance, there is set of atomic numbers (input_ids), 3-D coordinates (coords), a set of forces per atom (forces), the total and formation energy per
system… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/s2ef-15m.pi07_cable_three_vector_v1
Three-holder cable routing with vector goals
Real-robot demonstrations for a vector-conditioned low-level policy: place three holders and route a cable through each holder. This is the validated LeRobot v3 dataset prepared for the first Pi0.7 four-camera-goal cable policy.
Property
Value
Source recordings
140
Subtask episodes
840
Frames
335,897
Sampling rate
100 Hz
Robot
ARX bimanual
Recorded state / action
14 / 14 dimensions
Video resolution
448 × 448… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/pi07_cable_three_vector_v1.route_red_yellow_vector_subtasks_pi05
Route Red-Yellow Vector Subtasks for pi0.5
This is a LeRobot v2.1 transformation of DistantSky/route at commit aced1e96f6b8bf98ffaa0754407636e82442b084.
The dataset contains 212 real-robot episodes, 135699 frames, three cameras, and 14-dimensional actions at 100 Hz.
Conditioning
Only observation.images.video_overhead is modified. The left and right videos are byte-identical to the source dataset.
The overhead image receives one fixed selected-connector pose glyph… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/route_red_yellow_vector_subtasks_pi05.MobileGym-ConAct-Trajectories
MobileGym-ConAct-Trajectories
Dataset Viewer · MobileGym · MemGUI-Agent · Paper
Abstract
MobileGym-ConAct-Trajectories is a release of successful mobile GUI-agent rollouts collected in the MobileGym simulator. Each trajectory is selected from judge-verified rollouts using a deterministic per-task rule: retain the shortest structurally valid success, then break ties by source run and episode ID. The release preserves screenshots, the rendered prompt supplied to the… See the full description on the dataset page: https://huggingface.co/datasets/Ma-Vector/MobileGym-ConAct-Trajectories.trash_pickup_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 25,
"total_frames": 7500,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/vectorcrumb/trash_pickup_v1.wiki-image-vectorsVector embeddings of images on Wikipedia. Images were embedded with "google/siglip2-base-patch16-384"
The "002" partition contains all of the quality images.
shamela-vectors
🇬🇧 English
The largest open-source vector database of the complete al-Maktaba al-Shamela (المكتبة الشاملة) Islamic text corpus, prepared for semantic search and Retrieval-Augmented Generation (RAG). Includes the full Quran text (6,236 verses, Hafs 'an 'Asim).
Statistics:
11,482,164 chunks from 8,589 classical Islamic books
6,236 Quran verses (one verse = one chunk, included in total)
40 categories covering the full breadth of Islamic scholarship
Period: 1st century AH to 15th… See the full description on the dataset page: https://huggingface.co/datasets/Kandil7/shamela-vectors.hackernews-vector-search-datasetThe Hacker News dataset contains 28.74 million postings and their vector embeddings. The embeddings were generated using SentenceTransformers model all-MiniLM-L6-v2. The dimension of each embedding vector is 384.
Created by clickhouse more info: https://clickhouse.com/docs/getting-started/example-datasets/hackernews-vector-search-dataset
Qwen3-0.6B-pts-steering-vectors
PTS Steering Vectors Dataset
A dataset of activation-based steering vectors created using the Pivotal Token Search (PTS) technique.
Details
Source: Generated using the PTS tool
Model: Qwen/Qwen3-0.6B
Dataset Structure
This dataset contains:
steering_vectors.jsonl: The main file with token-level steering vectors
Usage
These steering vectors can be used for activation-based steering during inference to guide language models toward particular… See the full description on the dataset page: https://huggingface.co/datasets/codelion/Qwen3-0.6B-pts-steering-vectors.autonomous-db-internals-vector-search-suite
⚡ Autonomous Database Internals, Vector Search Engines & Distributed Storage Suite (2026)
A Production-Grade, Verifiable Synthetic Corpus for Training Autonomous Database Kernel & Vector Retrieval LLMs
💼 Get Full 12,500-Row Enterprise Suite on Gumroad →
Full 10,000 SFT + 2,500 DPO Rows • 254.6 MB Pre-Indexed SQLite DB • RLVR/GRPO Sandboxed Testbed • Commercial License
⚡ Overview & Industry Problem
Deploying autonomous AI… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/autonomous-db-internals-vector-search-suite.vector-9-17-sft1frequent-stock-fccb17
frequent-stock-fccb17
Synthetic sensors test data: 39 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/vectorJoseph/frequent-stock-fccb17.details_Delta-Vector__Odin-9B
Dataset Card for Evaluation run of Delta-Vector/Odin-9B
Dataset automatically created during the evaluation run of model Delta-Vector/Odin-9B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Delta-Vector__Odin-9B.africa-synth-malaria-vector-surveillance-control-all
Vector Surveillance & Control | Africa (World Health Organization)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets help researchers examine… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-malaria-vector-surveillance-control-all.Semantic-Search-Engine-with-Vectorized-DB
Semantic Search Engine with Vectorized DB — Artifacts
This repository hosts the pre-computed on-disk index artifacts for the 20,000,000 vector database (OpenSubtitles_en_20M_emb_64.dat), built for the Advanced Database Systems project (Cairo University, Faculty of Engineering).
📁 Repository Structure
semantic-search-artifacts/
│
├── README.md # Repository documentation & usage guide
│
├── production/
│ ├── m1_ivf_k4096/… See the full description on the dataset page: https://huggingface.co/datasets/Final-Progs/Semantic-Search-Engine-with-Vectorized-DB.open-web-vectors-manifest
Open Web Vector Initiative — Site Manifest
Per-site metadata for every site in the Open Web Vector Initiative, including
what each site told us about AI use on the day we asked.
The initiative — how the permission gate works, and what we will and will
not publish: https://divinci.ai/open-web-vectors/
The live directory — search the corpus, chat with any site in it, or claim
your own: https://divinci.ai/www-rag/
This dataset contains no page text and no embeddings. That is… See the full description on the dataset page: https://huggingface.co/datasets/Divinci-AI/open-web-vectors-manifest.international-delivery-68f942
international-delivery-68f942
Synthetic sensors test data: 34 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting… See the full description on the dataset page: https://huggingface.co/datasets/vectorDawn/international-delivery-68f942.unhappy-discipline-7f43f3
unhappy-discipline-7f43f3
Synthetic weather test data: 58 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Vector-Xueyong/unhappy-discipline-7f43f3.dirty-occasion-d21fee
dirty-occasion-d21fee
Synthetic sensors test data: 34 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/vectorRemy/dirty-occasion-d21fee.Delta-Vector__Control-8B-details
Dataset Card for Evaluation run of Delta-Vector/Control-8B
Dataset automatically created during the evaluation run of model Delta-Vector/Control-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Delta-Vector__Control-8B-details.vector-sft2llm-eval-requestssignificant-table-7d9bce
significant-table-7d9bce
Synthetic sensors test data: 37 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/vectorridge/significant-table-7d9bce.Vector-QM24_DFT_all
Cite this dataset Khan, D., Benali, A., Kim, S. Y. H., Rudorff, G. F., and Lilienfeld, O. A. Vector-QM24 DFT all. ColabFit, 2025. https://doi.org/None
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_typ49f9r0b6v_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Vector-QM24_DFT_all.scoutieDataset_russian_language_grammar_and_rules_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning the Russian language. This dataset contains grammar, syntax, spelling and punctuation rules.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_russian_language_grammar_and_rules_vectorized.
