datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vector-100k
VectorOS Vector 100k SimSat VLM Dataset
VectorOS Vector 100k is a high-fidelity multimodal instruction dataset for fine-tuning vision-language models on geospatial epidemiology tasks. It was built for the VectorOS hackathon project and targets LiquidAI/LFM2.5-VL-450M.
The dataset contains 100,000 chat-style examples derived from 10,000 geospatial chips across 30 AOIs. Every accepted chip has a real SimSat Sentinel-2 true-color view, a real SimSat Sentinel-2 NIR-red-green false-color… See the full description on the dataset page: https://huggingface.co/datasets/Alfaxad/vector-100k.Dutch-Judiciary-Court-Cases-Netherlands-Rechtspraak-Vector-V3Tauri-RL-Plaintext-System-V2Tauri-RL-Markdown-System-V2Tauri-Opus-Accepted-GPT-Rejected-Opus-Writing-PromptsQwen3-0.6B-pts-steering-vectors
PTS Steering Vectors Dataset
A dataset of activation-based steering vectors created using the Pivotal Token Search (PTS) technique.
Details
Source: Generated using the PTS tool
Model: Qwen/Qwen3-0.6B
Dataset Structure
This dataset contains:
steering_vectors.jsonl: The main file with token-level steering vectors
Usage
These steering vectors can be used for activation-based steering during inference to guide language models toward particular… See the full description on the dataset page: https://huggingface.co/datasets/codelion/Qwen3-0.6B-pts-steering-vectors.Hydrus-Claude-Instruct-5K
kalo made dis
Thanks to Kubernetes bad for filtering + converting this to sharegpt
Mix of Opus and 3.5 for data
This is a combined set of
uncurated-raw-gens-og-test-filtered
uncurated-raw-gens-opus-jul-31-filtered
opus_jul10_test-filtered
uncurated_opus_jul8-filtered
microduck-policy-golden-vectors
Microduck policy golden vectors
Observation → action pairs recorded from Pollen Robotics' trained
Microduck policies, so that anybody writing their
own runner can check it against the same numbers instead of against a video.
This is a conformance fixture, not a model and not a dataset to train on. It contains no
weights. If you want the networks, they are Pollen's, in
pollen-robotics/microduck and
pollen-robotics/microduck_rl.
What is in it
golden_policies.json… See the full description on the dataset page: https://huggingface.co/datasets/craigm26/microduck-policy-golden-vectors.vector-9-17-sft1Semantic-Search-Engine-with-Vectorized-DB
Semantic Search Engine with Vectorized DB — Artifacts
This repository hosts the pre-computed on-disk index artifacts for the 20,000,000 vector database (OpenSubtitles_en_20M_emb_64.dat), built for the Advanced Database Systems project (Cairo University, Faculty of Engineering).
📁 Repository Structure
semantic-search-artifacts/
│
├── README.md # Repository documentation & usage guide
│
├── production/
│ ├── m1_ivf_k4096/… See the full description on the dataset page: https://huggingface.co/datasets/Final-Progs/Semantic-Search-Engine-with-Vectorized-DB.open-web-vectors-manifest
Open Web Vector Initiative — Site Manifest
Per-site metadata for every site in the Open Web Vector Initiative, including
what each site told us about AI use on the day we asked.
The initiative — how the permission gate works, and what we will and will
not publish: https://divinci.ai/open-web-vectors/
The live directory — search the corpus, chat with any site in it, or claim
your own: https://divinci.ai/www-rag/
This dataset contains no page text and no embeddings. That is… See the full description on the dataset page: https://huggingface.co/datasets/Divinci-AI/open-web-vectors-manifest.ps4guard-safety-vector-resultsDelta-Vector__Control-8B-details
Dataset Card for Evaluation run of Delta-Vector/Control-8B
Dataset automatically created during the evaluation run of model Delta-Vector/Control-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Delta-Vector__Control-8B-details.vector-sft2llm-eval-requestsTauri-KTO-Instruct-Mixassistant-axis-vectors
Assistant Axis Vectors for gemma-3-27b-it
This dataset contains pre-computed role vectors and the assistant axis for gemma-3-27b-it.
Overview
These vectors were computed using the methodology from the paper "The Assistant Axis"
by Christina Lu et al. The vectors can be used for activation steering to control model behavior along the
"assistant-like" to "role-playing" spectrum.
Contents
gemma-3-27b-it/assistant_axis.pt - The computed assistant axis (principal… See the full description on the dataset page: https://huggingface.co/datasets/massines3a/assistant-axis-vectors.vector-sft3vector-9-14-sft2steering_vector_distillation_datavector-9-14unbias-plus-dataset
Unbias Dataset
This dataset contains configurations used for the Unbias project at the Vector Institute:
train_4 (config, default): Our newest and highest quality training split.
other_splits (config): Contains the earlier splits below.
train_1: Training split sourced from VLDBench (regenerated version).
train_2: Another training split.
train_3: Another training split.
test_set: Test split sourced from BABE Golden 500.
⭐ train_4 is our newest, highest quality… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/unbias-plus-dataset.fairlens
FairLens: Benchmarking Bias in Vision-Language Models Across High-Stakes Domains
FairLens evaluates fairness and evidential validity in vision-language model (VLM) responses to high-stakes questions about people, across three domains: hiring, legal, and healthcare.
Each question is designed around one idea: a face image alone often cannot justify a judgment about someone's qualifications, threat level, illness, or professional role.… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/fairlens.vector_dbHydrus-The-Elder-Scrolls-Instruct
TES instruct
First attempt at mass-data-gen, wanted to get some instruct data related to games and such so i used a scrape for the elder scrolls wiki's including oblivion, morrowind, skyrim and then i deduped for the most unique possible entries + So i don't get killed by the guy who uses my deepseek API key
1. Initial Data gen
The base file containing source texts (e.g., lore articles) was used as input. For each entry, deepseek V3.1 was prompted to generate a… See the full description on the dataset page: https://huggingface.co/datasets/Delta-Vector/Hydrus-The-Elder-Scrolls-Instruct.Hydrus-Task-JudgementDelta-Vector__Henbane-7b-attempt2-details
Dataset Card for Evaluation run of Delta-Vector/Henbane-7b-attempt2
Dataset automatically created during the evaluation run of model Delta-Vector/Henbane-7b-attempt2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Delta-Vector__Henbane-7b-attempt2-details.Ursa-Completion-LIThttps://huggingface.co/datasets/AquaV/Lit
What i did was i converted each book to it's own JSONL with each line in the JSONL being its own chapters, after that was a simple merge between them keeping things in order and i ended out with this
sde-bench
sde-bench — does memory help a coding agent?
61 bug-fix tasks on a real codebase where every task hinges on a non-guessable,
project-specific decision: the obvious fix passes the visible repro test and fails a held-out
hidden test, because the project long ago decided the rule the obvious fix violates. The decision
lives in the repo's git history (28 tasks), a past developer conversation
(27), or a conversation later amended (6 — a cross-chat consolidation test).
Whether a… See the full description on the dataset page: https://huggingface.co/datasets/vectorize-io/sde-bench.tdk-vector-store
