datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lora-training-datasetssole_training_data
This is the training dataset for SOLE-R1-8B
SOLE-R1-8B is a video-language reward reasoning model for robotics. It is designed to estimate task progress from robot video frames and a natural-language task description, producing both per-timestep reasoning traces and scalar progress predictions that can be used as rewards for online robot reinforcement learning.
This dataset accompanies the paper “SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot RL” by Philip… See the full description on the dataset page: https://huggingface.co/datasets/Philip-MIT/sole_training_data.Bee-Training-Data-Stage2
Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs
[🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code]
Introduction
We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality.
Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Bee-Training-Data-Stage2.e2e-stream-slam-training-dataset
e2e-stream-slam training assets
Reproducibility bundle for the V4 SLAMFormer ablation suite.
Code: https://github.com/SlamMate/e2e-semantic-SLAM/tree/submap (commit 3195a7a)
Contents
Checkpoints
File
Size
Role
checkpoints/v1_paper_ckpt10.pth
3.6 GB
SLAMFormer paper base ckpt (10 ep on the paper datasets). PRETRAINED init for V3 Scale Token training.
checkpoints/v3_scale_token_ckpt2.pth
3.8 GB
V3 Scale Token epoch-2 (3 ep, 3×A6000… See the full description on the dataset page: https://huggingface.co/datasets/qizhangslam/e2e-stream-slam-training-dataset.gaussian_training_datasets
Gaussian Training Datasets (COLMAP) for msplat
COLMAP-format multi-view scenes for training 3D Gaussian Splatting models,
packaged for msplat — a Metal-native 3DGS
trainer for Apple Silicon. Also includes pre-trained .ply splats under
tested_outputs/.
All scenes are redistributed from third-party datasets. Full credit goes to
their original authors — see Licensing & credits and please
cite the original papers. This repo only repackages them in COLMAP layout for
convenience.… See the full description on the dataset page: https://huggingface.co/datasets/alexmkwizu/gaussian_training_datasets.GEWDiff_training_dataset
GEWDiff Training & Evaluation Dataset
📘 Overview
The GEWDiff Training & Evaluation Dataset is derived from the EnMAP Champion and MDAS hyperspectral datasets.It is designed for image enhancement, super-resolution, restoration, and generative remote sensing tasks.The dataset includes Low-Quality (LQ) low-resolution images, corresponding Ground-Truth (GT) high-resolution images, and optional structure information such as masks and edges (partially provided;… See the full description on the dataset page: https://huggingface.co/datasets/zhu-xlab/GEWDiff_training_dataset.DeepVRM-Training-DataBee-Training-Data-Stage2
Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs
[🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code]
Introduction
We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality.
Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/toilaluan/Bee-Training-Data-Stage2.GEWDiff_training_dataset
GEWDiff Training & Evaluation Dataset
📘 Overview
The GEWDiff Training & Evaluation Dataset is derived from the EnMAP Champion and MDAS hyperspectral datasets.It is designed for image enhancement, super-resolution, restoration, and generative remote sensing tasks.The dataset includes Low-Quality (LQ) low-resolution images, corresponding Ground-Truth (GT) high-resolution images, and optional structure information such as masks and edges (partially provided;… See the full description on the dataset page: https://huggingface.co/datasets/zhang5566/GEWDiff_training_dataset.Chest_CT_GMPO_Training_Datasetsegmentation_training_data
Build Change augmented segmentation dataset
Paired images and segmentation masks, including augmented (rotated, flipped,
transformed) copies of each original photo, paired by filename. Every row links
to the original photo it was made from whenever that original is in the dataset.
Load it
No Hugging Face account or token is needed:
from datasets import load_dataset
ds = load_dataset("buildchange/segmentation_training_data", split="train")
row = ds[0]
row["image"]… See the full description on the dataset page: https://huggingface.co/datasets/buildchange/segmentation_training_data.VideoChat-Flash-Training-Data-subsetwebui-training-dataSOC-Training-Data-Visualization
Paper Link
SOS: Synthetic Object Segments Improve Detection, Segmentation, and Grounding
Code repo
Code for Generation
Citation
@misc{huang2025sossyntheticobjectsegments,
title={SOS: Synthetic Object Segments Improve Detection, Segmentation, and Grounding},
author={Weikai Huang and Jieyu Zhang and Taoyang Jia and Chenhao Zheng and Ziqi Gao and Jae Sung Park and Ranjay Krishna},
year={2025},
eprint={2510.09110},
archivePrefix={arXiv}… See the full description on the dataset page: https://huggingface.co/datasets/weikaih/SOC-Training-Data-Visualization.lingbot-top2400-training-data
LingBot Top2400-Compatible Training Data
Restricted internal training data. This dataset repository is private and is not intended for redistribution.
This is the self-contained training-data bundle referenced by the archived LingBot WoRL run. It contains the portable manifest together with every referenced first frame and camera-conditioning array.
Dataset summary
Item
Value
Portable training records
2,218
First-frame images
2,218
Camera arrays
4… See the full description on the dataset page: https://huggingface.co/datasets/qyoo/lingbot-top2400-training-data.Meta-CoT-Training-Data-ExampleMed_training_dataTinyLLAVA-Training-DataTrainingData_Stage3_VideoCases
Stage3 视频/多图人工查看样例
这是 18条人工查看样例,不是新训练集或评测集。直接向下浏览 QA 和拼图,
也可在 Dataset Viewer 中查看 frames 图片列。训练 QA 与原数据逐字保留;答案是数据集标注,不是人工确认的视觉真值。
来源
已确认视频类
待判断多图
CA-VQA
3(有序帧)
0
SpaceVista
3(有序帧)
3
VSI-590K
3(视频文件)
0
SIMS-VSI
3(视频文件)
0
ViCA-322K
3(视频文件)
0
选样:从正式 Small/train 中按固定 ID 顺序选择,优先不同任务、不同来源场景;不是随机代表性统计,也没有按答案是否正确挑样。
每张拼图从左到右、从上到下阅读。视频文件全程均匀抽取最多16帧;帧序列保留全部输入帧。
展示缩放不改变正式训练数据;拼图不是模型训练输入。多图未知项不计为视频。
来源:AnchorSR/TrainingData_Stage3,
固定提交… See the full description on the dataset page: https://huggingface.co/datasets/AnchorSR/TrainingData_Stage3_VideoCases.ilids-faint-patterns-training-datasetscta-htr-training-data
Guiding Principles of this Dataset
Segmentation
The goal is here to segment the "main text" while ignoring paratext. This means page numbers, catch words, marginal annotations, are purposefully not segmented
Transcription
The guiding transcription principles are to
expand abbreviations
preserve orthography
Andrew_Alpha_training_data
Bark Beetle Grouped Images for AI Classification
Dataset Summary
This dataset comprises high-resolution photographs of bark and ambrosia beetles captured under controlled laboratory conditions. Each image contains multiple beetle specimens arranged on a uniform white background while submerged in 70% ethanol. This approach speeds up data collection and ensures reproducible imaging conditions. Individual beetle images can later be extracted from these grouped photographs… See the full description on the dataset page: https://huggingface.co/datasets/ChristopherMarais/Andrew_Alpha_training_data.Brain-Tumor-MRI-Dataset-TrainingBee-Training-Data-Stage1
Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs
[🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code]
Introduction
We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality.
Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Bee-Training-Data-Stage1.LayoutLM_Training_datasettraining
Training Data for Detecting Drones
This dataset was generated for a bachelor degree in computer science.
The aim was to improve the detection score on the drone versus bird dataset.
The images in each folder are from videos with a frame rate of 25 FPS.
They are intended to be used for the training of recurrent computer vision models.
herislab-ca-training-data
CA_Training_Data -- Convolutional Autoencoder (Track A)
Curated dataset for training and evaluating the Convolutional Autoencoder anomaly detection model.
Approach
The autoencoder is trained only on normal (no-fault) images. At inference, high reconstruction error indicates an anomaly/fault.
Structure
train/normal/ -- Normal images for autoencoder training
electric_motor/ -- 168 PNG (Electric Motor Thermal Fault Diagnosis, no_fault class)… See the full description on the dataset page: https://huggingface.co/datasets/Ryanflash/herislab-ca-training-data.laser_gui_grounding_training_datazalo-ai-2025-training-data-v2
Zalo AI Challenge 2025 - RoadBuddy Training Data V2set
This dataset contains training data v2 for the RoadBuddy – Understanding the Road through Dashcam AI challenge from Zalo AI Challenge 2025.
Dataset Description
The challenge aims to build a driving assistant capable of understanding video content from dashcams to quickly answer questions about traffic signs, signals, and driving instructions in Vietnam.
Dataset Structure
Files
frames/: Directory… See the full description on the dataset page: https://huggingface.co/datasets/OpenHay/zalo-ai-2025-training-data-v2.AIDA_ocr_training_data
OCR training data from AIDA-project
Dataset Summary
The zip file contains textlines and their annotations from AIDA-project. There are ~ 166k textlines that are mainly in Finnish language, but contain a little Swedish and
English and little French and German textlines. The textlines contains typewritten and also handwritten lines. Roughly 24 % of the annotated lines are handwritten and the rest are
typewritten. The dataset also contains 120 000 synthetic… See the full description on the dataset page: https://huggingface.co/datasets/Kansallisarkisto/AIDA_ocr_training_data.
