datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ProcVQA-20M-annotations
ProcVQA-20M Annotations
Project Page |
arXiv |
Code |
Model |
Media
This repository contains the text annotations for the ProcVQA-20M dataset. The full image files are hosted separately on ProcVQA-20M-media.
Overview
This dataset is constructed from over 26 embodied datasets, comprising:
20M QA pairs for training
330K original trajectories
50M annotated frames from ~5,000 hours of manipulation data
200+ different tasks
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/ce-amtic/ProcVQA-20M-annotations.The-Oxford-IIIT-Pet-Dataset-With-Annotationsforest-fire-annotations
Forest Fire Detection Dataset — Auto-Annotated
Bounding-box annotated version of touati-kamel/forest-fire-dataset,
built for training forest-fire / smoke / fog object detection models.
Overview
This dataset contains video frames auto-labeled with bounding boxes for fire and
smoke-related visual phenomena, using a zero-shot open-vocabulary object detector
(Grounding DINO). It is derived from the original touati-kamel/forest-fire-dataset image
classification dataset… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/forest-fire-annotations.ToothXpert.MM-OPG-Annotations4k-video-annotations
4K Video Annotations — Shot Segmentation and Camera Motion
This dataset contains 12 frame-accurate shot clips segmented from five short cinematic video sequences. Every clip is paired with a detailed, manually reviewed annotation covering visible content, subject actions, shot scale, camera angle, camera movement, movement direction, stabilization, composition, lighting, color, pacing, transitions, timecodes, and technical properties.
The footage depicts a tense nighttime… See the full description on the dataset page: https://huggingface.co/datasets/LianeMarilin/4k-video-annotations.od-syn-page-annotations
📦 Dhivehi Synthetic Document Layout + Textline Dataset
This dataset contains synthetically generated image-document pairs with detailed layout annotations and ground-truth Dhivehi text extractions.It’s designed for document layout analysis , visual document understanding , OCR fine-tuning, and related tasks specifically for Dhivehi script.
📋 Dataset Summary
Total Examples: ~58,738
Image Content: Synthetic Dhivehi documents generated to simulate real-world layouts… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/od-syn-page-annotations.od-syn-page-annotations-com
📦 Dhivehi Synthetic Document Layout + Textline Dataset
This dataset contains synthetically generated image-document pairs with detailed layout annotations and ground-truth Dhivehi text extractions.It’s designed for document layout analysis, visual document understanding, OCR fine-tuning, and related tasks specifically for Dhivehi script.
Note: this version image are compressed.
Raw version 📁 Repository: Hugging Face Datasets
📋 Dataset Summary
Total Examples: ~58… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/od-syn-page-annotations-com.usd-side-coco-annotations
USD Side Detection Dataset (Front/Back)
A refined COCO-format dataset for detecting US Dollar currency and classifying whether the front or back side is visible.
Dataset Summary
Total Images: 3,618
Total Annotations: 3,746
Format: COCO + HuggingFace JSONL
Classes: 24 (denominations × front/back × authentic/counterfeit)
Classification Accuracy: 100% (all Front/Back classified)
Split
Images
Annotations
Train
2,671
2,738
Valid
597
627
Test
350
381… See the full description on the dataset page: https://huggingface.co/datasets/ebowwa/usd-side-coco-annotations.orena-frame-annotations
ORena FOCUS 2026 — FRAME supplementary annotations (public half)
Supplementary VQA annotations produced by MLO-Lab for the ORena FOCUS 2026 FRAME track.
This is the openly releasable half; the LapChole-FOCUS half is withheld under that dataset's
usage agreement until the organisers publish it.
rows
vqa/heico_derived.jsonl — HeiCo-FOCUS
7,608
vqa/hernia_mesh.jsonl — hernia videos
1,350
total
8,958
Also included: raw/hernia_mesh_annotations/ (12 frame-level… See the full description on the dataset page: https://huggingface.co/datasets/Machine-Learning-Oncology/orena-frame-annotations.bower-waste-annotations
Dataset Card for waste annotations made by the recycling solution Bower
The data offered by Bower (Sugi Group AB) in collaboration with Google.org
Dataset Summary
The bower-waste-annotations dataset consists of 1440 images of waste and various consumer items taken by consumer phone cameras. The images are annotated with Material type and Object type classes, listed below.
The images and annotations has been manually reviewed to ensure correctness. It is assumed… See the full description on the dataset page: https://huggingface.co/datasets/BowerApp/bower-waste-annotations.orena-segment-annotations
ORena FOCUS 2026 — SEGMENT supplementary annotations (public half)
Supplementary VQA annotations produced by MLO-Lab for the ORena FOCUS 2026 SEGMENT
track. This is the openly releasable half; the LapChole-FOCUS half is withheld under that
dataset's usage agreement until the organisers publish it.
rows
vqa/heico_derived.jsonl — HeiCo-FOCUS
1,300
vqa/hernia_mesh.jsonl — hernia videos
140
total
1,440
Also included: raw/hernia_mesh_annotations/ (12 frame-level… See the full description on the dataset page: https://huggingface.co/datasets/Machine-Learning-Oncology/orena-segment-annotations.pokemon-cards-image-and-annotationsreal-resumes-section-detection-annotationsamz-image-annotationsannotations_only_10pct_gpt5_miniImageIn_annotations_resized_images
Dataset Card for ImageIn_annotations_resized_images
More Information needed
ImageIn_annotationsInitial annotated dataset derived from ImageIN/IA_unlabelled
new_year-50-annotationsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "openarms_follower",
"total_episodes": 23,
"total_frames": 42588,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:23"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jadechoghari/new_year-50-annotations.multilingual-image-annotations
Multilingual Image Annotations
Image annotations across 7 languages (en, es, fr, hi, zh, ar, pt) generated by google/gemma-4-31B-it via the Hugging Face Router. Each row pairs an image with an English description, multilingual descriptions, 21 VQA pairs (3 per language), and conditional object detections with normalized bounding boxes. When detections are present, a derivative image with rectangles drawn is included as boxed_image.
Stats
Images: 464… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/multilingual-image-annotations.TerraVis-Annotations
TerraVis Annotations
Human ratings of world-grounded visual consistency for 4,500 AI-generated images. World consistency asks whether the entities, structures, interactions and phenomena in an image are visually plausible with respect to the real world, regardless of the prompt or the visual style. These are the human reference scores used to validate the TerraVis metric in our NeurIPS 2026 paper.
📄 Paper: TerraVis: Towards Evaluation of World-Grounded Visual Consistency in… See the full description on the dataset page: https://huggingface.co/datasets/ShyFoo/TerraVis-Annotations.TriConflict-hallucination-annotationsleicester_loaded_annotations_binary
Dataset Card for "leicester_loaded_annotations_binary"
More Information needed
task2-1_original-annotationsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 45,
"total_frames": 37489,
"total_tasks": 3,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:45"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jasontchan/task2-1_original-annotations.ImageIn_annotationsleicester_loaded_annotations
Dataset Card for "leicester_loaded_annotations"
More Information needed
test-phd-annotations
PhD Hallucination Annotations
This dataset contains hallucination annotations for the PhD dataset.
Usage
from datasets import load_dataset
dataset = load_dataset("alita01/test-phd-annotations")
print(dataset)
# View a sample
sample = dataset['phd_ccs'][0]
print(sample['question'])
sample['image'].show()
Fields
image: Original image (PIL Image)
question: Input question
model_output: Model's generated response
ground_truth: Ground truth answer… See the full description on the dataset page: https://huggingface.co/datasets/alita01/test-phd-annotations.TriConflict-annotationsNamgyal-OCR-AnnotationsMultimodal-Hallucination-Annotationsreasoning_annotations_002
