datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
XLRS-Bench_visual_grounding_en
🐙GitHub
Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench
📜Dataset License
Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from:
DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply).
ITCVDLicensed under CC-BY-NC-SA-4.0.
MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/XLRS-Bench_visual_grounding_en.HoliSpatial-3D-GroundingXLRS-Bench_visual_grounding_zh
🐙GitHub
Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench
📜Dataset License
Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from:
DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply).
ITCVDLicensed under CC-BY-NC-SA-4.0.
MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/XLRS-Bench_visual_grounding_zh.easyr1-grounding-dataset-30k-not_grounded-SE-GUI-3B-2MPFACTS-grounding-public
FACTS Grounding 1.0 Public Examples
860 public FACTS Grounding examples from Google DeepMind and Google Research
FACTS Grounding is a benchmark from Google DeepMind and Google Research designed to measure the performance of AI Models on factuality and grounding.
▶ FACTS Grounding Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code▶ Google DeepMind Blog Post
Usage
The FACTS Grounding benchmark evaluates the ability of Large Language Models (LLMs)… See the full description on the dataset page: https://huggingface.co/datasets/google/FACTS-grounding-public.breakpoint-grounding-55m
Breakpoint Grounding 55M
Quick start
from datasets import load_dataset
ds = load_dataset("BreakpointAI/breakpoint-grounding-55m", split="train")
ds[0] # {'image': <PIL.Image>, 'image_caption': ..., 'object_captions': [...], 'normalized_boxes': [...], 'img_size_wh': [...]}
Dataset summary
Breakpoint Grounding 55M is, to our knowledge, the largest instance-grounded image–text dataset
released publicly. Every image comes with an image-level caption… See the full description on the dataset page: https://huggingface.co/datasets/BreakpointAI/breakpoint-grounding-55m.sequential-3d-grounding
Sequential 3D Visual Grounding Dataset
5 indoor datasets (ScanNet, HM3D, 3RScan, ARKitScenes, MultiScan) · 10301 scenes · 117885 sequences · 579872 steps
Overview
Dataset
Scenes
Train
Val
Test
ScanNet
1513
14501
1700
1808
HM3D
2302
33302
7663
4232
3RScan
1381
13826
4864
1937
ARKitScenes
4834
24349
1394
2653
MultiScan
271
4133
858
665
Total
10301
90111
16479
11295
Annotation
Each step has one target (the object to locate)… See the full description on the dataset page: https://huggingface.co/datasets/Ziyannn/sequential-3d-grounding.FiftyOne-GUI-Grounding-Train
Dataset Card for FiftyOne GUI Grounding Training Set
Dataset Details
Dataset Description
This dataset contains 739 annotated GUI screenshots designed for training computer vision models to understand and interact with graphical user interfaces. The dataset uses the specialized COCO4GUI format, which extends the standard COCO detection format to handle GUI-specific features, interaction sequences, and rich metadata.
The dataset captures real user… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/FiftyOne-GUI-Grounding-Train.GroundingCapv2grounding_dataset
Grounding Dataset
A comprehensive, high-quality dataset for GUI element grounding tasks, curated from multiple authoritative sources to provide diverse, well-annotated interface interactions.
Overview
This dataset combines and standardizes annotations from five major GUI interaction datasets:
Aria-UI
OmniAct
Widget Caption
UI-Vision
OS-Atlas
Dataset Schema
Each sample contains the following fields:
Field
Type
Description
Example
dataset
string… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/grounding_dataset.grounding_datasetFor desktop and web datasets in GUI grounding, the data is generally collected via screenshots alongside accessibility tools like A11y or HTML parsers to extract element structure and bounding boxes. However, these bounding boxes may sometimes be misaligned with the visual rendering due to UI animations or timing inconsistencies. In our work, we primarily rely on datasets curated from Aria-UI and OS-Atlas, which we found to be cleaner and better aligned than alternative data collections.
To… See the full description on the dataset page: https://huggingface.co/datasets/HelloKKMe/grounding_dataset.easyr1-agent-grounding-dataUI-Grounding-Benchmarks
UI-Grounding-Benchmarks
This is a collection of UI grounding benchmarks:
ScreenSpot
ScreenSpot-V2
ScreenSpot-Pro
OS-World-G
UI-Vision
Thanks for their great work!
This benchmark collection is used in the paper:
FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
🖼️ Project Page: https://showlab.github.io/FocusUI/
🏠 Github Repo: https://github.com/showlab/FocusUI
📝 Paper: https://arxiv.org/pdf/2601.03928
Model Zoo
Model
Backbone
🤗… See the full description on the dataset page: https://huggingface.co/datasets/yyyang/UI-Grounding-Benchmarks.GroundingME
(CVPR 2026) GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation
🔍 Overview
Visual grounding—localizing objects from natural language descriptions—represents a critical bridge between language and vision understanding. While multimodal large language models (MLLMs) achieve impressive scores on existing benchmarks, a fundamental question remains: can MLLMs truly ground language in vision with human-like… See the full description on the dataset page: https://huggingface.co/datasets/lirang04/GroundingME.PanoCaps
PanoCaps: A Human-Annotated Benchmark for Panoptic Grounded Captioning
PanoCaps is a benchmark for panoptic grounded captioning: a model writes a full-scene caption and grounds every mentioned entity, things and stuff alike, to pixel-level masks.
It contains 3,470 images and 34K panoptic regions, averaging ~9 grounded entities per image, with >99% of regions grounded. Captions are human-written and verified, cover the entire visible scene, use open-vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/Panorama-grounding/PanoCaps.grounding-data
grounding_data
The annotation tree behind a grounding/segmentation training stack, plus the images and
video frames that live alongside it. Unlike the companion
royguw/mm-olmo-images — which is
pixels only — this repo carries the annotations: the parquet caches, JSON label files
and vocabularies under each dataset's cache/ directory, which is where masks, boxes,
referring expressions, captions and category vocabularies actually live.
9,331,994 files / 1.20 TiB, packed as 255 tar… See the full description on the dataset page: https://huggingface.co/datasets/royguw/grounding-data.FiftyOne-GUI-Grounding-Train-with-Synthetic
Dataset Card for FiftyOne GUI Grounding Training Set with Synthetic Augmentation
Dataset Details
Dataset Description
This dataset represents a significant expansion of the original FiftyOne GUI Grounding Training Set, growing from 739 real GUI screenshots to 4,036 total samples through systematic synthetic data generation. The dataset combines authentic GUI interactions with carefully crafted synthetic variants designed to improve model robustness… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/FiftyOne-GUI-Grounding-Train-with-Synthetic.PR1-Datasets-GroundingECG-Grounding
ECG-Grounding dataset
ECG-Grounding provides more accurate, holistic, and evidence-driven interpretations with diagnoses grounded in measurable ECG features. Currently, it contains 30,000 instruction pairs annotated with heartbeat-level physiological features. This is the first high-granularity ECG grounding dataset, enabling evidence-based diagnosis and improving the trustworthiness of medical AI. We will continue to release more ECG-Grounding data and associated… See the full description on the dataset page: https://huggingface.co/datasets/LANSG/ECG-Grounding.web-ui-grounding-jsonminecraft-grounding-action-datasetblip3-grounding-smallgrounding-atlas
grounding-atlas: verifiable-signal pairs
Matched (representation, verifiable-property) pairs for measuring whether a
language model grounds the content of a scientific representation (a SMILES
string, a protein/DNA/RNA sequence, an expression vector, a spectrum, an image)
or merely its name. Each property is either an experimentally measured endpoint
or a closed-form function of the representation, so the representation is the
ground truth and grounding becomes directly… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/grounding-atlas.PointCloud-groundingfineweb_synth_dense_ocr_gradients_with_grounding_Bvst_3d_grounding_benchmark
Visual Spatial Tuning
This dataset is the 3D object detection benchmark used in VST.
Please follow VST/benchmarkto evaluate the models.
git clone https://github.com/Yangr116/VST.git
python tools/download_hf_data.py --repo_id="rayruiyang/vst_3d_grounding_benchmark" --local_dir $YOUR_LOCAL_PATH
cd $YOUR_LOCAL_PATH/vst_3d_grounding_benchmark
tar -zxvf arkit/arkit_omni3d_test_640x480.tar.gz -C arkit
tar -zxvf sunrgbd/sunrgbd_val.tar.gz -C sunrgbd… See the full description on the dataset page: https://huggingface.co/datasets/rayruiyang/vst_3d_grounding_benchmark.synthetic-grounding-images
Synthetic grounding images
3,162 images generated to extend visual grounding to concrete words that no
photograph dataset covers, for Augustinian BabyLM
(paper, code).
How they were made
Starting from 1,986 concrete words with no image support, an LLM
(claude-sonnet-4-6, temperature 0.8) wrote short scene descriptions
placing as many target words as fit naturally into one scene. Each of the
1,054 resulting descriptions was rendered three times with SDXL-Turbo
(2… See the full description on the dataset page: https://huggingface.co/datasets/augustinian-babylm/synthetic-grounding-images.highlevel_thinking_with_grounding_annotation_split1000_v2fineweb_synth_dense_ocr_colors_with_groundingECG-Protocol-Guided-Grounding-CoT
ECG-R1: Protocol-Guided and Modality-Agnostic MLLM for Reliable ECG Interpretation
If you find this project useful, please give us a star🌟.
Jiarui Jin, Haoyu Wang, Xingliang Wu, Xiaocheng Fang, Xiang Lan, Zihan Wang
Deyun Zhang, Bo Liu, Yingying Zhang, Xian Wu, Hongyan Li, Shenda Hong
Introduction
Electrocardiography (ECG) serves as an indispensable diagnostic tool in clinical practice, yet existing multimodal large language models (MLLMs) remain unreliable… See the full description on the dataset page: https://huggingface.co/datasets/PKUDigitalHealth/ECG-Protocol-Guided-Grounding-CoT.
