Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01wentao-yuan /robopoint-data RoboPoint Dataset Card Dataset details This dataset contains 1432K image-QA instances used to fine-tune RoboPoint, a VLM for spatial affordance prediction. It consists of the following parts: 347K object reference instances from a synthetic data pipeline; 320K free space reference instances from a synthetic data pipeline; 100K object detection instaces from LVIS; 150K GPT-generated instruction-following instances from liuhaotian/LLaVA-Instruct-150K; 515K general-purpose… See the full description on the dataset page: https://huggingface.co/datasets/wentao-yuan/robopoint-data.image1M<n<10M13 likes5.9k downloads2y agoHugging Face02Yale-BIDS-Chen /medpmc-11m-dataset_jun24_baseline MedPMC WebDataset MedPMC is a large-scale medical image-text dataset curated from articles in the PubMed Central (PMC) collection. This release contains approximately 11 million image-text pairs collected from the June 2024 PMC baseline. MedPMC is an ongoing effort, and future releases will continue to expand the dataset with newly published literature, improved annotations, and additional resources. This dataset is presented in the paper MedPMC: A Systematic Framework for… See the full description on the dataset page: https://huggingface.co/datasets/Yale-BIDS-Chen/medpmc-11m-dataset_jun24_baseline.imagezero-shot-image-classification1M<n<10M3 likes4.9k downloads3mo agoHugging Face03nvidia /describe-anything-dataset Describe Anything: Detailed Localized Image and Video Captioning NVIDIA, UC Berkeley, UCSF Long Lian, Yifan Ding, Yunhao Ge, Sifei Liu, Hanzi Mao, Boyi Li, Marco Pavone, Ming-Yu Liu, Trevor Darrell, Adam Yala, Yin Cui [Paper] | [Code] | [Project Page] | [Video] | [HuggingFace Demo] | [Model/Benchmark/Datasets] | [Citation] Dataset Card for Describe Anything Datasets Datasets used in the training of describe anything models (DAM). The datasets are in tar files. These… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/describe-anything-dataset.imageimage-to-text100K<n<1M59 likes3.6k downloads1y agoHugging Face04allenai /olmoearth_pretrain_datasetThis is the pre-training dataset for training the OlmoEarth pre-trained remote sensing foundation models. Documentation is on GitHub at https://github.com/allenai/olmoearth_pretrain/blob/main/docs/Pretraining-Dataset.md The dataset is released under CC BY 4.0. It includes data from the following sources: Sentinel-2 L2A imagery from the European Space Agency, available under the Copernicus Sentinel Data and Service Legal Notice Sentinel-1 GRD IW vv+vh imagery from the European Space Agency… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmoearth_pretrain_dataset.image100M<n<1B19 likes3.3k downloads11mo agoHugging Face05LPY /BridgeVLA_COLOSSEUM_EVAL_DATAarxiv: https://arxiv.org/abs/2506.07961 image1M<n<10M0 likes1.7k downloads1y agoHugging Face06OPPOer /X2Edit-Dataset X2Edit   Introduction X2Edit Dataset is a comprehensive image editing dataset that covers 14 diverse editing tasks and exhibits substantial advantages over existing open-source datasets including AnyEdit, HQ-Edit, UltraEdit, SEED-Data-Edit, ImgEdit and OmniEdit. For the relevant data construction scripts, model training and inference scripts, please refer to X2Edit. News 2025/09/16: We are about to release a dataset constructed by Qwen-Image and… See the full description on the dataset page: https://huggingface.co/datasets/OPPOer/X2Edit-Dataset.image10M<n<100M20 likes1.5k downloads10mo agoHugging Face07nvidia /Harmonizer-Datasetgated HARMONIZER DATASET Dataset Description Training dataset for DiffusionHarmonizer: a generative AI model for image and video enhancement bridging neural reconstruction and photorealistic simulation . Model checkpoints: https://huggingface.co/nvidia/Harmonizer/Training code: https://github.com/NVIDIA/harmonizer/ The dataset was curated to support the following functions of the model: 3D reconstruction artifact removal Harmonization of inserted objects to blend… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Harmonizer-Dataset.image100K<n<1M4 likes1.2k downloads3mo agoHugging Face08Uni-Edit /Uni-Edit-Train-Data Uni-Edit Training Data: Uni-Edit-148k Project Page | GitHub Repository | Paper 👀 Intro We introduce Uni-Edit, an intelligent image editing task that serves as the first general task for Unified Multimodal Model (UMM) tuning. Unlike conventional mixed multi-task training that suffers from inherent task conflicts and requires complex multi-stage pipelines, Uni-Edit breaks this paradigm. It achieves true mutual reinforcement by improving image… See the full description on the dataset page: https://huggingface.co/datasets/Uni-Edit/Uni-Edit-Train-Data.imageany-to-any1K<n<10K3 likes885 downloads5mo agoHugging Face09taki0112 /hf_dataimage1M<n<10M0 likes811 downloads4mo agoHugging Face10Dodon /plotqa-dataset Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Dodon/plotqa-dataset.image100K<n<1M1 likes777 downloads3y agoHugging Face11Wolowolo /DF_DiFF_FAS_dataset_in_FSFM_FSVFM FSFM / FS-VFM Downstream Datasets Processed downstream fine-tuning datasets for cross-dataset deepfake detection, cross-domain face anti-spoofing, and unseen diffusion-generated face detection used with FSFM and FS-VFM. This archive is intended for research use with the FSFM/FS-VFM release scripts. image1M<n<10M1 likes745 downloads4mo agoHugging Face12JWRoboticsVision /HO-Cap-Datasetimagen<1K0 likes716 downloads1y agoHugging Face13turing-motors /STRIDE-QA-Dataset STRIDE-QA Dataset 📦 Dataset STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning in autonomous driving. Constructed from 100 hours of multi-sensor driving data in Tokyo, it offers 16 M QA pairs over 270 K frames with dense annotations including 3D bounding boxes, segmentation masks, and multi-object tracks. Category Description Object-centric Spatial QA Spatial relations between two… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/STRIDE-QA-Dataset.imagevisual-question-answering100K<n<1M9 likes692 downloads9mo agoHugging Face14zpschang /PIG-Nav-Dataset-PretrainThe pretraining dataset for our paper PIG-Nav: Key Insights for Pretrained Image-Goal Navigation Models. Description of the dataset: The pretraining dataset include GoStanford, RECON, CoryHall, Berkeley DeepDrive, SCAND, TartanDrive, and SACSoN. Please follow our github repo https://github.com/zpschang/PIG-Nav for detailed use. image10M<n<100M1 likes533 downloads1y agoHugging Face15onepiece1999 /correctvla_dataimage1M<n<10M0 likes527 downloads1y agoHugging Face16wjn922 /Recap-Datacomp-1B_tars_part7Final size: 7,236,721, samples per tar: 10000 image1M<n<10M0 likes462 downloads1y agoHugging Face17arimalabs /waifu-preprocessed-datasetimage10M<n<100M1 likes449 downloads11mo agoHugging Face18CodeGoat24 /UnifiedReward-2.0-T2X-score-data Dataset Summary UnifiedReward-2.0-T2X-score-data is added for our UnifiedReward-2.0-qwen-[3b/7b/32b/72b] training. This dataset enables UnifiedReward-2.0 introducing several new capabilities: Pairwise scoring for image and video generation assessment on Alignment, Coherence, Style dimensions. Pointwise scoring for image and video generation assessment on Alignment, Coherence/Physics, Style dimensions. Welcome to try the latest version, and the inference code is available at… See the full description on the dataset page: https://huggingface.co/datasets/CodeGoat24/UnifiedReward-2.0-T2X-score-data.image100K<n<1M0 likes390 downloads1y agoHugging Face19ductai199x /image-manipulation-dataset-compilationimage10K<n<100K0 likes338 downloads2y agoHugging Face20Leonardo6 /datacomp10mimage1M<n<10M0 likes332 downloads1y agoHugging Face21cornuHGF /datacomp-medium-12mimage10M<n<100M2 likes323 downloads1y agoHugging Face22DivisonOfficer /Pixel-aligned_RGB-NIR_stereo_dataset Pixel-aligned RGB-NIR Stereo Imaging and Dataset for Robot Vision CVPR 2025Jinnyeong Kim, Seung-Hwan BaekPOSTECH[arXiv] • [Code] • [Video] • [Dataset on HuggingFace] Overview This repository provides the code and dataset accompanying our CVPR 2025 paper: "Pixel-aligned RGB-NIR Stereo Imaging and Dataset for Robot Vision" We propose a novel robotic vision system equipped with two pixel-aligned RGB-NIR stereo cameras and a LiDAR sensor mounted on a mobile robot. Our… See the full description on the dataset page: https://huggingface.co/datasets/DivisonOfficer/Pixel-aligned_RGB-NIR_stereo_dataset.image10K<n<100K1 likes321 downloads2y agoHugging Face23notawhalel /cluster_VisNIR_data_copiesimage100K<n<1M1 likes317 downloads3mo agoHugging Face24RichardErkhov /The_Million_Song_Datasetimage1M<n<10M0 likes302 downloads2y agoHugging Face25ismatsamadov /azerbaijan-court-data Azerbaijan Court System Dataset The most comprehensive open dataset of Azerbaijan's judicial system — 1.64 million structured records and 1.54 million court decision PDFs (~160 GB) covering court decisions, active cases, scheduled hearings, court registries, judges, lawyers, and mediator organizations. Built for AI engineers, legal tech startups, and researchers who need real-world legal data at scale. Quick Start Load with Hugging Face datasets from datasets… See the full description on the dataset page: https://huggingface.co/datasets/ismatsamadov/azerbaijan-court-data.imagetext-classification1M<n<10M2 likes270 downloads6mo agoHugging Face26MelosY /TextMonkey_Dataimage10K<n<100K4 likes264 downloads2y agoHugging Face27qihoo360 /VRF-datasetsimage100K<n<1M0 likes263 downloads6mo agoHugging Face28mchelali /forbin_dataset Forbin Dataset: A collection of historical photographs with archival metadata This repository hosts the Forbin Dataset, a large-scale collection of historical photographs taken or collected by Victor Forbin (1868–1947). This HuggingFace dataset version provides: COCO-style annotations (segmentation polygons) Archival metadata (Box ID, description, notes, dates when available) A lightweight explorer interface (HTML/JS) to preview images and annotations:… See the full description on the dataset page: https://huggingface.co/datasets/mchelali/forbin_dataset.imageobject-detection100K<n<1M1 likes253 downloads10mo agoHugging Face29gsarch /vigorl_datasets ViGoRL Datasets This repository contains the official datasets associated with the paper "Grounded Reinforcement Learning for Visual Reasoning (ViGoRL)", by Gabriel Sarch, Snigdha Saha, Naitik Khandelwal, Ayush Jain, Michael J. Tarr, Aviral Kumar, and Katerina Fragkiadaki. Dataset Overview These datasets are designed for training and evaluating visually grounded vision-language models (VLMs). Datasets are organized by the visual reasoning tasks described in the ViGoRL… See the full description on the dataset page: https://huggingface.co/datasets/gsarch/vigorl_datasets.imagevisual-question-answering100K<n<1M1 likes246 downloads1y agoHugging Face30SFXX /provision_datacomp_tarsimage100K<n<1M0 likes226 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.