Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01builddotai /Egocentric-100Kgated Egocentric-100K is the largest dataset of manual labor. You can visualize the dataset here. Egocentric-100K is state-of-the-art in hand visibility and active manipulation density compared to previous in-the-wild egocentric datasets. The complete 30,000 frame evaluation set is available at Egocentric-100K-Evaluation. Dataset Statistics Attribute Value Total Hours 100,405 Total Frames 10.8 billion Video Clips 2,010,759 Median Clip Length 180.0 seconds Mean… See the full description on the dataset page: https://huggingface.co/datasets/builddotai/Egocentric-100K.text1M<n<10M144 likes181k downloads8mo agoHugging Face02InternRobotics /OmniWorld[ICLR 2026] OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling         🎉NEWS [2026.3.21] 🔥 OmniWorld-Game with Metric Scale is now released! Check out our latest model Pi3X (an enhanced version of Pi3), which leverages this data to achieve better performance! [2026.1.26] 🎉 OmniWorld was accepted by ICLR 2026! [2026.1.7] Update OmniWorld-Game, release RH20T-Robot, RH20T-Human, Ego-Exo4D, EgoDex, Epic-Kitchens. [2025.11.11] The OmniWorld is… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/OmniWorld.imagetext-to-video1B<n<10B97 likes82k downloads6mo agoHugging Face03clip-benchmark /wds_objectnetimage1K<n<10K4 likes81k downloads4y agoHugging Face04agibot-world /AgiBotWorld-Betagated Key Features 🔑 1 million+ trajectories from 100 robots, with a total duration of 2976.4 hours. 100+ real-world scenarios across 5 target domains. Cutting-edge hardware: visual tactile sensors / 6-DoF dexterous hand / mobile dual-arm robots 200+ types of tasks: Contact-rich manipulation Long-horizon planning Multi-robot collaboration 87 types of Atomic Skills, including Tie, OpenJar, Peel, Sweep etc. Your… See the full description on the dataset page: https://huggingface.co/datasets/agibot-world/AgiBotWorld-Beta.textother100M<n<1B83 likes76k downloads1y agoHugging Face05bop-benchmark /hot3d HOT3D-Clips This Hugging Face repository hosts HOT3D-Clips, a set of curated sub-sequences of the HOT3D dataset. Download instructions for HOT3D-Clips and the full HOT3D dataset can be found here. See HOT3D Toolkit for documentation of the data format and for Python utilities (for loading, undistorting fisheye images, rendering using fisheye cameras, etc.). More details can be found in the HOT3D paper and BOP 2024 report. image100K<n<1M8 likes66k downloads1y agoHugging Face06LanguageBind /Open-Sora-Plan-v1.1.0 Annotation We resized the dataset to 1080p for easier uploading. Therefore, the original annotation file might not match the video names. Please refer to this https://github.com/PKU-YuanGroup/Open-Sora-Plan/issues/312#issuecomment-2197312973 Pexels Pexels consists of multiple folders, but each folder exceeds the size limit for Huggingface uploads. Therefore, we divided each folder into 5 parts. You need to merge the 5 parts of each folder first, and then extract each… See the full description on the dataset page: https://huggingface.co/datasets/LanguageBind/Open-Sora-Plan-v1.1.0.text100K<n<1M46 likes64k downloads2y agoHugging Face07nvidia /PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes Dataset Card Dataset Description PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes is a large-scale synthetic dataset of physically-simulated multi-object interaction scenes, generated using NVIDIA Isaac Sim and the PhysX physics engine. It is designed to train and evaluate AI models on physical reasoning, rigid body dynamics, optical flow, depth estimation, and scene understanding. Each clip… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes.image100M<n<1B46 likes40k downloads4mo agoHugging Face08pixparse /cc12m-wds Dataset Card for Conceptual Captions 12M (CC12M) Dataset Summary Conceptual 12M (CC12M) is a dataset with 12 million image-text pairs specifically meant to be used for visionand-language pre-training. Its data collection pipeline is a relaxed version of the one used in Conceptual Captions 3M (CC3M). Usage This instance of Conceptual Captions is in webdataset .tar format. It can be used with webdataset library or upcoming releases of Hugging Face datasets.… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/cc12m-wds.imageimage-to-text10M<n<100M45 likes38k downloads3y agoHugging Face09xwm /WildGUI WildGUI This repository hosts a personally reprocessed annotation release for WildGUI, the dataset introduced by Video2GUI. The original Video2GUI project builds WildGUI from large-scale Internet tutorial videos for GUI agent pretraining. This repository focuses on the open annotation artifacts: the records were regenerated and cleaned following the full annotation workflow, then reformatted to make the data easier to inspect, reuse, and reproduce. It also ships the screenshot… See the full description on the dataset page: https://huggingface.co/datasets/xwm/WildGUI.image10M<n<100M8 likes35k downloads4mo agoHugging Face10amphion /Emilia-Datasetgated Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline. News 🔥 2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.audiotext-to-speech10M<n<100M492 likes30k downloads2y agoHugging Face11agibot-world /AgiBotWorld-Alphagated ⚠️Important Notice !!! Dear Users, The Alpha Dataset has been updated as follows: Frame Loss Data Removal: Several episodes with frame loss issues have been removed. For the complete list of removed episode IDs, please refer to this document. Changes in Episode Count: The updated Alpha Dataset retains the original 36 tasks. The new version has been enriched with additional interactive objects, extending the total duration from 474.12… See the full description on the dataset page: https://huggingface.co/datasets/agibot-world/AgiBotWorld-Alpha.textrobotics10M<n<100M239 likes29k downloads1y agoHugging Face12Lakonik /laion-3m WebDataset shards Each sample contains only: {image_hash}.png {image_hash}.json image1M<n<10M1 likes24k downloads9mo agoHugging Face13InternRobotics /InternData-M1gated InternData-M1 InternData-M1 is a comprehensive embodied robotics dataset containing 244K simulation demonstrations with rich frame-based information including 2D/3D boxes, trajectories, grasp points, and semantic masks, with comprehensive annotations. Your browser does not support the video tag. Changelog 📋 Previous versions remain available in the branch version name. v0.1 : Initial version(26-07-2025) Added simulated/agilex… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/InternData-M1.textother1M<n<10M31 likes21k downloads10mo agoHugging Face14vaishaal /ImageNetV2image10K<n<100K9 likes21k downloads4y agoHugging Face15pixparse /cc3m-wds Dataset Card for Conceptual Captions (CC3M) Dataset Summary Conceptual Captions is a dataset consisting of ~3.3M images annotated with captions. In contrast with the curated style of other image caption annotations, Conceptual Caption images and their raw descriptions are harvested from the web, and therefore represent a wider variety of styles. More precisely, the raw descriptions are harvested from the Alt-text HTML attribute associated with web images. To arrive at the… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/cc3m-wds.imageimage-to-text1M<n<10M60 likes20k downloads3y agoHugging Face16clip-benchmark /wds_imagenet_sketchimage10K<n<100K1 likes20k downloads4y agoHugging Face17ma-xu /fine-t2i Fine-T2I: An Open, Large-Scale, and Diverse Dataset for High-Quality T2I Fine-Tuning [arxiv] by Xu Ma, Yitian Zhang, Qihua Dong, Yun Fu Northeastern Univeristy Please see our [Dataset Explore] to view detailed samples (loading is slow, be patient). 🆕 What's New [2026.02.20]: Fine-T2I reaches the #1 spot among Hugging Face Datasets Trending list ⭐️⭐️⭐️ [2026.02.16]: Fine-T2I tops the Hugging Face Datasets Trending list, reaching the #2 spot and #1… See the full description on the dataset page: https://huggingface.co/datasets/ma-xu/fine-t2i.imageimage-to-text100K<n<1M126 likes20k downloads8mo agoHugging Face18Intelligent-Systems /BEDLAM-depthgated Dataset Mirror of BEDLAM Dataset (Depth Data Subset) Project site: https://bedlam.is.tuebingen.mpg.de/ Please register at project site for additional information and data (Download section) Related Hugging Face dataset mirror: BEDLAM Dataset Information Depth maps (EXR, 32-bit, 3.8TB) Camera ground truth information is not included but can be found in the BEDLAM dataset mirror Image/video data with motion blur is not included but can be found in the BEDLAM dataset… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Systems/BEDLAM-depth.text1M<n<10M0 likes19k downloads7mo agoHugging Face19wendlerc /RenderedTextThis dataset has been created by Stability AI and LAION. This dataset contains 12 million 1024x1024 images of handwritten text written on a digital 3D sheet of paper generated using Blender geometry nodes and rendered using Blender Cycles. The text has varying font size, color, and rotation, and the paper was rendered under random lighting conditions. Note that, the first 10 million examples are in the root folder of this dataset repository and the remaining 2 million are in ./remaining (due… See the full description on the dataset page: https://huggingface.co/datasets/wendlerc/RenderedText.imagetext-to-image10M<n<100M59 likes19k downloads1y agoHugging Face20Vchitect /Vchitect_T2V_DataVerse Vchitect-T2V-Dataverse Vchitect Team1  1Shanghai Artificial Intelligence Laboratory  Paper | Project Page | Data Overview The Vchitect-T2V-Dataverse is the core dataset used to train our text-to-video diffusion model, Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models. It comprises 14 million high-quality videos collected from the Internet, each paired with detailed textual… See the full description on the dataset page: https://huggingface.co/datasets/Vchitect/Vchitect_T2V_DataVerse.texttext-to-video1M<n<10M11 likes19k downloads2y agoHugging Face21mlfoundations /MINT-1T-PDF-CC-2023-14 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.imageimage-to-text1M<n<10M6 likes17k downloads2y agoHugging Face22BLIP3o /BLIP3o-Pretrain-Long-Caption BLIP3o Pretrain Long-Caption Dataset This collection contains 27 million images, each paired with a long (~120 token) caption generated by Qwen/Qwen2.5-VL-7B-Instruct. Download from huggingface_hub import snapshot_download snapshot_download( repo_id="BLIP3o/BLIP3o-Pretrain-Long-Caption", repo_type="dataset" ) Load Dataset without Extracting You don’t need to unpack the .tar archives, use WebDataset support in 🤗datasets instead: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/BLIP3o/BLIP3o-Pretrain-Long-Caption.image10M<n<100M74 likes16k downloads1y agoHugging Face23BAAI-Humanoid /DECO-50 DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter DECO-50 is a bimanual dexterous manipulation dataset with tactile sensing, comprising 50 hours of teleoperated data across 4 scenarios and 28 subtasks, totaling over 5 million frames collected on real dual-arm robots. Dataset Structure DECO-50/ ├── task1/ │ ├── sub_task_1/ │ │ ├── episode_000000/ │ │ │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/BAAI-Humanoid/DECO-50.imagerobotics10M<n<100M7 likes16k downloads8mo agoHugging Face24TencentARC /TimeLens-100K TimeLens-100K 📑 Paper | 💻 Code | 🏠 Project Page | 🤗 Model & Data ✨ Dataset Description TimeLens-100K is a large-scale, diverse, and high-quality training dataset for video temporal grounding. It was proposed in our paper TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs and used for training TimeLens models. The annotation process was conducted using an automated pipeline powered by Gemini-2.5-Pro. 📊 Dataset Statistics Total Videos:… See the full description on the dataset page: https://huggingface.co/datasets/TencentARC/TimeLens-100K.textvideo-text-to-text10K<n<100K7 likes16k downloads10mo agoHugging Face25JoeLeelyf /OVO-Bench OVO-Bench: How Far is Your Video-LLMs from Real-World Online VideO Understanding? 🔥🔥OVO-Bench is accepted by CVPR 2025!🔥🔥 Important Note: Current codebase is modified compared to our initial arXiv paper. We strongly recommend that any use of OVO-Bench should be based on current edition. Introduction 🌟 Three distinct problem-solving modes Backward Tracing: trace back to past events to answer the question.Real-Time Visual Perception:… See the full description on the dataset page: https://huggingface.co/datasets/JoeLeelyf/OVO-Bench.textvideo-text-to-text1K<n<10K10 likes15k downloads2y agoHugging Face26Spawning /pd12m-fullThis dataset is the downloaded variant of Spawning/PD12M. More specifically, this dataset is compatible with webdataset. It was made public after obtaining permission from the original authors of the dataset. You can use the following to explore the dataset with webdataset: import webdataset as wds dataset_path = "pipe:curl -s -f -L https://huggingface.co/datasets/sayakpaul/pd12m-full/resolve/main/{00155..02480}.tar" dataset = ( wds.WebDataset(dataset_path… See the full description on the dataset page: https://huggingface.co/datasets/Spawning/pd12m-full.image10M<n<100M21 likes14k downloads2y agoHugging Face27speechcolab /gigaspeech2gated Dataset Card for GigaSpeech 2 Dataset Description GigaSpeech 2 is an evolving, large-scale, multi-domain, and multilingual ASR corpus focusing on low-resource languages. GigaSpeech 2 raw comprises about 30,000 hours of automatically transcribed speech, across Thai, Indonesian, and Vietnamese. GigaSpeech 2 refine consists of 10,000 hours of Thai, 6,000 hours each for Indonesian and Vietnamese. Repository: https://github.com/SpeechColab/GigaSpeech2 Paper:… See the full description on the dataset page: https://huggingface.co/datasets/speechcolab/gigaspeech2.audioautomatic-speech-recognition10M<n<100M72 likes13k downloads7mo agoHugging Face28pixelprose /pixelprose-shards PixelProse Sharding Tars arXiv | public-released version: pixelprose | JSON-only version: pixelprose-jsons summary Each tar file is approximately 500-600 MB, friendly for fast on-the-fly sampling, filtering, and loading in dataloaders. Each tar file contains triplets of images, text, and JSON files. The *.txt files contain the raw original captions, while the *.json files include all the relevant information. Due to Gemini-1.0 internal version changes during the… See the full description on the dataset page: https://huggingface.co/datasets/pixelprose/pixelprose-shards.image1M<n<10M2 likes13k downloads10mo agoHugging Face29tom-jerry-123 /Physical-AI-AV-US PhysicalAI-AV-SFT Supervised fine-tuning (SFT) dataset for an autonomous-vehicle vision-language waypoint-prediction model. Contains 2,789,773 samples from 150 000 driving scenes (18 seconds per scene, sampled at 1 Hz) recorded in the United States. Format WebDataset — 100 uncompressed .tar shards, each containing pairs of files per sample: Entry Description {key}.png Front-facing wide-angle camera frame (640 × 360 px) {key}.json Metadata (see schema below)… See the full description on the dataset page: https://huggingface.co/datasets/tom-jerry-123/Physical-AI-AV-US.imagerobotics1M<n<10M0 likes12k downloads7mo agoHugging Face30clip-benchmark /wds_imagenet-rimage10K<n<100K0 likes12k downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.