Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01inclusionAI /OpenAoE-2000h Open-AoE — Egocentric Hand Manipulation Dataset Release Roadmap Tier Duration Status nano ~3 h ✅ Released tiny ~100 h ✅ Released full 2000 h 🚧 Uploading Release notes 2026-07-30: Removed samples flagged in PR #1 for camera-intrinsics vs. video-resolution mismatches. 2026-07-31: Uploaded ~323h of data. 2026-08-12: Uploaded ~694h of data. 2026-09-03: Uploaded ~189h of data. Additional data for the full ~2000h release is still… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/OpenAoE-2000h.image39 likes254k downloads6h agoHugging Face02MIC-DKFZ /OpenMind The OpenMind Dataset: A large-scale Head-And-Neck 3D MRI Dataset for self-supervised learning Description The OpenMind Dataset is a large-scale 3D MRI dataset of the head and neck region featuring 114k MRI Images. Its purpose is to provide access of large amounts of 3D medical imaging data to accelerate the development of self-supervised learning methods for 3D medical imaging. This data was pooled from exactly 800 datasets from the OpenNeuro platform and… See the full description on the dataset page: https://huggingface.co/datasets/MIC-DKFZ/OpenMind.3dimage-feature-extractionn<1K56 likes176k downloads2y agoHugging Face03Goku-OpenLab /gpt-image-2-prompts-datasets 🖼️ GPT Image 2 Prompt Dataset 🖼️ The ultimate GPT Image 2 prompt dataset (5GB+). 15,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators. This project is a massive collection of prompts used for OpenAI's GPT Image 2 model and the resulting generated images. The entire dataset exceeds 5GB and contains 15,000+ images, all structured into a comprehensive dataset. Due to… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/gpt-image-2-prompts-datasets.imagetext-to-image10K<n<100K8 likes111k downloads5d agoHugging Face04Open-Bee /Honey-Data-15M Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs [🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code] Introduction We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality. Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Honey-Data-15M.imageimage-text-to-text10M<n<100M120 likes40k downloads7mo agoHugging Face05opendatalab /OmniDocBench OmniDocBench English | 简体中文 OmniDocBench is an evaluation dataset for diverse document parsing in real-world scenarios, with the following characteristics: Diverse Document Types: The evaluation set contains 1651 PDF pages, covering 10 document types, 5 layout types and 5 language types. Coverage includes academic literature, research and financial reports, newspapers, textbooks, exam papers, magazines, handwritten notes, historical documents, and more. Rich Annotations:… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/OmniDocBench.image1K<n<10K113 likes27k downloads4mo agoHugging Face06Goku-OpenLab /nano-banana-pro-prompts-datasets 🖼️ Nano Banana Pro Prompt Dataset 🖼️ The ultimate Nano Banana Pro prompt dataset (6GB+). 26,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators. This project is a massive collection of prompts used for Nano Banana Pro AI image model and the resulting generated images. The entire dataset exceeds 6GB and contains 26,000+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/nano-banana-pro-prompts-datasets.imagetext-to-image10K<n<100K3 likes26k downloads2mo agoHugging Face07OpenGVLab /MVBench MVBench Important Update [18/10/2024] Due to NTU RGB+D License, 320 videos from NTU RGB+D need to be downloaded manually. Please visit ROSE Lab to access the data. We also provide a list of the 320 videos used in MVBench for your reference. We introduce a novel static-to-dynamic method for defining temporal-related tasks. By converting static tasks into dynamic ones, we facilitate systematic generation of video tasks necessitating a wide range of temporal abilities, from… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/MVBench.imagevisual-question-answering1K<n<10K48 likes20k downloads2y agoHugging Face08ComplexDataLab /OpenFake Dataset Card for OpenFake OpenFake is a dataset and benchmark for detecting AI-generated images, with a focus on politically and socially salient content where misinformation risk is highest. It pairs real photographs with synthetic counterparts produced by a wide range of frontier proprietary generators, open-source diffusion models, and community fine-tunes. A separate in-the-wild test set is sourced from Reddit to evaluate detector performance on naturally circulated… See the full description on the dataset page: https://huggingface.co/datasets/ComplexDataLab/OpenFake.imageimage-classification1M<n<10M35 likes17k downloads11d agoHugging Face09liu-lz /OpenPathNet OpenPathNet Dataset This README describes the OpenPathNet dataset (the release referred to as Link 1 in the OpenPathNet project documentation). The dataset is generated by the OpenPathNet toolchain from real-world Miami and Boston urban areas based on OpenStreetMap (OSM), and then simulated with NVIDIA Sionna ray tracing for RF multipath propagation / channel modeling research and AI tasks. The dataset is also carefully cleaned to ensure good building coverage in every scene.… See the full description on the dataset page: https://huggingface.co/datasets/liu-lz/OpenPathNet.3d1K<n<10K1 likes14k downloads10mo agoHugging Face10OpenGVLab /GUI-Odyssey Dataset Card for GUI Odyssey News⭐️ A new and improved version of the GUIOdyssey dataset has been released! 🎉🎉 👉 Please use the latest version and refer to the updated README for the most up-to-date information. We highly recommend using the new version for all training and evaluation! Repository: https://github.com/OpenGVLab/GUI-Odyssey Latest Version of Dataset: hflqf88888/GUIOdyssey Paper: https://arxiv.org/pdf/2406.08451 Introduction GUI Odyssey is… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/GUI-Odyssey.image1K<n<10K26 likes13k downloads1y agoHugging Face11acane2 /Open4DHOI Open-Source Release Manifest Generated from upload_records.json for records with annotation_progress == 4. Required release items video.mp4: source video to publish. obj_init.obj: object mesh to publish. mask_dir/: object masks. human_mask_dir/: human masks. motion/result.pt: reconstructed human motion. motion/hand_pose.npz: SMPL-X hand pose parameters. kp_record_new.json: point annotations. Generated files release_manifest.json: full per-record… See the full description on the dataset page: https://huggingface.co/datasets/acane2/Open4DHOI.image100K<n<1M1 likes9.6k downloads5mo agoHugging Face12bshada /open-schematics Open Schematics Dataset The largest dataset of electronic schematics and PCB layouts on the internet, built as an engineering reference for schematic and PCB layout work. It's a self-growing, autonomous dataset that continuously scans the web for new engineering designs and updates itself accordingly. Dataset Description Each record corresponds to one schematic file and includes the raw source, rendered images, structured metadata, and all associated PCB files… See the full description on the dataset page: https://huggingface.co/datasets/bshada/open-schematics.imagetext-generation10K<n<100K192 likes7.6k downloads3mo agoHugging Face13finedet /openimages FineOpenimages — Open Images V7 boxed subset in the unified detection format Source: official Open Images bbox CSVs + CVDF-hosted image tars (open-images-dataset S3 bucket). Converted by the finedet project into a unified, AutoTrain-compatible layout: image / width / height / objects{bbox, category} with COCO-format [x, y, w, h] boxes in absolute pixels. Boxes are clipped to the image and empty boxes dropped; category ids are densified per the category tables below.… See the full description on the dataset page: https://huggingface.co/datasets/finedet/openimages.imageobject-detection1M<n<10M0 likes5.2k downloads2mo agoHugging Face14Asklv /OpenMath-Vision-CoT-10kimage10K<n<100K1 likes5k downloads9mo agoHugging Face15Goku-OpenLab /messy-prompt-datasets 🎨 Messy Prompt Dataset 🎨 A mixed collection of AI image prompts (500+). A bit of everything — raw and uncurated. Truly open source: No login, no ads, no redirection. Just pure data for AI creators. This project is a growing collection of diverse image generation prompts gathered from social platforms like Twitter/X. The entire dataset contains 500+ images, all structured into a comprehensive dataset. Due to GitHub's limitations with large file storage, the full dataset… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/messy-prompt-datasets.image1K<n<10K1 likes4.7k downloads2mo agoHugging Face16nebula /OpenSDI_trainThis repository contains the OpenSDI training dataset, presented in the paper OpenSDI: Spotting Diffusion-Generated Images in the Open World. Code: https://github.com/iamwangyabin/OpenSDI imageimage-classification100K<n<1M2 likes4.2k downloads2y agoHugging Face17opencompass /TextEdit TextEdit: A High-Quality, Multi-Scenario Text Editing Benchmark for Generation Models Danni Yang, Sitao Chen, Changyao Tian If you find our work helpful, please give us a ⭐ or cite our paper. See the InternVL-U technical report appendix for more details. 🎉 News [2026/03/06] TextEdit benchmark released. [2026/03/06] Evaluation code and initial baselines released. [2026/03/06] Leaderboard updated with latest models. 📖 Introduction… See the full description on the dataset page: https://huggingface.co/datasets/opencompass/TextEdit.imageimage-to-image1K<n<10K9 likes3.9k downloads7mo agoHugging Face18jake-song-openmmo /onlinerpg-assets3d0 likes3.9k downloads5h agoHugging Face19openbmb /VisRAG-Ret-Train-Synthetic-data Dataset Description This dataset is the synthetic part of the training set of VisRAG it includes 239,358 Query-Document (Q-D) Pairs from a synthetic dataset made up of pages from web-crawled PDF documents and augmented with VLM-generated (GPT-4o) pseudo-queries. Our training data is organized with a batch size of 128, ensuring that all data within the same batch comes from the same dataset. Name Source Description # Pages Textbooks https://openstax.org/ College-level… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/VisRAG-Ret-Train-Synthetic-data.image100K<n<1M20 likes3.9k downloads2y agoHugging Face20data-is-better-together /open-image-preferences-v1 Open Image Preferences Prompt: Anime-style concept art of a Mayan Quetzalcoatl biomutant, dystopian world, vibrant colors, 4K. Image 1 Image 2 Prompt: 8-bit pixel art of a blue knight, green car, and glacier landscape in Norway, fantasy style, colorful and detailed. Image 1… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/open-image-preferences-v1.imagetext-to-image1K<n<10K31 likes3.8k downloads2y agoHugging Face21Vishwa-307 /OpenLaneV2image10K<n<100K0 likes3.8k downloads6mo agoHugging Face22Goku-OpenLab /open-models-prompt-datasets 🖼️ Open Models Prompt Dataset 🖼️ The ultimate open models image prompt dataset (10GB+). 5400+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators. This project is a massive collection of prompts used for various open-source AI image models and the resulting generated images. The entire dataset exceeds 10GB and contains 5400+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/open-models-prompt-datasets.image1K<n<10K2 likes3.5k downloads3mo agoHugging Face23nebula /OpenSDI_test OpenSDI: Spotting Diffusion-Generated Images in the Open World This dataset is designed to address the OpenSDI challenge: spotting diffusion-generated images in realistic, open-world scenarios. It is described in the paper: Project Page: https://iamwangyabin.github.io/OpenSDI/ OpenSDID Dataset Highlights: User Diversity: Simulates a wide range of user intentions and creative styles using diverse text prompts generated by VLMs. Model Innovation: Includes images from multiple… See the full description on the dataset page: https://huggingface.co/datasets/nebula/OpenSDI_test.imageimage-classification100K<n<1M1 likes3k downloads2y agoHugging Face24joonyeongs /cosmos3-ap-openarm-wam-robot-ours-aug-v2v-b256-8kimage1K<n<10K0 likes2.9k downloads20d agoHugging Face25vector-institute /open-pmc-18m OPEN-PMC Arxiv: Arxiv     |     Code: Open-PMC Github     |     Model Checkpoint: Hugging Face Dataset Summary This dataset consists of image-text pairs extracted from medical papers available on PubMed Central. It has been curated to support research in medical image understanding, particularly in natural language processing (NLP) and computer vision tasks related to medical imagery. The dataset includes: Extracted images from research articles.… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/open-pmc-18m.image10M<n<100M7 likes2.9k downloads5mo agoHugging Face26austinpatel /libero_gen_goal_chain_train_openpiimage1M<n<10M0 likes2.8k downloads5mo agoHugging Face27Lichess /chess-openings Dataset Card for Lichess Puzzles Dataset Description 3704 chess openings with their Encyclopaedia of Chess Openings (ECO) classification. Dataset Creation Creating this dataset is detailed in its original GitHub repository. Updates to the original repo will also be reflected in this version. Dataset last updated on May 20th, 2026. Dataset Usage Using the datasets library: from datasets import load_dataset dset =… See the full description on the dataset page: https://huggingface.co/datasets/Lichess/chess-openings.image1K<n<10K24 likes2.8k downloads5mo agoHugging Face28joonyeongs /cosmos3-ap-openarm-wam-robot-ours-aug-v2v-b256-4kimage1K<n<10K0 likes2.7k downloads20d agoHugging Face29Open-Bee /Bee-Training-Data-Stage2 Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs [🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code] Introduction We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality. Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Bee-Training-Data-Stage2.imageimage-to-text10M<n<100M6 likes2.7k downloads7mo agoHugging Face30LAYEK-143 /Open-Pixel-1T 🌌 Open-Pixel-1T (Visual Atlas) A Large-Scale, High-Entropy Synthetic Image Dataset for Foundational Pre-Training 📑 Dataset Summary Open-Pixel-1T is a monumental open-source initiative designed to create a "Visual Atlas" of stochastic imagery. Unlike traditional datasets scraped from social media which contain inherent human bias, Open-Pixel-1T is constructed using high-entropy random seeds to generate unique, diverse visual signals. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/LAYEK-143/Open-Pixel-1T.imagetext-to-image10M<n<100M8 likes2.7k downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.