Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nebula /OpenSDI_test OpenSDI: Spotting Diffusion-Generated Images in the Open World This dataset is designed to address the OpenSDI challenge: spotting diffusion-generated images in realistic, open-world scenarios. It is described in the paper: Project Page: https://iamwangyabin.github.io/OpenSDI/ OpenSDID Dataset Highlights: User Diversity: Simulates a wide range of user intentions and creative styles using diverse text prompts generated by VLMs. Model Innovation: Includes images from multiple… See the full description on the dataset page: https://huggingface.co/datasets/nebula/OpenSDI_test.imageimage-classification100K<n<1M1 likes3k downloads2y agoHugging Face02paulpacaud /rlbenchfail_test_dataset Guardian: RLBench-Fail Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_test_dataset.tabularvisual-question-answering1K<n<10K0 likes892 downloads8mo agoHugging Face03Yunncheng /Mirage-Test 🌊 Mirage-Test Dataset Mirage-Test is a modern test-only dataset for benchmarking AI-generated image detection models. It contains real (0_real) and fake (1_fake) images across five distinct content domains, designed to evaluate generalization across diverse visual semantics. The fake images are generated using state-of-the-art generative models specifically optimized for perceptual realism and visual fidelity. 📌 This dataset is for evaluation only. No training split is… See the full description on the dataset page: https://huggingface.co/datasets/Yunncheng/Mirage-Test.imageimage-classification10K<n<100K3 likes783 downloads10mo agoHugging Face04paulpacaud /ur5fail_test_dataset Guardian Failure Detection Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Guardian introduces an automated failure generation approach that procedurally perturbs successful robot trajectories to produce diverse planning failures and execution failures, each… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/ur5fail_test_dataset.tabularvisual-question-answeringn<1K1 likes203 downloads8mo agoHugging Face05jrw2989 /testing-goldstandard-cuthill Dataset Card for Curated Gold Standard Hoyal Cuthill Dataset Dataset Description Dorsal full body images of subspecies of Heliconius erato and Heliconius melpomene (18 subspecies total). There are 960 images with 320 specimens (3 images of each specimen: Original/ Bird transformed/ Butterfly transformed) The original images are low-resolution RGB photographs (photographs were "cropped and resized to a height of 64 pixels (maintaining the original image aspect ratio and… See the full description on the dataset page: https://huggingface.co/datasets/jrw2989/testing-goldstandard-cuthill.imageimage-classificationn<1K0 likes175 downloads1y agoHugging Face06Futuremark /winml-test-setgated WinML Test Set Dataset Summary WinML Test Set is an evaluation‑only collection for validating model accuracy and stability on Windows ML / DirectML / ONNX Runtime pipelines. It aggregates several permissively‑licensed sources and harmonizes schema for reproducible, regression‑grade testing across backends and versions. Not intended for training. Intended Use Accuracy and regression benchmarking of Windows ML / DirectML / ONNX Runtime pipelines.… See the full description on the dataset page: https://huggingface.co/datasets/Futuremark/winml-test-set.imageimage-classification1K<n<10K0 likes124 downloads7mo agoHugging Face07paulpacaud /bdv2fail_test_dataset Guardian: BridgeDataV2-Fail Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data derived from the BridgeDataV2 real-robot dataset for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Failures are produced by an automated pipeline that perturbs successful real-robot trajectories offline (without re-executing actions)… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/bdv2fail_test_dataset.tabularvisual-question-answering1K<n<10K0 likes104 downloads8mo agoHugging Face08OmarK211 /photo-test Anime vs. Live Action Film Image Classification OmarK211/photo-test A binary image classification dataset designed to distinguish between Anime (Label 0) and Live Action (Label 1) film frames/imagery. Images are prepared as standardized square RGB files with multi-pass synthetic training variants. Source and task The images were manually imported from the web into my local computer. Then manually uploaded to Google Colab Preparation source: 24-679 Image Data… See the full description on the dataset page: https://huggingface.co/datasets/OmarK211/photo-test.imageimage-classificationn<1K0 likes94 downloads24d agoHugging Face09ZihengZ /testset Dataset Card for TreeOfLife-10M Captions This dataset consists of generated captions, Wikipedia-derived descriptions and format examples for the TreeOfLife-10M. These captions were generated using InternVL3-38B based on biological contexts that help the model generate more accurate captions. It was used to train BioCAP, a CLIP-based model. Dataset Details This dataset is comprised of captions for the images in TreeOfLife-10M that were generated using InternVL3 38B.… See the full description on the dataset page: https://huggingface.co/datasets/ZihengZ/testset.textimage-classification1M<n<10M0 likes86 downloads1y agoHugging Face10Krows7 /test Tiny overlapping ImageFolder demo A tiny synthetic repository for testing Hugging Face Dataset Viewer. 12 lossless WebP images, each exactly 128×128. full: 8 train, 2 validation, 2 test. core: 4 train, 1 validation, 1 test. core is a subset of full. Images exist once in images/. manifest.csv is the canonical manifest. splits/ contains lightweight source indices. Root full_*.csv and core_*.csv are materialized metadata files used by ImageFolder. Expected Viewer… See the full description on the dataset page: https://huggingface.co/datasets/Krows7/test.imageimage-classificationn<1K0 likes73 downloads3mo agoHugging Face11FatimahEmadEldin /GenAI-RealEstate-TestSet 🏙️ GenAI Real Estate Test Set (Track B) Dataset for the MenaML Winter School 2026 Challenge. 📊 Dataset Structure This dataset contains 1,000 images split evenly between: Authentic: Real estate photography from the Places365 dataset. Manipulated: Synthetically generated deepfake artifacts (Inpainting, Diffusion Noise, GAN Grids). 🕵️ How to Use This dataset is designed for testing forensic detection models. imageimage-classification1K<n<10K0 likes57 downloads9mo agoHugging Face12abhisheklalwani96 /diagram-eval-access-test Diagram Evaluation Access Test This public one-image dataset tests the external-access workflow planned for the Diagram Evaluation project. It is not a research dataset release. Dataset structure The repository uses Hugging Face's ImageFolder layout: data/ train/ metadata.csv sample-diagram.png The metadata records the image's provenance, license, and intended use. A production release can use the same contract with sharded Parquet or WebDataset files.… See the full description on the dataset page: https://huggingface.co/datasets/abhisheklalwani96/diagram-eval-access-test.imageimage-classificationn<1K0 likes47 downloads4d agoHugging Face13thxplz /HowToEat-test HowToEat: Hand-Object Interaction and Eating Action in Eating Scenarios HowToEat is an image dataset for analysing eating behaviour. It provides: Hand-object interaction + eating face detection (hand_object_detection): 95,190 images with 190,333 hand instances (box, left/right side, contact state, and the box and category of the held object) and 151,620 face instances (box, eating / not eating). Eating action recognition (eating_recognition): 6,280 manually labelled faces… See the full description on the dataset page: https://huggingface.co/datasets/thxplz/HowToEat-test.imageobject-detectionn<1K0 likes45 downloads7h agoHugging Face14BuildNg /astrobridge-yse-test-dataset-v2 AstroBridge YSE external test dataset v2 This dataset contains 266 object-disjoint, spectroscopically labeled YSE DR1 transients. The broad-class counts are SN II: 71, SN Ia: 180, SN Ibc: 15. V2 shortens the YSE forced-photometry time coverage to resemble the alert-photometry coverage of the AstroBridge BTS training dataset. For each object, it retains the smallest inclusive time interval that contains every positive measurement with flux/uncertainty at least 5 and at least five… See the full description on the dataset page: https://huggingface.co/datasets/BuildNg/astrobridge-yse-test-dataset-v2.imageimage-classificationn<1K0 likes43 downloads1mo agoHugging Face15Robo531 /ai-detector-benchmark-test-data 🎯 AI Detector Benchmark Test Dataset A comprehensive benchmark dataset for testing AI image detection models. 📊 Dataset Summary Total Images: 700 AI-Generated: 250 images (from 5 different generators) Real Images: 450 images (from 9 diverse datasets) Perfect for: ✅ Testing AI detection models ✅ Creating leaderboards ✅ Comparing model performance ✅ Benchmarking new approaches 🤖 AI Generators Included Generator Images Accuracy Baseline FLUX… See the full description on the dataset page: https://huggingface.co/datasets/Robo531/ai-detector-benchmark-test-data.imageimage-classificationn<1K0 likes37 downloads7mo agoHugging Face16Maxscha /test Zoo Animal Re-Identification Dataset A dataset for animal re-identification with 2,705 body images, 1,192 face crops, and 6 configurations. Configurations 1. face_and_body Individual frames with both face and body crops. Features: date: Date of capture (YYYY-MM-DD) time: Time of capture (HH:MM:SS) class: Animal name video: Source video filename frame_number: Frame number in video camera: Camera ID face_image: Cropped face image body_image: Cropped body/full… See the full description on the dataset page: https://huggingface.co/datasets/Maxscha/test.imageobject-detection1K<n<10K0 likes23 downloads8mo agoHugging Face17harpreetsahota /VGGSound-EG2-100-test Dataset Card for t1-vggsound-eg2-xmodal This is a FiftyOne dataset with 100 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("harpreetsahota/VGGSound-EG2-100-test") # Launch the App session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/VGGSound-EG2-100-test.textimage-classificationn<1K0 likes23 downloads22h agoHugging Face18BuildNg /astrobridge-yse-test-dataset AstroBridge YSE external test dataset This dataset contains 266 spectroscopically labeled YSE DR1 transients that do not overlap the final AstroBridge BTS training dataset by normalized TNS identity or a two-arcsecond transient-coordinate match. The broad-class counts are SN II: 71, SN Ia: 180, SN Ibc: 15. Each row has at least five valid ZTF g and five valid ZTF r observations, ordered by time. The light-curve flux arrays are in nJy. The ATCAT arrays retain SNANA FLUXCAL at… See the full description on the dataset page: https://huggingface.co/datasets/BuildNg/astrobridge-yse-test-dataset.imageimage-classificationn<1K0 likes22 downloads1mo agoHugging Face19ash12321 /ai-detector-benchmark-test-data 🎯 AI Detector Benchmark Test Dataset A comprehensive benchmark dataset for testing AI image detection models. 📊 Dataset Summary Total Images: 700 AI-Generated: 250 images (from 5 different generators) Real Images: 450 images (from 9 diverse datasets) Perfect for: ✅ Testing AI detection models ✅ Creating leaderboards ✅ Comparing model performance ✅ Benchmarking new approaches 🤖 AI Generators Included Generator Images Accuracy Baseline FLUX… See the full description on the dataset page: https://huggingface.co/datasets/ash12321/ai-detector-benchmark-test-data.imageimage-classificationn<1K0 likes19 downloads9mo agoHugging Face20mkzheng /test Overview lalalalalalalalanewwwwwwhhhhjyfjyf This project aims to introduce and release a comprehensive food image dataset designed specifically for computer vision tasks, particularly food recognition, classification, and nutritional analysis. We hope this dataset will provide a reliable resource for researchers and developers to advance the field of food AI. By publishing on Hugging Face, we expect to foster community collaboration and accelerate innovation in applications… See the full description on the dataset page: https://huggingface.co/datasets/mkzheng/test.imageimage-classification10K<n<100K0 likes16 downloads1y agoHugging Face21Quazitron420 /video-dataset-pre1_test Video Dataset - pre1_test Dataset Description This dataset contains video frames extracted from annotated video segments, along with annotations, transcriptions, and corresponding video clips. Dataset Structure frames/ — extracted frames grouped by role (start, middle, end) segments/ — video clips for each annotation interval annotations/ — original JSON annotation transcriptions/ — transcription files (full_transcription.txt + per segment) dataset.csv —… See the full description on the dataset page: https://huggingface.co/datasets/Quazitron420/video-dataset-pre1_test.imageimage-classificationn<1K0 likes14 downloads11mo agoHugging Face22Quazitron420 /video-dataset-test2040 Video Dataset - test2040 Dataset Description This dataset contains video frames extracted from annotated video segments, along with annotations, transcriptions, and corresponding video clips. Dataset Structure frames/ — extracted frames grouped by role (start, middle, end) segments/ — video clips for each annotation interval annotations/ — original JSON annotation transcriptions/ — transcription files (full_transcription.txt + per segment) dataset.csv —… See the full description on the dataset page: https://huggingface.co/datasets/Quazitron420/video-dataset-test2040.imageimage-classificationn<1K0 likes13 downloads11mo agoHugging Face2313point5 /lineex-test LineEX Test Split This dataset is a structured Hugging Face mirror of the test split released with the LineEX paper, "LineEX: Data Extraction from Scientific Line Charts" (WACV 2023). Repo: 13point5/lineex-test Split: test Rows: 20,000 Source format: PNG images plus COCO-style chart-element and line annotations Columns image_id file_name image width height data_type chart_elements lines chart_elements fields annotation_id category_id category_name… See the full description on the dataset page: https://huggingface.co/datasets/13point5/lineex-test.imageimage-classification10K<n<100K0 likes13 downloads7mo agoHugging Face24UngLong /radiology-test-v2 Radiology Test v2 Evaluation dataset for the Radiology agent of an AI Medical Department, paired with UngLong/radiology-ready-v2 (training set). CT images are in UngLong/openm3chest-npy-v2. ⚠️ Before using: CT scans for this test set must be uploaded to openm3chest-npy-v2 first. See scripts/test_npy_needed.txt (154 scan keys) for the list to upload via build_npy_hub.py. Dataset Summary Total rows 154 Screening rows 74 (all 8 screening tasks)… See the full description on the dataset page: https://huggingface.co/datasets/UngLong/radiology-test-v2.textvisual-question-answeringn<1K0 likes12 downloads4mo agoHugging Face25dbabnigg /botanical-vision-test Botanical Vision Fine-grained flowering-plant classification dataset: 495 research-grade iNaturalist photos across 5 species (all flowering plants with at least 2,000 observations). Built for Advanced Computer Vision (UChicago ADSP 32023). Splits split images train 345 val 75 test 75 Split is stratified within each species (70/15/15). Exact and cross-species duplicate images were removed before splitting. Fields image — the… See the full description on the dataset page: https://huggingface.co/datasets/dbabnigg/botanical-vision-test.imageimage-classificationn<1K0 likes12 downloads3mo agoHugging Face26Quazitron420 /video-dataset-anonim_test Video Dataset - anonim_test Dataset Description This dataset contains video frames extracted from annotated video segments, along with annotations, transcriptions, and corresponding video clips. Dataset Structure frames/ — extracted frames (first frame from each segment) segments/ — video clips for each annotation interval annotations/ — original JSON annotation transcriptions/ — transcription files (full_transcription.txt + per segment) dataset.csv — mapping… See the full description on the dataset page: https://huggingface.co/datasets/Quazitron420/video-dataset-anonim_test.imageimage-classificationn<1K0 likes11 downloads10mo agoHugging Face27Mouwiya /image-in-Words400_DOCCI_Testimageimage-to-textn<1K0 likes10 downloads2y agoHugging Face28mlrvv /edgeimpulse-test-image-classification Edgeimpulse Test Image Classification This dataset is an integration-test fixture for Edge Impulse's "Import from Hugging Face" flow. Structure Splits: train, validation, test Main fields: image, label Extra metadata columns (from metadata.csv): source_split source_file source_stem source_path Important note Label source mode: source-metadata. imageimage-classificationn<1K0 likes9 downloads7mo agoHugging Face29ChonJohn171105 /openm3chest-cardiology-test-200 OpenM3Chest Cardiology Test 200 This dataset is a sampled test subset prepared from UngLong/openm3chest-labels. Source Source repo: UngLong/openm3chest-labels Source split: test Subsets/tasks: ['CVD_diagnosis', 'CVD_mortality'] Sampling Number of samples: 200 Sampling mode: primary_ratio Primary subset: CVD_diagnosis Positive label: 1 Positive ratio: 0.5 Seed: 42 Columns keys: CT scan key / series identifier from source dataset.… See the full description on the dataset page: https://huggingface.co/datasets/ChonJohn171105/openm3chest-cardiology-test-200.textquestion-answeringn<1K0 likes9 downloads4mo agoHugging Face30egrace479 /fish-vista-testgated Dataset Card for Fish-Visual Trait Analysis (Fish-Vista) Dataset Deetails Dataset Description The Fish-Visual Trait Analysis (Fish-Vista) dataset is a large, annotated collection of 60K fish images spanning 1900 different species; it supports several challenging and biologically relevant tasks including species classification, trait identification, and trait segmentation. These images have been curated through a sophisticated data processing… See the full description on the dataset page: https://huggingface.co/datasets/egrace479/fish-vista-test.documentimage-classification100K<n<1M0 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.