Team Ai
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yanggangu /SMAT-CLIP8 SMAT · CLIP8 The frozen train/dev/test splits for Table 2 of SMAT: Simple and Efficient Merge-Aware Training. Eight image-classification tasks, with 200 development examples per task. Original image bytes, labels and class metadata are preserved. Paper · GitHub & usage This is a benchmark repackaging, not newly collected data. Original dataset rights and usage terms apply; sources and attribution. No additional rights to the images are granted. imageimage-classification100K<n<1M0 likes1.4k downloads10d agoHugging Face02llama-farm /military-labeled-clip Military-Labeled CLIP Crops (DVIDS sourced) Per-object crops extracted from DVIDS military imagery, each accompanied by a Gemini-VLM caption suitable for CLIP fine-tuning or zero-shot evaluation. Files crops/dvids_image_{id}_{class}_{idx}.jpg — 4,844 cropped objects captions.jsonl — per-crop metadata: {crop_path, image_id, class_name, class_id, bbox, caption, image_caption, branch, source, source_url} Classes (12) — Distribution Class Crops… See the full description on the dataset page: https://huggingface.co/datasets/llama-farm/military-labeled-clip.imageimage-classification1K<n<10K1 likes442 downloads4mo agoHugging Face03Voxel51 /getting-started-validation-clip-pred Dataset Card for labeled_validation_predicted_clip This is a FiftyOne dataset with 143 samples. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("TheSteve0/getting-started-validation-clip-pred") # Launch the App session =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/getting-started-validation-clip-pred.imageimage-classificationn<1K0 likes404 downloads2y agoHugging Face04moneyzz432 /CLIP-FMoE-Evaluation CLIP-FMoE Evaluation Data Evaluation datasets used by the CLIP-FMoE repository. Large raw image directories are stored as uncompressed .tar files. This avoids uploading millions of individual image files and makes download/extraction substantially faster. Repository layout clip_benchmark/wds_<dataset>/ Existing CLIP_benchmark WebDataset shards retrieval/docci_iiw/ Metadata + docci_arr_new/images_aar.tar retrieval/dci/ Annotations +… See the full description on the dataset page: https://huggingface.co/datasets/moneyzz432/CLIP-FMoE-Evaluation.image-classification0 likes205 downloads2mo agoHugging Face05AMANMP0007 /military-labeled-clip Military-Labeled CLIP Crops (DVIDS sourced) Per-object crops extracted from DVIDS military imagery, each accompanied by a Gemini-VLM caption suitable for CLIP fine-tuning or zero-shot evaluation. Files crops/dvids_image_{id}_{class}_{idx}.jpg — 4,844 cropped objects captions.jsonl — per-crop metadata: {crop_path, image_id, class_name, class_id, bbox, caption, image_caption, branch, source, source_url} Classes (12) — Distribution Class Crops… See the full description on the dataset page: https://huggingface.co/datasets/AMANMP0007/military-labeled-clip.imageimage-classification1K<n<10K0 likes137 downloads4mo agoHugging Face06Fujitsu-FRE /clide_synthetic_datasets CLIDE Synthetic Image Datasets 📄 Paper • 💻 Code • 🌐 Webpage • 🎥 Video A collection of synthetic images generated by modern text-to-image models, organized by domain and generator. The dataset is designed to support analysis and evaluation of generated-image detection methods under domain and generator shifts. 🗂️ Dataset Structure The dataset contains two visual domains: 💥🚗 Damaged Cars Synthetic images of damaged cars generated by multiple… See the full description on the dataset page: https://huggingface.co/datasets/Fujitsu-FRE/clide_synthetic_datasets.imageimage-classification1K<n<10K1 likes117 downloads8mo agoHugging Face07marco-willi /clip-cues-artifacts CLIP-Cues — frozen input snapshot Cached CLIP features and the checksummed inputs behind "Synthetic Image Detection with CLIP: Understanding and Assessing Predictive Cues" (Willi, Mathys & Graber). This repository holds no images: it is the set of frozen arrays that lets every number and figure in the paper be reproduced on a CPU in minutes, without re-extracting features from 33k photographs. Code: https://github.com/marco-willi/clip-cues · Image datasets: synthclic ·… See the full description on the dataset page: https://huggingface.co/datasets/marco-willi/clip-cues-artifacts.image-classification100K<n<1M0 likes114 downloads2mo agoHugging Face08p14ton /military-labeled-clip Military-Labeled CLIP Crops (DVIDS sourced) Per-object crops extracted from DVIDS military imagery, each accompanied by a Gemini-VLM caption suitable for CLIP fine-tuning or zero-shot evaluation. Files crops/dvids_image_{id}_{class}_{idx}.jpg — 4,844 cropped objects captions.jsonl — per-crop metadata: {crop_path, image_id, class_name, class_id, bbox, caption, image_caption, branch, source, source_url} Classes (12) — Distribution Class Crops… See the full description on the dataset page: https://huggingface.co/datasets/p14ton/military-labeled-clip.imageimage-classification1K<n<10K0 likes103 downloads3mo agoHugging Face09mlnomad /imagenet-1k-224-clip-embeddings ImageNet-1k-224 CLIP Embeddings Pre-computed CLIP image embeddings for every image in mlnomad/imagenet-1k-224. Columns Column Type Description original_index int Row index in the source dataset for cross-referencing label int (0–999) ImageNet class index embedding List[float] L2-normalised CLIP image embedding (768D) Stats Source: mlnomad/imagenet-1k-224 (train split) Total images: 1281167 Embedding dim: 768 CLIP model:… See the full description on the dataset page: https://huggingface.co/datasets/mlnomad/imagenet-1k-224-clip-embeddings.image-classification1M<n<10M0 likes56 downloads7mo agoHugging Face10thaotien /movies_CLIP_ViT-L14 🎬 Movie Frame & Caption Dataset 📖 Introduction This dataset was created from multiple movies across 10 genres, with approximately 3 movies per genre.From each movie, frames were extracted periodically, and AI-generated captions (BLIP) were assigned to each frame.A total of 93,813 frames were extracted. This dataset can be used for tasks such as: Video understanding Multimodal learning (image + text) Image captioning Vision-language retrieval 📂 Data… See the full description on the dataset page: https://huggingface.co/datasets/thaotien/movies_CLIP_ViT-L14.imageimage-classification10K<n<100K0 likes31 downloads1y agoHugging Face11s-emanuilov /coco-clip-vit-l-14 COCO Dataset Processed with CLIP ViT-L/14 Overview This dataset represents a processed version of the '2017 Unlabeled images' subset of the COCO dataset (COCO Dataset), utilizing the CLIP ViT-L/14 model from OpenAI. The original dataset comprises 123K images, approximately 19GB in size, which have been processed to generate 786-dimensional vectors. These vectors can be utilized for various applications like semantic search systems, image similarity assessments, and more.… See the full description on the dataset page: https://huggingface.co/datasets/s-emanuilov/coco-clip-vit-l-14.textimage-classification100K<n<1M2 likes24 downloads3y agoHugging Face12resoajoe /clinical-attire-labels clinical-attire-labels 4,000 Places365 images labelled for clinical attire and role. Every label is machine-generated. scrubs · surgical_gown · patient_gown · lab_coat · street · mask · none plus derived roles: role_staff · role_patient · role_visitor · ppe_mask Why this exists There is no public labelled dataset for medical scrubs, hospital gowns or lab coats. Fashionpedia has none of them among its 46 garment categories; dataset searches return face-mask corpora… See the full description on the dataset page: https://huggingface.co/datasets/resoajoe/clinical-attire-labels.image-classification1K<n<10K0 likes24 downloads1mo agoHugging Face13yeeecheng /OD-CLIP OD-CLIP Training Dataset Training data for OD-CLIP: a degradation-aware CLIP variant that jointly predicts degradation type and LPIPS-calibrated perceptual severity, used for blind image super-resolution. The dataset provides paired ground-truth (GT) and low-quality (LQ) image crops together with per-image degradation metadata, covering four synthetic degradation types (Gaussian blur, Gaussian noise, JPEG compression, and downsampling) at a dense grid of physical severity… See the full description on the dataset page: https://huggingface.co/datasets/yeeecheng/OD-CLIP.image-to-image100K<n<1M0 likes16 downloads2mo agoHugging Face14yashm /clip-insect-sex-data Gryllus bimaculatus Insect Sex Dataset Image dataset for binary classification: male and female. Species Gryllus bimaculatus Structure augmented_data/ male/ female/ Labels male female Source and curation Images were captured by the dataset owner. Augmented variants were generated from owner-captured source images for training. Access and permission terms This dataset is shared for viewing/research reference. Reuse… See the full description on the dataset page: https://huggingface.co/datasets/yashm/clip-insect-sex-data.imageimage-classificationn<1K0 likes15 downloads6mo agoHugging Face15wltjr1007 /cifar100_clipimage-classification10K<n<100K0 likes14 downloads3y agoHugging Face16Mobiusi /Agricultural-Climate-Adaptation-Research-Dataset Agricultural Climate Adaptation Research Dataset Agriculture is currently facing challenges posed by climate change, particularly the increasing impact of drought on crop yields. Existing research data often lacks detailed analysis under specific climate conditions, leading to ineffective agricultural management measures. This dataset aims to fill this gap by including images of farmland drought and vegetation recovery, assisting AI models in researching agriculture's ability to… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Agricultural-Climate-Adaptation-Research-Dataset.textimage-classificationn<1K0 likes11 downloads7mo agoHugging Face17Outerview /urban-climate-green-infrastructure Urban Climate & Green Infrastructure Visual Dataset Rows: 18,107 Build with Outerview Geographic Feature List — browse geographic features available through Outerview Outerview — programmable infrastructure for Earth Developer Documentation — APIs, tools, guides, and examples SDK — build geographic capabilities directly into applications CLI — work with geographic search and spatial data from the terminal MCP — connect geographic search and spatial tools to AI… See the full description on the dataset page: https://huggingface.co/datasets/Outerview/urban-climate-green-infrastructure.imageimage-classification10K<n<100K0 likes10 downloads20h agoHugging Face18wltjr1007 /cifar10_clipimage-classification10K<n<100K0 likes7 downloads3y agoHugging Face19cliptrace-2026 /cliptrace-baseline-datagated CLIPTrace 2026 Baseline Data Participant data for the reproducible CLIPTrace 2026 baseline. The repository mirrors the full-size Imagenette train and validation images used by the baseline and preserves the original class-directory layout. This repository is intended to be public with access gating. Before downloading, users must accept the repository terms and the upstream image-use conditions configured by the organizers on the Hugging Face settings page. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/cliptrace-2026/cliptrace-baseline-data.imageimage-classification10K<n<100K0 likes1 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.