Team Ai
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01DUDE-Framework /Real-UI-Clickboxes RUC: Real UI Clickboxes Click carefully, even when the page is trying to trick you! 👀 Official Hugging Face release for RUC: Real UI Clickboxes, the dataset accompanying our ACL 2026 paper Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces on deceptive UI understanding for web agents. ACL Anthology: https://aclanthology.org/2026.acl-long.310/ PDF: https://aclanthology.org/2026.acl-long.310.pdf DOI: https://doi.org/10.18653/v1/2026.acl-long.310… See the full description on the dataset page: https://huggingface.co/datasets/DUDE-Framework/Real-UI-Clickboxes.imageimage-text-to-text1K<n<10K1 likes683 downloads3mo agoHugging Face02setrsoft /climbing-holds [!IMPORTANT] This dataset is in construction. The current files are raw scans intended for establishing the structure. Using them? Help us clean them up or identify the brands by consulting the CONTRIBUTING.md guide. GUI for contributions https://setrsoft.github.io/holds-dataset-hub/ Or send your files here Climbing Holds 3D dataset (SetRsoft) 📋 Project Overview This dataset is a community-driven open-source dataset of 3D-scanned climbing holds… See the full description on the dataset page: https://huggingface.co/datasets/setrsoft/climbing-holds.3dn<1K1 likes656 downloads6mo agoHugging Face03thaotien /movies_CLIP_ViT-L14 🎬 Movie Frame & Caption Dataset 📖 Introduction This dataset was created from multiple movies across 10 genres, with approximately 3 movies per genre.From each movie, frames were extracted periodically, and AI-generated captions (BLIP) were assigned to each frame.A total of 93,813 frames were extracted. This dataset can be used for tasks such as: Video understanding Multimodal learning (image + text) Image captioning Vision-language retrieval 📂 Data… See the full description on the dataset page: https://huggingface.co/datasets/thaotien/movies_CLIP_ViT-L14.imageimage-classification10K<n<100K0 likes102 downloads1y agoHugging Face04Arabic-Clip /Arabic_3M_5M_ViT-B-16-SigLIP-512 Loading the training split as follows: from datasets import load_dataset ds_train = load_dataset("Arabic-Clip/Arabic_3M_5M_ViT-B-16-SigLIP-512", split="train") ds_train # Dataset({ # features: ['index', 'url', 'en_caption', 'embeddings_en', 'caption_ar'], # num_rows: 2000000 # }) Loading the validation split as follows: from datasets import load_dataset ds_validation = load_dataset("Arabic-Clip/Arabic_3M_5M_ViT-B-16-SigLIP-512", split="validation")… See the full description on the dataset page: https://huggingface.co/datasets/Arabic-Clip/Arabic_3M_5M_ViT-B-16-SigLIP-512.image1M<n<10M0 likes31 downloads2y agoHugging Face05Arabic-Clip /Arabic_3M_5M_ViT-B-16-plus-240 Loading the training split as follows: from datasets import load_dataset ds_train = load_dataset("Arabic-Clip/Arabic_3M_5M_ViT-B-16-plus-240", split="train") ds_train # Dataset({ # features: ['index', 'url', 'en_caption', 'embeddings_en', 'caption_ar'], # num_rows: 200000 # }) Loading the validation split as follows: from datasets import load_dataset ds_validation = load_dataset("Arabic-Clip/Arabic_3M_5M_ViT-B-16-plus-240", split="validation")… See the full description on the dataset page: https://huggingface.co/datasets/Arabic-Clip/Arabic_3M_5M_ViT-B-16-plus-240.image100K<n<1M0 likes25 downloads2y agoHugging Face06Arabic-Clip-Archive /Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-SigLIP-512_validationimage1K<n<10K0 likes19 downloads3y agoHugging Face07Kyzzen /nft-females-generated-cli-test-3-clone nft females generated cli test 3 Automated NFT generation report Number of characters: 10 Model: midorimae-characters-female Guidance: 6.0 LORA scale: 0.95 Total time taken: 0h 0m 51s Average time per prompt: 2.39 seconds Average time per image: 2.74 seconds Average time per image: 5.13 seconds imagen<1K0 likes19 downloads1y agoHugging Face08Arabic-Clip-Archive /Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-plus-240-fulldata-v2DatasetDict({ train: Dataset({ features: ['index', 'embeddings', 'en_caption', 'ar_caption', 'nr_words', 'url'], num_rows: 12166802 }) }) image1M<n<10M0 likes17 downloads3y agoHugging Face09Arabic-Clip-Archive /Arabic_dataset_1M_translated_jsonl_format_ViT-B-16-plus-240This translation done using https://huggingface.co/Helsinki-NLP/opus-mt-en-ar image100K<n<1M0 likes12 downloads3y agoHugging Face10Arabic-Clip-Archive /ImageCaptions-7M-Translations-Arabic-subset-150000image100K<n<1M0 likes11 downloads3y agoHugging Face11Arabic-Clip-Archive /Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-plus-240This dataset repo contains the dataset (CC3M+CC12M+SBU) translated using opus-mt-en-ar and cleaned. Its size about 13M image1M<n<10M0 likes11 downloads3y agoHugging Face12Arabic-Clip-Archive /arabic_dataset_translated_v2_ViT-B-16-SigLIP-512image1M<n<10M0 likes6 downloads3y agoHugging Face13Arabic-Clip-Archive /xtd_arabicimage1K<n<10K0 likes5 downloads3y agoHugging Face14Arabic-Clip-Archive /ccs_synthetic_ar_1M-Arabic_dataset_1M_translated_jsonl_formatimage1M<n<10M0 likes5 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.