Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jmhessel /newyorker_caption_contest Dataset Card for New Yorker Caption Contest Benchmarks Dataset Summary See capcon.dev for more! Data from: Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption Contest @inproceedings{hessel2023androids, title={Do Androids Laugh at Electric Sheep? {Humor} ``Understanding'' Benchmarks from {The New Yorker Caption Contest}}, author={Hessel, Jack and Marasovi{\'c}, Ana and Hwang, Jena D. and Lee, Lillian and… See the full description on the dataset page: https://huggingface.co/datasets/jmhessel/newyorker_caption_contest.imageimage-to-text100K<n<1M76 likes24k downloads3y agoHugging Face02quarterturn /danbooru-1024-eq-captioned Danbooru 1024 e/q Captioned Dataset 59,495 high-resolution (1024px) anime-style images from Danbooru's explicit and questionable rated pools. Each image includes comprehensive JSON captions generated via MiniMax-M3 with structured per-character state-of-dress inventories, camera notes, mood palettes, and post-processing detections. Directory Structure danbooru-1024-eq-captioned.parquet <- consolidated metadata manifest originals/ <-… See the full description on the dataset page: https://huggingface.co/datasets/quarterturn/danbooru-1024-eq-captioned.image10K<n<100K6 likes22k downloads2mo agoHugging Face03BLIP3o /BLIP3o-Pretrain-Long-Caption BLIP3o Pretrain Long-Caption Dataset This collection contains 27 million images, each paired with a long (~120 token) caption generated by Qwen/Qwen2.5-VL-7B-Instruct. Download from huggingface_hub import snapshot_download snapshot_download( repo_id="BLIP3o/BLIP3o-Pretrain-Long-Caption", repo_type="dataset" ) Load Dataset without Extracting You don’t need to unpack the .tar archives, use WebDataset support in 🤗datasets instead: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/BLIP3o/BLIP3o-Pretrain-Long-Caption.image10M<n<100M74 likes16k downloads1y agoHugging Face04google-research-datasets /conceptual_captions Dataset Card for Conceptual Captions Dataset Summary Conceptual Captions is a dataset consisting of ~3.3M images annotated with captions. In contrast with the curated style of other image caption annotations, Conceptual Caption images and their raw descriptions are harvested from the web, and therefore represent a wider variety of styles. More precisely, the raw descriptions are harvested from the Alt-text HTML attribute associated with web images. To arrive at the… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/conceptual_captions.imageimage-to-text1M<n<10M111 likes13k downloads2y agoHugging Face05BLIP3o /BLIP3o-Pretrain-Short-Caption BLIP3o Pretrain Short-Caption Dataset This collection contains 5 million images, each paired with a short (~20 token) caption generated by Qwen/Qwen2.5-VL-7B-Instruct. Download from huggingface_hub import snapshot_download snapshot_download( repo_id="BLIP3o/BLIP3o-Pretrain-Short-Caption", repo_type="dataset" ) Load Dataset without Extracting You don’t need to unpack the .tar archives, use WebDataset support in 🤗datasets instead: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/BLIP3o/BLIP3o-Pretrain-Short-Caption.image1M<n<10M10 likes9.3k downloads1y agoHugging Face06lambda /pokemon-blip-captionsgated Notice of DMCA Takedown Action We have received a DMCA takedown notice from The Pokémon Company International, Inc. In response to this action, we have taken down the dataset. We appreciate your understanding. imagetext-to-imagen<1K314 likes8k downloads3y agoHugging Face07CaptionEmporium /pexels-568k-internvl2 Dataset Card for pexels-568k-internvl2 Dataset Summary This is 567,573 synthetic captions for the images found in ptx0/photo-concept-bucket. The captions were produced using OpenGVLab/InternVL2-40B-AWQ. The dataset was grounded for captioning using the tags originally listed. Languages The text is in English, but occasionally text in images in other languages is transcribed. Intended Usage Training text-to-image models and other machine learning… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/pexels-568k-internvl2.imagetext-to-image100K<n<1M21 likes6.5k downloads2y agoHugging Face08jxie /coco_captions Dataset Card for "coco_captions" More Information needed image100K<n<1M18 likes6.3k downloads3y agoHugging Face09lmms-lab-encoder /COCO-Caption Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of COCO-Caption-2014-version. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @misc{lin2015microsoft, title={Microsoft COCO: Common Objects in Context}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/COCO-Caption.image10K<n<100K15 likes4.1k downloads3y agoHugging Face10lmms-lab-encoder /COCO-Caption2017 Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of COCO-Caption-2017-version. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @misc{lin2015microsoft, title={Microsoft COCO: Common Objects in Context}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/COCO-Caption2017.image10K<n<100K24 likes3.9k downloads3y agoHugging Face11hanlincs /InternVL-SA1B-Caption-WebDatasetThis repo contains the recaptioned SA1B images in webdataset format. The recaptioned prompts are from https://huggingface.co/datasets/OpenGVLab/InternVL-SA-1B-Caption image10M<n<100M1 likes3.6k downloads1y agoHugging Face12laion /conceptual-captions-12m-webdatasetimage10K<n<100K34 likes3.1k downloads5y agoHugging Face13GeroldMeisinger /laion2b-en-a65_cogvlm2-4bit_captions Abstract This dataset contains image captions for the laion2B-en aesthetics>=6.5 image dataset using CogVLM2-4bit with the "laion-pop"-prompt to generate captions which were "likely" used in Stable Diffusion 3 training. From these image captions new synthetic images were generated using stable-diffusion-3-medium (batch-size=8). The synthetic images are best viewed locally by cloning this repo with: git lfs install git clone… See the full description on the dataset page: https://huggingface.co/datasets/GeroldMeisinger/laion2b-en-a65_cogvlm2-4bit_captions.imageimage-classification1K<n<10K6 likes3k downloads2y agoHugging Face14Ryan-sjtu /ffhq512-captionimage10K<n<100K7 likes2.3k downloads3y agoHugging Face15Felldude /Gradients_Gradients_and_Text_Full_Logic_Captionsimage1K<n<10K2 likes1.8k downloads1mo agoHugging Face16limingcv /Captioned_COCOStuffimage100K<n<1M2 likes1.8k downloads3y agoHugging Face17Borise /CaptionQA 📌 CaptionQA Benchmark A high-density, taxonomy-grounded benchmark for evaluating image caption quality and the alignment between image information and generated captions 📄 Paper: CaptionQA: Is Your Caption as Useful as the Image Itself? 📦 Evaluation Code: GitHub Repository Sample Usage You can load the dataset using the Hugging Face datasets library: from datasets import load_dataset # Load the entire dataset dataset = load_dataset("Borise/CaptionQA") # Load a… See the full description on the dataset page: https://huggingface.co/datasets/Borise/CaptionQA.imageimage-text-to-textn<1K9 likes1.7k downloads10mo agoHugging Face18diffusers /pokemon-gpt4-captions Dataset Card for "pokemon-gpt4-captions" This dataset is just lambdalabs/pokemon-blip-captions but the captions come from GPT-4 (Turbo). Code used to generate the captions: import base64 from io import BytesIO import requests from PIL import Image def encode_image(image): buffered = BytesIO() image.save(buffered, format="JPEG") img_str = base64.b64encode(buffered.getvalue()) returnimg_str.decode("utf-8") def create_payload(image_string): payload = {… See the full description on the dataset page: https://huggingface.co/datasets/diffusers/pokemon-gpt4-captions.imagetext-to-imagen<1K42 likes1.4k downloads3y agoHugging Face19Multimodal-Fatima /COCO_captions_train Dataset Card for "COCO_captions_train" More Information needed image100K<n<1M7 likes1.3k downloads4y agoHugging Face20unography /movie-scenes-captionedimage100K<n<1M0 likes1.3k downloads2y agoHugging Face21pranked03 /flowers-blip-captions Dataset Card for "flowers-blip-captions" More Information needed image1K<n<10K7 likes1.2k downloads4y agoHugging Face22ProGamerGov /synthetic-dataset-1m-dalle3-high-quality-captions Dataset Card for Dalle3 1 Million+ High Quality Captions Alt name: Human Preference Synthetic Dataset Example grids for landscapes, cats, creatures, and fantasy are also available. Description: This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/ProGamerGov/synthetic-dataset-1m-dalle3-high-quality-captions.imagetext-to-image1M<n<10M154 likes1.2k downloads2y agoHugging Face23JustInEchoes /Senko-san-NAI-V5-Curated-HD-ULTRA-Captions Senko-san | NovelAI V5 Curated HD ULTRA | Captions This is a Senko-san LoRA training dataset for The Helpful Fox Senko-san. It contains 1,000 native 1024×1024 PNG images generated with NovelAI V5 Curated, paired with natural-language captions. The HD ULTRA collection brings together 598 base compositions regenerated in V5 Curated and 402 additional activity scenes. Use it to train your own Senko-san LoRA, bring the character to another base model, or experiment with your… See the full description on the dataset page: https://huggingface.co/datasets/JustInEchoes/Senko-san-NAI-V5-Curated-HD-ULTRA-Captions.imagetext-to-image1K<n<10K1 likes1.1k downloads3d agoHugging Face24svjack /Chinese_Children_Image_Captioning_Dataset_Split0 CODP-1200:Children Oral Description of Picture(Chinese-Child-Captions) CODP-1200: An AIGC based benchmark for assisting in child language acquisition 数据集介绍 目前已知最大的儿童图像描述数据集,children image captioning 共有1200张图片 每张图片对应五个中文描述,每两张图片为一组 描述文字600*5=3000 如果使用CODP-1200数据集,请引用以下文章 @article{LENG2024102627, title = {CODP-1200: An AIGC based benchmark for assisting in child language acquisition}, journal = {Displays}, volume = {82}, pages = {102627}, year =… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Chinese_Children_Image_Captioning_Dataset_Split0.image1K<n<10K0 likes1.1k downloads2y agoHugging Face25clip-benchmark /wds_mscoco_captionsimage10K<n<100K4 likes992 downloads4y agoHugging Face26Obscure-Entropy /CONCEPTUAL_CAPTIONS_HU_FILTEREDimage1M<n<10M0 likes974 downloads2y agoHugging Face27shunk031 /STAIR-Captions Dataset Card for STAIR-Captions Dataset Summary STAIR Captions is a large-scale dataset containing 820,310 Japanese captions. This dataset can be used for caption generation, multimodal retrieval, and image generation. Supported Tasks and Leaderboards [More Information Needed] Languages The language data in JDocQA is in Japanese (BCP-47 ja-JP). Dataset Structure Data Instances [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/shunk031/STAIR-Captions.imageimage-to-text100K<n<1M6 likes970 downloads2y agoHugging Face28fwd4xl /enhanced_image_captionsimage100K<n<1M0 likes887 downloads2y agoHugging Face29graph-based-captions /GBC10M Graph-based captioning (GBC) is a new image annotation paradigm that combines the strengths of long captions, region captions, and scene graphs GBC interconnects region captions to create a unified description akin to a long caption, while also providing structural information similar to scene graphs. ** The associated data point can be found at demo/water_tower.json Description and data format The GBC10M dataset, derived from the original images in CC12M, is… See the full description on the dataset page: https://huggingface.co/datasets/graph-based-captions/GBC10M.imageimage-to-text10M<n<100M35 likes832 downloads2y agoHugging Face30alexandrainst /nordjylland-news-image-captioning Dataset Card for "nordjylland-news-image-captioning" Dataset Summary This dataset is a collection of image-caption pairs from the Danish newspaper TV2 Nord. Supported Tasks and Leaderboards Image captioning is the intended task for this dataset. No leaderboard is active at this point. Languages The dataset is available in Danish (da). Dataset Structure An example from the dataset looks as follows. { "file_name": "1.jpg", "caption":… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/nordjylland-news-image-captioning.imageimage-to-text10K<n<100K4 likes811 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.