Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CaptionEmporium /pexels-568k-internvl2 Dataset Card for pexels-568k-internvl2 Dataset Summary This is 567,573 synthetic captions for the images found in ptx0/photo-concept-bucket. The captions were produced using OpenGVLab/InternVL2-40B-AWQ. The dataset was grounded for captioning using the tags originally listed. Languages The text is in English, but occasionally text in images in other languages is transcribed. Intended Usage Training text-to-image models and other machine learning… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/pexels-568k-internvl2.imagetext-to-image100K<n<1M21 likes6.5k downloads2y agoHugging Face02friedrichor /ActivityNet_Captions About ActivityNet Captions contains 20K long-form videos (180s as average length) from YouTube and 100K captions. Most of the videos contain over 3 annotated events. We follow the existing works to concatenate multiple short temporal descriptions into long sentences and evaluate ‘paragraph-to-video’ retrieval on this benchmark. We adopt the official split: Train: 10,009 videos, 10,009 captions (concatenate from 37,421 short captions) Test (Val1): 4,917 videos, 4,917 captions… See the full description on the dataset page: https://huggingface.co/datasets/friedrichor/ActivityNet_Captions.texttext-to-video10K<n<100K16 likes4.3k downloads1y agoHugging Face03AudioVisual-Caption /ASID-1M ASID-1M: Attribute-Structured and Quality-Verified Audiovisual Instructions [🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code] Introduction We introduce ASID-1M, a large-scale audiovisual instruction dataset built to support universal video understanding with fine-grained, controllable supervision. Most existing video-instruction data represents complex audiovisual content as a single, monolithic caption. This often leads to incomplete coverage (missing audio… See the full description on the dataset page: https://huggingface.co/datasets/AudioVisual-Caption/ASID-1M.textimage-text-to-text100K<n<1M85 likes1.3k downloads7mo agoHugging Face04wchai /Video-Detailed-Caption Video Detailed Caption Benchmark Resources Website arXiv: Paper GitHub: Code Huggingface: AuroraCap Model Huggingface: VDC Benchmark Huggingface: Trainset Features Benchmark Collection and Processing We building VDC upon Panda-70M, Ego4D, Mixkit, Pixabay, and Pexels. Structured detailed captions construction pipeline. We develop a structured detailed captions construction pipeline to generate extra detailed descriptions from various… See the full description on the dataset page: https://huggingface.co/datasets/wchai/Video-Detailed-Caption.textvideo-text-to-text1K<n<10K17 likes574 downloads2y agoHugging Face05neurlang /Minecraft-Skins-Captioned-1M Dataset Card for Minecraft Skins Dataset Summary This dataset contains 981,079 unique Minecraft player skins collected from various sources. Each skin is stored as a base64-encoded image with a unique identifier. Dataset Structure Data Fields This dataset includes the following fields: hash: A data dependent hash. These hashes are generated from raw bytes and will be same if the skin is identical. image: The skin image encoded in base64 format.… See the full description on the dataset page: https://huggingface.co/datasets/neurlang/Minecraft-Skins-Captioned-1M.textimage-classification1M<n<10M8 likes409 downloads1y agoHugging Face06Obscure-Entropy /conceptual_captions_jsonimage1M<n<10M0 likes392 downloads2y agoHugging Face07CaptionEmporium /anime-caption-danbooru-2021-sfw-5m-hq Dataset Card for anime-caption-danbooru-2021-sfw-5m-hq Dataset Summary This is 5.71 M captions of 1.43 M images from a safe-for-work (SFW) filtered subset of the Danbooru 2021 dataset. There are 4 captions per image: 1 by CogVLM, 1 by llava-v1.6-34b, 1 llava-v1.6-34b cleaned, and 1 llava-v1.6-34b shortened. See the sections below for how they were generated. Most captions are substantially larger than 77 tokens and are unsuitable for discrimination using current… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/anime-caption-danbooru-2021-sfw-5m-hq.textimage-to-text1M<n<10M29 likes382 downloads2y agoHugging Face08yauheniya-adesso /icongenai-svg-captions IconGenAI SVG Captions Captioned SVG icons from the Iconify corpus, intended for fine-tuning text-to-SVG generation models. Part of the IconGenAI research project. Files Two files are provided at different stages of the processing pipeline: File Records Purpose icons_captioned_merged.jsonl 275,912 Full license-filtered corpus with VLM-generated captions and collection metadata icons_training_captioned.jsonl227,821 Quality-filtered, normalised subset… See the full description on the dataset page: https://huggingface.co/datasets/yauheniya-adesso/icongenai-svg-captions.tabulartext-to-image100K<n<1M2 likes285 downloads6mo agoHugging Face09bdsqlsz /Tom_and_Jerry_captions3SRC46HSXPXFC63X34QGLDUCIXGZNOXR tabularn<1K2 likes271 downloads2y agoHugging Face10vvwangvv /emilia-captions-v3text10M<n<100M0 likes258 downloads9mo agoHugging Face11embedding-data /coco_captions_quintets Dataset Card for "coco_captions" Dataset Summary COCO is a large-scale object detection, segmentation, and captioning dataset. This repo contains five captions per image; useful for sentence similarity tasks. Disclaimer: The team releasing COCO did not upload the dataset to the Hub and did not write a dataset card. These steps were done by the Hugging Face team. Supported Tasks Sentence Transformers training; useful for semantic search and sentence… See the full description on the dataset page: https://huggingface.co/datasets/embedding-data/coco_captions_quintets.textsentence-similarity10K<n<100K6 likes208 downloads4y agoHugging Face12embedding-data /flickr30k_captions_quintets Dataset Card for "flickr30k-captions" Dataset Summary We propose to use the visual denotations of linguistic expressions (i.e. the set of images they describe) to define novel denotational similarity metrics, which we show to be at least as beneficial as distributional similarities for two tasks that require semantic inference. To compute these denotational similarities, we construct a denotation graph, i.e. a subsumption hierarchy over constituents and their denotations… See the full description on the dataset page: https://huggingface.co/datasets/embedding-data/flickr30k_captions_quintets.text10K<n<100K4 likes201 downloads4y agoHugging Face13OpenGVLab /InternVL-SA-1B-Caption Dataset Card for InternVL-SA-1B-Caption Overview The InternVL-SA-1B-Caption Dataset is a bilingual dataset created using the InternVL2-Llama3-76B model. The dataset contains 12 million image-caption pairs in both English and Chinese. All images are sourced from Meta’s SA-1B dataset, and captions were generated using specific prompts designed to minimize hallucinations and ensure accurate descriptions based on visible image content. The dataset is intended for use in tasks… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/InternVL-SA-1B-Caption.tabular1M<n<10M25 likes193 downloads2y agoHugging Face14CaptionEmporium /conceptual-captions-cc12m-llavanext Dataset Card for conceptual-captions-cc12m-llavanext Dataset Summary This is a data of 21,930,344 synthetic captions for 10,965,172 images from conceptual_12m. In the interest of reproducibility, an archive found here on Huggingface was used (cc12m-wds). The captions were produced using llama3-llava-next-8b inferenced in float16, followed by cleanup and shortening with Meta-Llama-3-8B. Languages The captions are in English. Data Instances An… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/conceptual-captions-cc12m-llavanext.imagetext-to-image10M<n<100M28 likes182 downloads2y agoHugging Face15alinasdkey /graph-captioning-train-onlyimagen<1K0 likes139 downloads1y agoHugging Face16Sreevardhan1729 /ActivityNet_Captions About ActivityNet Captions contains 20K long-form videos (180s as average length) from YouTube and 100K captions. Most of the videos contain over 3 annotated events. We follow the existing works to concatenate multiple short temporal descriptions into long sentences and evaluate ‘paragraph-to-video’ retrieval on this benchmark. We adopt the official split: Train: 10,009 videos, 10,009 captions (concatenate from 37,421 short captions) Test (Val1): 4,917 videos, 4,917 captions… See the full description on the dataset page: https://huggingface.co/datasets/Sreevardhan1729/ActivityNet_Captions.texttext-to-video10K<n<100K0 likes129 downloads6mo agoHugging Face17Senqiao /LiDAR-LLM-Nu-Caption Dataset Details Dataset type: This is the nu-Caption dataset, a QA dataset designed for training MLLM models on caption tasks in autonomous driving scenarios. It is built upon the NuScenes dataset. Dataset keys: "answer" is the output of the VLM models using image data. "answer_lidar" uses GPT4O-mini to filter information that cannot be obtained from the image data. If you want to train the model like LiDAR-LLM, which only uses the LiDAR modality and does not use the vision modality… See the full description on the dataset page: https://huggingface.co/datasets/Senqiao/LiDAR-LLM-Nu-Caption.textquestion-answering100K<n<1M8 likes108 downloads2y agoHugging Face18AVoCaDO-Captioner /training_settext100K<n<1M1 likes105 downloads7mo agoHugging Face19Waterfront /social-media-captions Social Media Captions Based on the Instagram Influencer Dataset from Seungbae Kim, Jyun-Yu Jiang, and Wei Wang Extended with photo descriptions of ydshieh/vit-gpt2-coco-en model to create a dataset which can be used to finetune Llama-2. 20k smaller subset: Waterfront/social-media-captions-20k 10k smaller subset: Waterfront/social-media-captions-10k text10K<n<100K3 likes85 downloads3y agoHugging Face20DAMO-NLP-SG /Multi-Source-Video-Captioning Multi-source Video Captioning (MSVC) Dataset Card Dataset details Dataset type: MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities. Dataset detail: MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.textvisual-question-answering1K<n<10K7 likes85 downloads2y agoHugging Face21Waterfront /social-media-captions-10k Social Media Captions Based on the Instagram Influencer Dataset from Seungbae Kim, Jyun-Yu Jiang, and Wei Wang Extended with photo descriptions of ydshieh/vit-gpt2-coco-en model to create a dataset which can be used to finetune Llama-2. 60k complete dataset: Waterfront/social-media-captions 20k bigger subset: Waterfront/social-media-captions-20k text1K<n<10K3 likes83 downloads3y agoHugging Face22junwann /CSFM-ImageNet1K-Caption CSFM-ImageNet1K-Caption Dataset Project Page | Paper | Code This repository contains dataset associated with the paper "Better Source Better Flow: Learning Condition-Dependent Source Distribution for Flow Matching". This dataset is used for training and evaluating Condition-dependent Source Flow Matching (CSFM), a framework that learns condition-dependent source distributions for flow matching. We recaptioned the ImageNet-1K dataset using Qwen3-VL-8B Instruct, resulting in detailed… See the full description on the dataset page: https://huggingface.co/datasets/junwann/CSFM-ImageNet1K-Caption.texttext-to-image1M<n<10M4 likes82 downloads8mo agoHugging Face23SOTAagi2030 /Harbor-Photo-Captions Harbor Photo Captions This collection contains caption records for digitized waterfront photographs. Rights register Accession: HP-19 Rights statement: Public Domain Mark 1.0 Depositing archive: Tideglass Image Library Caption normalization follows the catalog's controlled vocabulary. textn<1K0 likes76 downloads1mo agoHugging Face24Tran1312 /Image-Caption-100k BLIP3o Long-Caption 100K Image-Text Subset This dataset is a locally reorganized subset of BLIP3o/BLIP3o-Pretrain-Long-Caption, containing approximately 100,000 image-text pairs selected from the original BLIP3o long-caption pretraining dataset [1]. The original BLIP3o long-caption collection contains approximately 27 million images, each paired with a long caption of roughly 120 tokens generated using Qwen2.5-VL-7B-Instruct [1]. The BLIP3-o project was introduced as part of a… See the full description on the dataset page: https://huggingface.co/datasets/Tran1312/Image-Caption-100k.textimage-to-text100K<n<1M0 likes73 downloads13d agoHugging Face25saifkhichi96 /mpii-human-pose-captions Dataset Card for MPII Human Pose Descriptions Dataset Summary The MPII Human Pose Descriptions dataset extends the widely-used MPII Human Pose Dataset with rich textual annotations. These annotations are generated by various state-of-the-art language models (LLMs) and include detailed descriptions of the activities being performed, the count of people present, and their specific poses. The dataset consists of the same image splits as provided in MMPose, with 14644… See the full description on the dataset page: https://huggingface.co/datasets/saifkhichi96/mpii-human-pose-captions.tabularzero-shot-classification10K<n<100K3 likes66 downloads2y agoHugging Face26CaptionEmporium /refined-anime-instruct-en-641k Dataset Card for refined-anime-instruct-en-641k Dataset Summary This is 641,497 instructions for an expert model that knows about the following things: Anime Manga Live Action Shows Children's Films Western Comics Agatha Christie Novels and Adaptations (not sure why this is over-represented) Video Games It is derived from Refined-Anime-Text by filtering out all ZH entries. According to their README.md, these outputs are completions derived from GPT3.5 and GPT4.… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/refined-anime-instruct-en-641k.textquestion-answering100K<n<1M5 likes48 downloads3y agoHugging Face27HamGangster /coco_2017_caption_traintextn<1K0 likes42 downloads3y agoHugging Face28HamGangster /coco_2017_caption_validationtextn<1K0 likes41 downloads3y agoHugging Face29CaptionEmporium /furry-e621-sfw-7m-hq Dataset Card for furry-e621-sfw-7m-hq Dataset Summary This is 6.92 M captions of the images from the safe-for-work (SFW) split of e621 ("e926"). It extends to January 2023, before the widespread advent of machine learning images. It includes captions created by LLMs and a custom multilabel classifier along with CogVLM. There are 8 LLM (mistralai/Mistral-7B-v0.1) and 1 CogVLM (THUDM/CogVLM) captions per image. Most captions are substantially larger than 77 tokens and are… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/furry-e621-sfw-7m-hq.textimage-to-text100K<n<1M7 likes41 downloads3y agoHugging Face30Waterfront /social-media-captions-20k Social Media Captions Based on the Instagram Influencer Dataset from Seungbae Kim, Jyun-Yu Jiang, and Wei Wang Extended with photo descriptions of ydshieh/vit-gpt2-coco-en model to create a dataset which can be used to finetune Llama-2. 60k complete dataset: Waterfront/social-media-captions 10k smaller subset: Waterfront/social-media-captions-10k text10K<n<100K0 likes39 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.