Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TommyBsk /Embodied-Captioning Embodied Image Captioning – Manually Annotated Test Set Paper: Embodied Image Captioning: Self-supervised Learning Agents for Spatially Coherent Image Descriptions (ICCV 2025)Authors: Tommaso Galliena, Tommaso Apicella, Stefano Rosa, Pietro Morerio, Alessio Del Bue, Lorenzo NataleAffiliations: Italian Institute of Technology (IIT), University of GenoaProject Website: https://hsp-iit.github.io/embodied-captioningCode: https://github.com/hsp-iit/embodied-captioning 📦… See the full description on the dataset page: https://huggingface.co/datasets/TommyBsk/Embodied-Captioning.tabularimage-to-text1K<n<10K0 likes8.8k downloads1y agoHugging Face02hf-internal-testing /fixtures-captioning\\n0 likes3k downloads1y agoHugging Face03svjack /Chinese_Children_Image_Captioning_Dataset_Split0 CODP-1200:Children Oral Description of Picture(Chinese-Child-Captions) CODP-1200: An AIGC based benchmark for assisting in child language acquisition 数据集介绍 目前已知最大的儿童图像描述数据集,children image captioning 共有1200张图片 每张图片对应五个中文描述,每两张图片为一组 描述文字600*5=3000 如果使用CODP-1200数据集,请引用以下文章 @article{LENG2024102627, title = {CODP-1200: An AIGC based benchmark for assisting in child language acquisition}, journal = {Displays}, volume = {82}, pages = {102627}, year =… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Chinese_Children_Image_Captioning_Dataset_Split0.image1K<n<10K0 likes966 downloads2y agoHugging Face04alexandrainst /nordjylland-news-image-captioning Dataset Card for "nordjylland-news-image-captioning" Dataset Summary This dataset is a collection of image-caption pairs from the Danish newspaper TV2 Nord. Supported Tasks and Leaderboards Image captioning is the intended task for this dataset. No leaderboard is active at this point. Languages The dataset is available in Danish (da). Dataset Structure An example from the dataset looks as follows. { "file_name": "1.jpg", "caption":… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/nordjylland-news-image-captioning.imageimage-to-text10K<n<100K4 likes950 downloads3y agoHugging Face05zhiqiulin /video_captioningvideo1K<n<10K0 likes714 downloads1y agoHugging Face06ituperceptron /image-captioning-turkish Türkçe Image Captioning Veri Seti Bu veri seti BLIP3o modelinin pretrain eğitiminde kullanılan BLIP3o-Pretrain-Long-Caption ve BLIP3o-Pretrain-Short-Caption veri setlerinin Türkçeye çevirilmiş bir alt parçasıdır. Orijinal veri setinin oluşturulması ile ilgili detaylı bilgiye BLIP-3o makalesi üzerinden ulaşabilirsiniz. Veri seti Image-to-Text modellerinin eğitilmesinde veya ince ayar sürecinde kullanılabilir. Veri seti, orijinal veri setinin lisansı olan Apache 2.0 altında… See the full description on the dataset page: https://huggingface.co/datasets/ituperceptron/image-captioning-turkish.imageimage-to-text1M<n<10M7 likes545 downloads9mo agoHugging Face07ekacare /IntraOral_Gingivitis_Image_Captioning A DENTAL INTRAORAL IMAGE DATASET OF GINGIVITIS FOR IMAGE CAPTIONING Dataset Description This dataset is a copy of A Dental IntraOral Image Dataset of Gingivitis for Image Captioning which is shared with the license CC BY 4.0. This dataset contains 1,096 samples organized across multiple splits. The dataset includes image data. Splits train: 732 samples test: 182 samples validation: 182 samples Dataset Creation This dataset was created using… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/IntraOral_Gingivitis_Image_Captioning.imageimage-classification1K<n<10K0 likes471 downloads1y agoHugging Face08PerRing /coco_captioning_complete_formatimage100K<n<1M0 likes399 downloads11mo agoHugging Face09laion /Segmentation-Captioning-Assistant-Tuning-Data1 likes383 downloads11mo agoHugging Face10gijs /tacos-captioningaudio10K<n<100K1 likes319 downloads1y agoHugging Face11TMICCProj /Closed_Captioning_Lecture_Datasetaudio10K<n<100K0 likes312 downloads7mo agoHugging Face12AKCIT /coco2017-captioningimage100K<n<1M0 likes309 downloads6mo agoHugging Face13gorovuha /ru_image_captioningimage1K<n<10K0 likes187 downloads2y agoHugging Face14MagiBoss /COCO-Image-Captioningimage100K<n<1M0 likes187 downloads2y agoHugging Face15fancyfeast /joy-captioning-20250328b Work In Progress I'm still going back through my data to add in the URLs. textvisual-question-answering1M<n<10M24 likes168 downloads2y agoHugging Face16MykMaks /nordjylland-news-image-captioningimagezero-shot-classification10K<n<100K1 likes160 downloads2y agoHugging Face17tavish-mishra /my_image_captioning_datasetimage10K<n<100K1 likes159 downloads1y agoHugging Face18alinasdkey /graph-captioning-train-onlyimagen<1K0 likes137 downloads1y agoHugging Face19TeeA /Pokemon-Captioning-Classification 2000+ download monthly. Really appreciate for all of you guys: Buy me a coffee: https://buymeacoffee.com/tridoan Disclaimer: This model is provided "as-is" without any warranties. The authors are not responsible for any misuse or damages arising from its use. image1K<n<10K2 likes134 downloads1mo agoHugging Face20KORMo-VL /coco_captioningfrom COCO val2014 image100K<n<1M0 likes119 downloads7mo agoHugging Face21fancyfeast /joy-captioning-20250408aThis is the dataset used to do the initial training for JoyCaption Beta One (https://huggingface.co/fancyfeast/llama-joycaption-beta-one-hf-llava), before post-training. Contents Most of the dataset focusses on descriptions and captions for images, with a smaller subset covering general VQA tasks. Some of the questions and answers are human written, some are automated, some are machine written. The is_human column is True when the answer text is human written. WARNING… See the full description on the dataset page: https://huggingface.co/datasets/fancyfeast/joy-captioning-20250408a.textvisual-question-answering100K<n<1M10 likes116 downloads7mo agoHugging Face22Image-Captioning-ML /ucf101-captioned-mappedtext1K<n<10K0 likes109 downloads1y agoHugging Face23roshbeed /ai-residency-multimodal-captioning-dataimagen<1K0 likes93 downloads2mo agoHugging Face24orzhan /minecraft-captioningimageimage-to-textn<1K0 likes88 downloads3y agoHugging Face25DAMO-NLP-SG /Multi-Source-Video-Captioning Multi-source Video Captioning (MSVC) Dataset Card Dataset details Dataset type: MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities. Dataset detail: MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.textvisual-question-answering1K<n<10K7 likes86 downloads2y agoHugging Face26ashless /ECG-Captioningtext10K<n<100K0 likes80 downloads2y agoHugging Face27MSEarth /MSEarth_Captioningimage1K<n<10K2 likes78 downloads1y agoHugging Face28SeyedAli /Persian-Image-Captioning Dataset Card for "Persian-Image-Captioning" More Information needed image10K<n<100K2 likes77 downloads3y agoHugging Face29vidore /tatdqa_test_captioningimagedocument-question-answering1K<n<10K0 likes77 downloads1y agoHugging Face30indrad123 /image-captioning-idimage1K<n<10K0 likes75 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.