Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BLIP3o /BLIP3o-Pretrain-Long-Caption BLIP3o Pretrain Long-Caption Dataset This collection contains 27 million images, each paired with a long (~120 token) caption generated by Qwen/Qwen2.5-VL-7B-Instruct. Download from huggingface_hub import snapshot_download snapshot_download( repo_id="BLIP3o/BLIP3o-Pretrain-Long-Caption", repo_type="dataset" ) Load Dataset without Extracting You don’t need to unpack the .tar archives, use WebDataset support in 🤗datasets instead: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/BLIP3o/BLIP3o-Pretrain-Long-Caption.image10M<n<100M74 likes16k downloads1y agoHugging Face02BLIP3o /BLIP3o-Pretrain-Short-Caption BLIP3o Pretrain Short-Caption Dataset This collection contains 5 million images, each paired with a short (~20 token) caption generated by Qwen/Qwen2.5-VL-7B-Instruct. Download from huggingface_hub import snapshot_download snapshot_download( repo_id="BLIP3o/BLIP3o-Pretrain-Short-Caption", repo_type="dataset" ) Load Dataset without Extracting You don’t need to unpack the .tar archives, use WebDataset support in 🤗datasets instead: from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/BLIP3o/BLIP3o-Pretrain-Short-Caption.image1M<n<10M10 likes9.3k downloads1y agoHugging Face03hanlincs /InternVL-SA1B-Caption-WebDatasetThis repo contains the recaptioned SA1B images in webdataset format. The recaptioned prompts are from https://huggingface.co/datasets/OpenGVLab/InternVL-SA-1B-Caption image10M<n<100M1 likes3.6k downloads1y agoHugging Face04laion /conceptual-captions-12m-webdatasetimage10K<n<100K34 likes3.1k downloads5y agoHugging Face05laion /captioned-ai-music-snippets Dataset Overview A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned with Gemini Flash 2.0. Designed specifically to train and evaluate audio captioning models. Source Clips are randomly cut from the songs referenced in the nyuuzyou/suno repository. Captioning All excerpts have been annotated using Gemini Flash 2.0 for high‑quality, human‑readable audio descriptions. License Apache 2.0 audio1M<n<10M16 likes2k downloads11mo agoHugging Face06clip-benchmark /wds_mscoco_captionsimage10K<n<100K4 likes992 downloads4y agoHugging Face07laion /timbre-audio-caption-pairsaudio100K<n<1M2 likes783 downloads10mo agoHugging Face08clip-benchmark /wds_mscoco_captions2017image10K<n<100K8 likes441 downloads3y agoHugging Face09mispeech /MECAT-CaptionMECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks 📖 Paper | 🛠️ GitHub | 🎧 Demo | 🔊 MECAT-QA (HF) Dataset Description MECAT (Multi-Expert Chain for Audio Tasks) is a comprehensive benchmark constructed on large-scale data to evaluate machine understanding of audio content through two core tasks: Audio Captioning: Generating textual descriptions for given audio Audio Question Answering: Answering questions about given audio Generated via… See the full description on the dataset page: https://huggingface.co/datasets/mispeech/MECAT-Caption.audioaudio-classification10K<n<100K4 likes348 downloads5mo agoHugging Face10mitermix /audioset-with-grounded-captionsaudio1M<n<10M4 likes256 downloads1y agoHugging Face11laion /freesound-commercially-permissive-subset-with-captionsaudio100K<n<1M2 likes195 downloads11mo agoHugging Face12data-archetype /ffhq_captioned_1024 ffhq_captioned_1024 A captioned bucketed-shards export of gaunernst/ffhq-1024-wds. This export contains 70,000 square face and portrait images from FFHQ, stored as JPEG TAR shards in a single 1024 x 1024 bucket. The source images are decoded from the original dataset, deterministically converted to RGB, and re-encoded as high-quality JPEG (quality=95, adaptive subsampling). Captions were generated with a Gemini 2.5 Flash Lite primary pass and a Mistral Medium 3.1 fallback. Intended… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/ffhq_captioned_1024.imagetext-to-image10K<n<100K0 likes193 downloads6mo agoHugging Face13wendlerc /CaptionedSynthTextThis dataset has been created by Stability AI and LAION. SynthText is a popular OCR dataset, where random texts are rendered into random locations in images based on depth maps. In this dataset, we additionally computed image captions using BLIP2. Caption: "a close up of a leopard's face with a blurry background" image100K<n<1M2 likes161 downloads3y agoHugging Face14TTS-AGI /majestrino-unified-detailed-captions Majestrino Unified Detailed Captions Filtered subset of laion/majestrino-data containing all samples with unified_detailed_caption. Stats 4,658,407 samples 932 tar files (~1.1 GB each) ~1,017 GB total Format Each tar contains paired .flac + .json files. JSON fields: caption — the unified detailed caption caption_type — always unified_detailed_caption transcription — speech transcription (when available, normalized from multiple source keys) duration — audio… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions.audioaudio-classification1M<n<10M3 likes142 downloads7mo agoHugging Face15PixArt-alpha /SAM-LLaVA-Captions10Mtext10M<n<100M65 likes129 downloads3y agoHugging Face16TTS-AGI /majestrino-unified-detailed-captions-temporal Majestrino Unified Detailed Captions with Temporal Aspects Filtered subset of laion/majestrino-data containing only samples with unified_detailed_caption_with_temporal_aspects. Stats 4,128,665 samples 826 tar files (~1.1 GB each) ~878 GB total Format Each tar contains paired .flac + .json files. JSON fields: caption — the unified detailed caption with temporal aspects caption_type — always unified_detailed_caption_with_temporal_aspects transcription — speech… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/majestrino-unified-detailed-captions-temporal.audioaudio-classification1M<n<10M0 likes100 downloads7mo agoHugging Face17DamianBoborzi /objaverse_processed_renders_and_captionsContains rendered views and captions from Objaverse XL objects. the objects are from the alignment and TRELLIS500K (over 1 Millionen processed objects) dataset. We downloaded and rendered 4 views of each object. We added TRELLIS and CAP3D Captions where available. If there were no captions we generated new captions with the large version of Florence 2. This is the base dataset we used to generate MeshFleet which is described in MeshFleet: Filtered and Annotated 3D Vehicle Dataset for Domain… See the full description on the dataset page: https://huggingface.co/datasets/DamianBoborzi/objaverse_processed_renders_and_captions.imageimage-to-text1M<n<10M0 likes93 downloads1y agoHugging Face18laion /audioset-with-captionsaudio1M<n<10M2 likes69 downloads11mo agoHugging Face19thisnick /nsfw-video-still-caption-grid-onlyimage10K<n<100K14 likes65 downloads2y agoHugging Face20Salmonnn /InternVL-SA-1B-Caption-512image10M<n<100M0 likes63 downloads1y agoHugging Face21ChristophSchuhmann /wikiart_with_BLIP_captionstext10K<n<100K2 likes59 downloads4y agoHugging Face22ngqtrung /full-modality-video-caption Full Modality Video Caption Dataset A large-scale multimodal video dataset with comprehensive vision, audio, and integrated captions. Dataset Description This dataset contains 55,940 video segments (10 seconds each) with three types of captions: Vision Caption: Visual description generated by GPT-4o Audio Caption: Audio/speech description generated by Qwen3-Omni-30B-A3B-Captioner Video Caption: Integrated multi-modal description combining vision and audio generated by… See the full description on the dataset page: https://huggingface.co/datasets/ngqtrung/full-modality-video-caption.text10K<n<100K0 likes54 downloads1y agoHugging Face23lingcarzy /synthetic-dataset-1m-dalle3-high-quality-captions Dataset Card for Dalle3 1 Million+ High Quality Captions Alt name: Human Preference Synthetic Dataset Example grids for landscapes, cats, creatures, and fantasy are also available. Description: This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/lingcarzy/synthetic-dataset-1m-dalle3-high-quality-captions.imagetext-to-image1M<n<10M0 likes52 downloads6mo agoHugging Face24CaptionEmporium /dalle3-llama3.2-11b Dataset Card for dalle3-llama3.2-11b Dataset Summary This is 3,577,716 new synthetic captions for the 1,192,572 images found in ProGamerGov/synthetic-dataset-1m-dalle3-high-quality-captions. The dataset was filtered for duplicates and then re-encoded with JPEGXL lossless or lossy depending on the source. The long captions were produced using meta-llama/Llama-3.2-11B-Vision-Instruct. Medium and short captions were produced from these captions using… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/dalle3-llama3.2-11b.texttext-to-image1M<n<10M0 likes47 downloads2y agoHugging Face25KhangTruong /NWPU-Captionimage10K<n<100K0 likes42 downloads2y agoHugging Face26shauray /aesthetic-cleaned-captionedimage1K<n<10K1 likes36 downloads4mo agoHugging Face27deepghs /midjourney_captioned_23m_fullgated Midjourney Captioned Full Dataset This is the full dataset of Midjourney Captioned 23M dataset. And all the original images are maintained here. Thanks to the contribution of a certain third-party data provider who wishes to remain anonymous. Information Images There are 23167456 images in total. The maximum ID of these images is 23167456. Last updated at 2024-12-01 12:11:43 UTC. These are the information of recent 50 images: id width height filename… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/midjourney_captioned_23m_full.imageimage-classification10M<n<100M36 likes30 downloads2y agoHugging Face286DammK9 /e621_2024-captions-1ktar E621 2024 captions only in 1k tar Raw captions jointed from lodestones/e621-captions It doesn't align to any dataset yet. meta_cap.json has been provided in compressed format if you want to train with kohyas triner. Currently I'm trying to merge this with my 2024 version. Core logic The script building this 1ktar There is not much choice, I don't have GPU to run for 1M captions with VLM so I just "take it or leave it". rearranged_tags = [row.regular_summary… See the full description on the dataset page: https://huggingface.co/datasets/6DammK9/e621_2024-captions-1ktar.textimage-classification1M<n<10M0 likes23 downloads2y agoHugging Face29Dant33 /WikiArt-81K-BLIP_2-captions WikiArt Enhanced Dataset Description This dataset contains 81,444 artistic images from WikiArt, organized into different artistic genres. It has undergone several improvements and corrections to optimize its use in machine learning tasks and computational art analysis. Credits to the original author of daset go to: WikiArt Enhancements 1. Encoding Issues Correction Fixed encoding issues in filenames and artist information. All filenames were renamed… See the full description on the dataset page: https://huggingface.co/datasets/Dant33/WikiArt-81K-BLIP_2-captions.image10K<n<100K0 likes23 downloads2y agoHugging Face30jacklishufan /soundnet-flux-captionimage100K<n<1M0 likes19 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.