VilaVision/Imageclassifier
This dataset consists of 994 images generated using DALL·E and Midjourney. Each image is annotated with detailed textual descriptions and the count of distinct objects present, using Nvidia's Nim VLM API. The dataset is designed for use in image captioning, text-to-image generation, and image segmentation tasks. Dataset Details Dataset Description This dataset includes images generated by AI models, specifically DALL·E and Midjourney. Each image is annotated with:… See the full description on the dataset page: https://huggingface.co/datasets/VilaVision/Imageclassifier.
Conversations for this repository live on Hugging Face.
Team Ai shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face