datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MM-Food-100K
Overview
This project aims to introduce and release a comprehensive food image dataset designed specifically for computer vision tasks, particularly food recognition, classification, and nutritional analysis. We hope this dataset will provide a reliable resource for researchers and developers to advance the field of food AI. By publishing on Hugging Face, we expect to foster community collaboration and accelerate innovation in applications such as smart recipe recommendations… See the full description on the dataset page: https://huggingface.co/datasets/Humanbased-AI/MM-Food-100K.text-2-image-human-preferences-2m
Text-to-image human preferences: 2M votes across 30 models
This dataset contains the complete voting record behind the
Datapoint Image Bench
leaderboard: 2,161,160 validated pairwise votes — exactly 10 for each of
216,116 image pairs. The votes compare 30 text-to-image models in a complete
round-robin on 500 prompts, judged by annotators from over 200 countries.
Every vote includes the annotator's trust score at the time the vote was
cast.
Built on the Datapoint annotation… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-human-preferences-2m.ad-creative-quality-human-vs-llm
Human Expert vs LLM Judge: Facebook Ad Creative Quality
500 real Facebook ads from 253 advertisers, each rated for creative quality by a human ad expert AND by a vision LLM — with the LLM's full reasoning.
The headline finding baked into this data: the human and the LLM agree on image quality only 26.8% of the time. The LLM judge rates 71.8% of ads "good"; the human expert rates only 20% "good". If you are using an LLM as a judge of ad creative (or any subjective visual quality)… See the full description on the dataset page: https://huggingface.co/datasets/AdControlCenter/ad-creative-quality-human-vs-llm.MJRecap_Human_filtered
Dataset Card
Introduction
In recent studies, there has been a significant concern regarding data privacy. While one can scrape data online or create one’s own datasets, there is a major problem concerning the privacy of any individuals depicted in a dataset, their consent to be published, as well as dataset copyright.
Because of this, the present dataset aims to focus on fully human-free and copyright-free material based on the existing published datasets created… See the full description on the dataset page: https://huggingface.co/datasets/RunningInTheVoid/MJRecap_Human_filtered.human-activity-pose_v4
🧍 Human Activity Pose Dataset (Split Version)
This dataset contains human pose landmarks extracted with MediaPipe Pose,
annotated with activity labels and textual descriptions in English.
Dataset structure
train/ — 80% of samples for training
validation/ — 20% of samples for validation
Each record includes:
33 pose keypoints (fields: x, y, z, visibility)
label: activity name (e.g., reading, dancing, office_work)
description: a short textual description of the action… See the full description on the dataset page: https://huggingface.co/datasets/guillherms/human-activity-pose_v4.
