AdControlCenter/ad-creative-quality-human-vs-llm
Human Expert vs LLM Judge: Facebook Ad Creative Quality 500 real Facebook ads from 253 advertisers, each rated for creative quality by a human ad expert AND by a vision LLM — with the LLM's full reasoning. The headline finding baked into this data: the human and the LLM agree on image quality only 26.8% of the time. The LLM judge rates 71.8% of ads "good"; the human expert rates only 20% "good". If you are using an LLM as a judge of ad creative (or any subjective visual quality)… See the full description on the dataset page: https://huggingface.co/datasets/AdControlCenter/ad-creative-quality-human-vs-llm.
Human Expert vs LLM Judge: Facebook Ad Creative Quality
500 real Facebook ads from 253 advertisers, each rated for creative quality by a human ad expert AND by a vision LLM — with the LLM's full reasoning.
The headline finding baked into this data: the human and the LLM agree on image quality only 26.8% of the time. The LLM judge rates 71.8% of ads "good"; the human expert rates only 20% "good". If you are using an LLM as a judge of ad creative (or any subjective visual quality), this dataset quantifies the positivity bias you are inheriting.
Built and released by AdControlCenter (ACC), an AI ad-creation platform. This is a slice of the internal corpus behind our report We analyzed 13,572 Facebook ads.
What each row is
One row = one ad that was live in the Meta Ad Library, identified by its public ad_archive_id. We publish derived features and quality labels only — no ad images, headlines, or body copy are included. To view the original creative, look the ad up in the Ad Library:
https://www.facebook.com/ads/library/?id=<ad_archive_id>(Meta removes inactive non-political ads from the library, so some ads may no longer be viewable there.)
Columns
How the sample was built
- Source pool: ACC's internal corpus of 13,572 Meta Ad Library ads from 311 curated advertisers across 11 verticals.
- 982 ads had both a human rating and an LLM rating; 955 of those had a definitive human image score (
bad/fair/good, not skipped). - From those 955, we drew a category-stratified proportional sample of 500, deterministic (ordered by
ad_archive_id) within each category.
Annotation process
- Human labels: a single expert annotator (ACC's founder, a practitioner who reviews ad creative daily), using a 3-point rubric per dimension (headline / image / audience-fit). Single-annotator is a real limitation — treat the human labels as one calibrated expert's judgment, not ground truth by committee.
- LLM labels:
claude-sonnet-4-6with vision, shown the ad creative and copy, using the same 3-point rubric, producing scores plus written reasoning. The LLM did not see the human's ratings.
Known limitations
- Single human annotator; no inter-annotator agreement stats possible.
- Category skew: 44% of rows are
developer_tools; the sample mirrors the labeled pool, not the ad ecosystem. - One judge model only (
claude-sonnet-4-6); the bias measurement is about that model family, not "LLMs" universally. days_activeis an estimate at collection time, not a verified spend or performance number.human_headline_score/human_fit_scoreare present on only 29 rows — usehuman_image_scoreas the primary human label.
Suggested uses
- LLM-as-judge calibration: measure and correct positivity bias of vision LLMs on subjective creative quality.
- Judge fine-tuning / prompt engineering: use
human_image_scoreas target,ai_reasoningas a critique corpus. - Creative-longevity analysis:
days_active× quality labels ×category. - Feature analysis: does
has_text/image_typepredict expert quality verdicts?
License & citation
Released under CC BY 4.0. Cite as:
AdControlCenter (2026). Human Expert vs LLM Judge: Facebook Ad Creative Quality (500 ads).
https://adcontrolcenter.com/learn/we-analyzed-13572-facebook-adsQuestions or want the bigger slice (3,260 human ratings / 13.5k ads with derived features)? Reach out via adcontrolcenter.com.
