Team Ai
Datasetpublic

AdControlCenter/ad-creative-quality-human-vs-llm

Human Expert vs LLM Judge: Facebook Ad Creative Quality 500 real Facebook ads from 253 advertisers, each rated for creative quality by a human ad expert AND by a vision LLM — with the LLM's full reasoning. The headline finding baked into this data: the human and the LLM agree on image quality only 26.8% of the time. The LLM judge rates 71.8% of ads "good"; the human expert rates only 20% "good". If you are using an LLM as a judge of ad creative (or any subjective visual quality)… See the full description on the dataset page: https://huggingface.co/datasets/AdControlCenter/ad-creative-quality-human-vs-llm.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
2likes55downloads
Dataset Card

Human Expert vs LLM Judge: Facebook Ad Creative Quality

500 real Facebook ads from 253 advertisers, each rated for creative quality by a human ad expert AND by a vision LLM — with the LLM's full reasoning.

The headline finding baked into this data: the human and the LLM agree on image quality only 26.8% of the time. The LLM judge rates 71.8% of ads "good"; the human expert rates only 20% "good". If you are using an LLM as a judge of ad creative (or any subjective visual quality), this dataset quantifies the positivity bias you are inheriting.

Built and released by AdControlCenter (ACC), an AI ad-creation platform. This is a slice of the internal corpus behind our report We analyzed 13,572 Facebook ads.

What each row is

One row = one ad that was live in the Meta Ad Library, identified by its public ad_archive_id. We publish derived features and quality labels only — no ad images, headlines, or body copy are included. To view the original creative, look the ad up in the Ad Library:

https://www.facebook.com/ads/library/?id=<ad_archive_id>

(Meta removes inactive non-political ads from the library, so some ads may no longer be viewable there.)

Columns

ColumnTypeCoverageDescription
ad_archive_idstring500/500Meta Ad Library public ID. Join key back to the original creative.
advertiserstring500/500Advertiser page name (public in the Ad Library). 253 unique advertisers, incl. Cloudflare, Datadog-class brands.
categorystring500/500Our vertical classification: developer_tools (220), ecommerce (155), crm_sales (64), handmade_etsy (26), health_wellness (20), saas_productivity (15).
ad_typestring~500/500feed or display.
image_typestring485/500Vision-classified composition: designed-ad (475), product-only (8), unclear (2), blank (15).
has_textbool~500/500Whether the image has baked-in text/typography (headline, offer, wordmark).
days_activeint495/500Estimated days the ad had been running when collected — a weak but real longevity/performance proxy (ads that keep running tend to keep working).
human_headline_scorestring29/500Human rating of the headline: bad / fair / good. Sparse — the human labeling pass was image-first.
human_image_scorestring500/500The core human label. Expert rating of the creative/image: bad (208) / fair (192) / good (100).
human_fit_scorestring29/500Human rating of headline↔image↔audience fit. Sparse, same reason as above.
ai_headline_scorestring500/500LLM judge rating of the headline: bad / fair / good.
ai_image_scorestring500/500LLM judge rating of the image: good (359) / fair (138) / bad (3).
ai_fit_scorestring500/500LLM judge rating of headline↔image↔audience fit.
ai_reasoningstring500/500The LLM judge's full written rationale (a paragraph per ad). May quote short excerpts of ad copy in commentary context.
ai_modelstring500/500The judge model. claude-sonnet-4-6 (vision) for all rows.
collected_atdate500/500Date the ad was pulled from the Ad Library (2026).

How the sample was built

  • —Source pool: ACC's internal corpus of 13,572 Meta Ad Library ads from 311 curated advertisers across 11 verticals.
  • —982 ads had both a human rating and an LLM rating; 955 of those had a definitive human image score (bad/fair/good, not skipped).
  • —From those 955, we drew a category-stratified proportional sample of 500, deterministic (ordered by ad_archive_id) within each category.

Annotation process

  • —Human labels: a single expert annotator (ACC's founder, a practitioner who reviews ad creative daily), using a 3-point rubric per dimension (headline / image / audience-fit). Single-annotator is a real limitation — treat the human labels as one calibrated expert's judgment, not ground truth by committee.
  • —LLM labels: claude-sonnet-4-6 with vision, shown the ad creative and copy, using the same 3-point rubric, producing scores plus written reasoning. The LLM did not see the human's ratings.

Known limitations

  • —Single human annotator; no inter-annotator agreement stats possible.
  • —Category skew: 44% of rows are developer_tools; the sample mirrors the labeled pool, not the ad ecosystem.
  • —One judge model only (claude-sonnet-4-6); the bias measurement is about that model family, not "LLMs" universally.
  • —days_active is an estimate at collection time, not a verified spend or performance number.
  • —human_headline_score / human_fit_score are present on only 29 rows — use human_image_score as the primary human label.

Suggested uses

  • —LLM-as-judge calibration: measure and correct positivity bias of vision LLMs on subjective creative quality.
  • —Judge fine-tuning / prompt engineering: use human_image_score as target, ai_reasoning as a critique corpus.
  • —Creative-longevity analysis: days_active × quality labels × category.
  • —Feature analysis: does has_text / image_type predict expert quality verdicts?

License & citation

Released under CC BY 4.0. Cite as:

AdControlCenter (2026). Human Expert vs LLM Judge: Facebook Ad Creative Quality (500 ads).
https://adcontrolcenter.com/learn/we-analyzed-13572-facebook-ads

Questions or want the bigger slice (3,260 human ratings / 13.5k ads with derived features)? Reach out via adcontrolcenter.com.