Team Ai
Datasetpublic

EthnicErotic/phenotype-catalog

Ethnic Erotic Phenotype Catalog A structured complement to Wikipedia for ethnographic data — 1,700+ ethnic groups indexed with normalized linguistic, geographic, cultural, and phenotype metadata, plus 23K+ notable-people references and 5K+ vision-grounded per-image phenotype observations. Curated from the live catalog at ethnicerotic.com and published as an open dataset for anthropological reference, AI training, and ethnographic research. What's in v6 Two columns… See the full description on the dataset page: https://huggingface.co/datasets/EthnicErotic/phenotype-catalog.

sourceHugging Facecc-by-4.0updated 22h agoView on Hugging Face
1likes211downloads
Dataset Card

Ethnic Erotic Phenotype Catalog

A structured complement to Wikipedia for ethnographic data — 1,700+ ethnic groups indexed with normalized linguistic, geographic, cultural, and phenotype metadata, plus 23K+ notable-people references and 5K+ vision-grounded per-image phenotype observations.

Curated from the live catalog at ethnicerotic.com and published as an open dataset for anthropological reference, AI training, and ethnographic research.

What's in v6

Two columns this card has always documented had never carried a value, and one column carried two types. All three are repaired, and the vocabularies now ship a measured reliability record.

  • —`observability` was empty on all 196 rows since the config shipped in v3. The exporter read a key (from_photo) that no dimension has ever had; the real key is from_photograph. This card described the column as populated the whole time. It is now populated: 120 high, 41 medium, 6 low, 29 not assessable.
  • —`unit` was empty on all 196 rows for the same class of reason: it read d.unit, a key that does not exist on any dimension in any vocabulary file. It is now derived from the dimension id for the numeric dimensions and left empty elsewhere, which is what this card always said it was.
  • —`birth_year`, `death_year`, `coverage_score` and `confidence` were mixed-type columns in JSONL. They used an empty string for null, so the published notable_people.jsonl carried 3,279 integers and 9,815 strings in birth_year. They now emit a real null. The CSVs were always correct and are unchanged.
  • —`atlas.description` shipped raw editorial HTML, Bootstrap classes and all, while every other text column in the dataset is flattened to plain text. Now consistent.
  • —New in `vocabulary_dimensions`: requires_unclothed, minimum_visible_extent, and four reliability columns (reliability_kappa, reliability_n, reliability_verdict, reliability_study). See the reliability section below.

No column was renamed or removed.

Vocabulary reliability, measured

The ten head-and-neck vocabularies now carry a per-dimension reliability record from a test-retest study run on 2026-08-31: 196 portraits across 29 ethnic groups, re-read with the identical prompt and compared against their stored codings, dimension by dimension.

  • —92.9% agreement over dimensions both readings answered.
  • —90.6% counting a dimension one reading declined as a miss. Both are given because the choice of denominator has reversed a conclusion on this corpus before.
  • —73 of 84 rankable dimensions reach Cohen's kappa 0.60 or better.
  • —Three are marked `uninformative`: they return a single value for nearly every portrait, so the two readings agree without the dimension distinguishing anything. Ranking on raw agreement would have put them at the top.

This measures reproducibility, not validity. It is one model agreeing with itself, which says the instrument is stable and not that it is correct. No human rater has scored these images, and no inter-rater agreement exists for this vocabulary. Do not cite these numbers as validation, accuracy or ground truth.

Kappa is fragile wherever one value dominates. An earlier 78-portrait read put six of these dimensions at kappa at or below zero; at 196 they read between 0.43 and 0.66, and one read a perfect 1.000 at n=72 and 0.535 at n=176. Every dimension that moved that far has chance agreement above 85%. Those carry estimate_is_fragile in the vocabulary JSON. Read the band, not the digits.

Full per-dimension counts, including the marginals the kappa was computed from, are in the pipeline repository under reliability/.

What's in v4

This release is a coverage expansion and a data-quality pass over notable_people.

  • —Notable-people coverage nearly doubled — 13,094 → 23,034 rows, and the groups carrying at least one entry went 291 → 499. The largest rosters went to the big diaspora groups (Ukrainian, Tamil, Serbian, Slovak, Scottish), which is where the demand is.
  • —Person-shape validation. The original scrape had no guard against non-people entries, and a "List of X people" title on Wikipedia is very often a redirect to the main ethnic article — so whole-article harvests contributed references, external links and prose where people were expected. Entries are now validated for person shape (no event, place, ethnonym or year-leading names; the item must open with the linked person), and any page that resolves to something other than a real list page is skipped. 107 groups were re-scraped under the corrected rules.
  • —Cleaner `known_for` values. Life-date parentheticals are removed properly rather than leaving an orphaned bracket, so descriptions read singer instead of 1930–1977), singer.

Known gaps, stated plainly: 190 image observations are not joinable to a notable_people row, because the person they describe was removed as a non-person entry during the quality pass. They are excluded from image_observations rather than published with a dangling example_id. Coverage of groups with at least one image observation is 233 (was 239 in v3) for the same reason.

What's in v3

This release adds two structured surfaces on top of v2:

  • —Controlled vocabularies — 22 per-anatomy-category JSON schemas covering 196 dimensions and 853 buckets across the full body (skin, head & face, eyes, lips & mouth, nose, ears, jaw & chin, head shape, head hair, body hair, neck, torso, breast, body shape, butt, arms, hands, legs, feet, plus three internal-only categories). Each dimension is grounded in a peer-reviewed scale (Fitzpatrick, Halls, Heath-Carter, Andre Walker, Manning 2D:4D, Hamilton-Norwood, Ludwig, Mendieta, Cavanagh-Rodgers, etc.) with its citation. The intent is for researchers to use this taxonomy as a re-implementable reference for vision-grounded phenotype analysis. New vocabulary_dimensions config (flat, ~196 rows) plus the full nested JSON schemas at data/vocabularies.jsonl.
  • —Coverage scores on every ethnic group — coverage_score (0–100, blended across data depth + image-observation count + per-dimension fill rate + atlas-category breadth) and coverage_breakdown (full per-component JSON). Surfaces SERP-side as the "Data Depth" card on each /ethnic/{slug} page; here it lets researchers filter for well-covered groups before training.

What's in v2

The previous release added three structured surfaces on top of the original ethnographic catalog:

  • —Synthesized phenotype profiles for 1,777 of 1,779 groups (ethnicities.phenotype_profile column) — editorial-anthropology prose covering skin tone (Fitzpatrick), hair, facial features, and build, drafted from public anthropological sources.
  • —Notable people references — 13K Wikipedia-sourced people names linked back to their ethnic group via ethnic_id, with image URLs and short "known for" descriptors. New notable_people config.
  • —Per-image phenotype observations — 5.6K Wikipedia public-domain photos analyzed via vision LLM (us.anthropic.claude-sonnet-4-6 Bedrock inference profile) into 14 structured phenotype fields per image: skin Fitzpatrick + undertone, hair color/texture/pattern, eye color/shape (incl. epicanthic-fold detection), facial features, build, image quality, confidence (0–1), and per-row source URL. New image_observations config.

The image-observations config is the headline addition of v2. It is, to our knowledge, the largest publicly-released dataset of structured per-image phenotype observations grounded in source-attributable photos. Its rows still carry the legacy 14 fields; re-running the analysis pipeline against v3's 22-category 196-dimension schema remains ahead of us, and is not part of v4.

Source

  • —Live catalog: https://ethnicerotic.com
  • —Browse by region: https://ethnicerotic.com/world
  • —Phenotype atlas: https://ethnicerotic.com/atlas
  • —Editorial reference articles: https://ethnicerotic.com/articles — companion long-form articles describing the peer-reviewed classification scales used in vocabulary_dimensions (see "Companion editorial reference" below)
  • —Reproducibility code (Apache 2.0): https://github.com/Agaveis/phenotype-catalog-pipeline
  • —Markdown corpus: https://ethnicerotic.com/llms-full.txt
  • —Sitemap: https://ethnicerotic.com/sitemap-index.xml

Each row's canonical_url field links back to the site's page for that entity, where you'll find the full editorial, image references, and community discussion.

Configurations

This dataset has six configs, each a separate config_name:

ethnicities (1,779 rows)

One row per ethnic group.

ColumnTypeDescription
idintStable internal identifier
namestringCommon English name of the ethnic group
homelandstringGeographic homeland (free text, may include country + region)
regionstringContinental sub-region (e.g., "East Asia", "Northern Europe")
subgroupstringSub-ethnic group or branch, when applicable
languagestringPrimary language(s) spoken
language_codesstringISO 639 language codes (when known)
iso_countriesstringISO 3166 country codes the group is associated with
religionstringPredominant religion(s)
descriptionstringSite-curated editorial paragraph
phenotype_profilestringNEW (v2) — synthesized phenotype prose: skin Fitzpatrick range, hair, facial features, build. Plain text, paragraph-broken.
image_observed_distributionstringNEW (v2) — aggregated summary of image_observations for the group: sample size, source breakdown, Fitzpatrick distribution, hair/eye distributions, epicanthic-fold percentages, sample-bias caveats. Plain text.
image_observed_atstringNEW (v2) — ISO timestamp of when the observed-distribution summary was last aggregated.
notable_people_countintNEW (v2) — count of rows in notable_people for this group.
coverage_scoreintNEW (v3) — 0–100 blended data-depth score (image count, dimension fill, atlas breadth, profile completeness). Surfaces as "Data Depth" on each /ethnic/{slug} page.
coverage_breakdownstringNEW (v3) — JSON-encoded per-component breakdown of how coverage_score was computed for this row.
coverage_updated_atstringNEW (v3) — ISO timestamp of the last coverage_score recomputation.
wiki_urlstringLink to a Wikipedia article on the group, when known
canonical_urlstringLink to the corresponding page on ethnicerotic.com (blank for corpus_only rows, which have no public page)
corpus_onlyintNEW (v4) — 1 = dataset-only "corpus" group with no public ethnicerotic.com page; 0/absent = normal public group.
Row coverage is uneven. The notable_people, image_observations, coverage_score, and image_observed_distribution assets cover the original ~484-group cohort. The country-composition, greenfield, and corpus_only rows added since v2 carry phenotype_profile prose but few or no notable-people / per-image-observation rows.

atlas (~21 rows)

The phenotype atlas — morphological reference categories (eyes, lips, nose, hair, skin, body composition).

ColumnTypeDescription
idintStable internal identifier
categorystringTop-level phenotype category
sub_categorystringSub-category or named variation
namestringSpecific variation name
descriptionstringEditorial description
value_typestringType of measurement / variation
canonical_urlstringLink to the corresponding atlas page

notable_people (~23K rows) — NEW in v2, expanded in v4

Wikipedia-sourced notable-people references, joined to ethnic groups. Sourced from each group's "List of {Ethnicity} people" article (where one exists). 499 of the 1,779 ethnic-group rows have at least one row here; the rest do not (small / obscure groups, plus the country-composition and corpus-only rows, for which no such Wikipedia list exists).

Every row is validated for person shape (see What's in v4). Groups whose "List of X people" title redirects to the main ethnic article are skipped rather than harvested, so references and prose from those articles do not appear here.

ColumnTypeDescription
idintStable identifier
ethnic_idintFK → ethnicities.id
ethnic_namestringDenormalized ethnic-group name
ethnic_canonical_urlstringLink to the group's catalog page
namestringPerson's common English name
known_forstringShort editorial descriptor (e.g., "Singer-songwriter, 1940s")
birth_yearintBirth year if known
death_yearintDeath year if known
reference_urlstringLink to the source page (typically Wikipedia)
image_urlstringWikipedia image URL when available
source_typestringSource of the reference (currently always wikipedia)

image_observations (~5.5K rows) — NEW in v2

Per-image phenotype observations. Each row is a single Wikipedia-sourced public-domain photo of a notable person, analyzed by a vision LLM into structured phenotype fields. Joinable to notable_people via example_id and to ethnicities via ethnic_id.

ColumnTypeDescription
idintStable identifier
example_idintFK → notable_people.id (the photographed person)
ethnic_idintFK → ethnicities.id
ethnic_namestringDenormalized ethnic-group name
person_namestringDenormalized person's name
image_urlstringSource image URL (typically upload.wikimedia.org/...)
reference_urlstringWikipedia article the image came from
skin_tonestringFitzpatrick range, e.g. Type II-III
skin_undertonestringcool, warm, neutral, or descriptive
hair_colorstringe.g. dark brown, blonde, black with grey
hair_texturestringstraight, wavy, curly, coily, etc.
hair_patternstringMore specific descriptor when present
eye_colorstringbrown, hazel, blue, etc.
eye_shapestringIncludes epicanthic-fold detection: almond, with epicanthic fold, etc.
facial_featuresstringFree-text description of distinctive features
buildstringBody build descriptor when visible
visible_extentstringWhich body parts are visible (head only, head and shoulders, etc.)
image_qualitystringhigh, medium, low, or very_low
obscurationsstringGlasses, hat, makeup, etc.
confidencefloat0.0–1.0 — the model's self-reported confidence in the analysis
analysis_modelstringModel identifier (currently us.anthropic.claude-sonnet-4-6)
analyzed_atstringISO timestamp

vocabulary_dimensions (~196 rows) — NEW in v3

A flat tabular view of the 22 controlled-vocabulary JSON schemas — one row per (category, dimension) pair. The full nested schemas are also published at data/vocabularies.jsonl (22 rows, full structure preserved with bucket definitions and citations).

ColumnTypeDescription
atlas_categorystringOne of 22 anatomical categories (skin, eyes, head-hair, body-shape, breast, arms, hands, legs, feet, torso, neck, face-proportions, lips-and-mouth, nose, ears, jaw-and-chin, head-shape, body-hair, butt, plus internal-only vulva, penis, pubic-region)
uberon_idstringUBERON ontology ID for the anatomical region (e.g. UBERON:0001460 for arms)
fma_idstringFMA (Foundational Model of Anatomy) ID for the region
vocabulary_versionstringSemver of the vocabulary file. The ten head-and-neck files are at 1.1.0 after gaining a reliability record; the other twelve are 1.0.0. No value set changed in that bump, so codings against either are directly comparable
vocabulary_licensestringCC-BY-4.0
dimension_idstringStable identifier for this dimension within its category (e.g. arm_length_proportional)
dimension_namestringHuman-readable name
dimension_typestringcategorical, ordinal, or numeric
scalestringThe peer-reviewed scale this dimension is grounded in (e.g. fitzpatrick_skin_type, halls_extended, andre_walker, cavanagh_rodgers, manning_2d4d_ratio, hamilton_norwood, ludwig_female_pattern, mendieta_buttock_classification, etc.)
scale_citationstringBibliographic citation for the scale, including DOIs where available
unitstringUnit for the 6 numeric dimensions (cm, kg); empty for categorical and ordinal. Empty on every row before v6, see above
descriptionstringOperational definition of the dimension
bucket_countintNumber of categorical/ordinal buckets defined for this dimension
bucket_idsstringPipe-delimited bucket IDs (e.g. `straight\wavy\curly\coily` for hair texture)
observabilitystringHow well the dimension reads from a single photograph: high, medium, low or not_assessable. Empty on every row before v6, see above
requires_unclothedboolWhether assessing this dimension needs an unclothed subject
minimum_visible_extentstringThe least framing that can carry it, e.g. head_only, head_shoulders, full_body
reliability_kappafloatCohen's kappa from the test-retest study; empty where the dimension was not measured, null where chance agreement is 1
reliability_nintPortraits where both readings gave a real value
reliability_verdictstringalmost-perfect, substantial, moderate, fair, slight, uninformative, or insufficient-sample
reliability_studystringThe study id the record belongs to, e.g. test-retest-2026-08-31

The internal-only categories (vulva, penis, pubic-region) are included in the schemas for completeness of the canonical taxonomy but are flagged as not analysed from the photo corpus, which is composed of Wikipedia-sourced public-figure portraits where genital anatomy is not visible.

reader_verdicts — NEW in v5

The only crowd-contributed table in this dataset. These rows are one-tap judgements cast by visitors to the live site, answering a single question about a portrait: does this look like this group?

It is not the only human input. The ethnographic catalog (ethnicities) and the controlled vocabularies are human-curated through an editorial interface, as the Methodology section below describes. image_observations is vision-LLM output and notable_people is scraped from Wikipedia. What is new here is contributions from readers rather than from an editor.

This is a preference label set over synthetic images. Only verdicts on portraits served from our own generated catalog are exported; anything pointing at an externally-hosted image is dropped, because the catalog also references some third-party photographs and a crowd judgement about the appearance of an identifiable real person does not belong in an openly-licensed dataset. That makes this a different corpus from the Wikipedia-sourced portraits behind image_observations. Do not conflate the two.

Generator metadata predates only part of the catalog, so image_generator_recorded marks the rows where the model that produced the portrait was recorded. A 0 there means the provenance was not captured at generation time, not that the image is a photograph.

image_url gives you the portrait each judgement is about. It is licensed separately from this dataset — the labels are CC BY 4.0, the images are not. Read the License section before doing anything with them beyond looking.

ColumnTypeDescription
row_idintSequential row number within this release
image_idintStable identifier of the judged portrait within the catalog
ethnic_idintJoins to ethnicities.id
ethnic_namestringCommon English name of the group the portrait claims to depict
ethnic_slugstringURL slug of the group
corpus_onlyint1 if the group has no public page (dataset-only), 0 otherwise
verdictint1 = reader judged it a good likeness, 0 = not
verdict_labelstringlooks_right or looks_off
featuresstringWhat a looks_off reader said was wrong, as comma-joined tags from a closed vocabulary of four: skin (skin tone), face, hair, setting (clothing or setting), always in that order. Collected since 2026-09-07 through one-tap chips offered after the "looks off" tap; a reader may pick any number or none, and the site drops anything outside the four. Empty means no tag was given, not that nothing was off: rows created before 2026-09-07 predate the ask entirely, and a later looks_off row with no tags is a reader who skipped the chips. Always empty on looks_right rows
surfacestringWhere the judgement was cast: ethnic_page (the group's own page, with its phenotype profile, observed distribution and notable people on screen), country_carousel (a portrait in a country page's faces slider, labelled only with a demographic weight), region_page or cluster_page (one unjudged portrait spotlighted on a region index or a language-family hub, since 2026-08-23), or inline_deck (a portrait handed to the reader in place right after their first tap, judged with no profile or country context in view, since 2026-09-07). Context is the whole variable in a perceived-likeness judgement, so keep these populations separate. Rows predating this column are ethnic_page by construction
rater_idstringPseudonym for one rater, stable within a release so inter-rater agreement is computable. See the privacy note below before treating it as anonymous
is_memberint1 if the rater had an account, 0 if anonymous
image_urlstringThe portrait that was judged. Licensed separately from this dataset — see License.
image_modelstringModel that generated the judged portrait, where recorded
image_providerstringGeneration provider, where recorded
image_generator_recordedint1 if this portrait carries explicit generator metadata, 0 if it predates provenance recording
image_deletedint1 if the portrait has since been removed from the catalog
created_atstringISO 8601 timestamp of the judgement

Privacy. No raw voter key, no IP-derived value and no account identifier is published, ever. The site stores a signed cookie value per browser and a keyed hash of the client IP purely as a vote-stuffing signal; neither is even read by the export.

rater_id is a salted hash, regenerated with a fresh salt on every build. It is a pseudonym, not an anonymity guarantee, and we would rather say so than let you assume otherwise: because each release is a cumulative re-export and created_at is published at millisecond precision, rows can be matched across releases on (image_id, created_at), which recovers the mapping between one release's pseudonyms and the next. Re-salting is not a defence against that. It is a defence against rater_id becoming a durable public identifier in its own right.

Known bias. Judgements come from search visitors to an adult-content site, overwhelmingly mobile, and they are non-expert snap judgements of perceived likeness rather than anthropological assessments. The sample is self-selected toward people who chose to tap.

Context differs by surface, which is why surface is published: an ethnic_page rater has the group's phenotype profile, observed distribution and notable-people list on screen, while a country_carousel rater sees a portrait captioned with a demographic weight and nothing else. Those are different questions in the reader's head. Filter on `surface` before pooling.

generator_renders and generator_pair_votes (published from the 2026-09-21 refresh)

Two configs from the generator comparison layer, in which several text-to-image generators are given the same portrait prompt for the same people and their renders are read by the same vision model that produced image_observations, then compared by readers in pairs. They are written by the build only while the layer is switched on; the layer was switched on 2026-09-16 and both configs are declared in this card's configs: list, so the first weekly refresh after that date carries the files. A refresh from before that date does not have them.

The people is the fixed axis in both tables: rows sort by people then generator, no column ranks peoples, and the comparison is always across generators for one people.

generator_renders, one row per render:

ColumnTypeDescription
render_idintStable identifier of the render
ethnic_id, ethnic_name, ethnic_slug, corpus_onlyAs in reader_verdicts
generatorstringRegistry slug of the generator (recraft-v3, flux-1-1-pro, flux-schnell, stability-core, ideogram-v3-turbo)
providerstringWhere it was called: replicate or bedrock
model_versionstringThe exact model string the call was made with
registerstringPrompt register: traditional, modern-diaspora, modern-urban or mixed
promptstringThe full prompt text, one per people per register
prompt_hashstringsha256 of the prompt, hex; a cell is one people, one register, one hash
seedint or nullThe seed where the call path accepts one
image_urlstringThe render. Licensed separately from this dataset, see License
width, heightintThe provider's output size
analyzedint1 once the vision model has read the render
scorefloat or nullShare of scored dimensions where the render's value fell inside the range the people's documented photographs show; null when the people has no usable baseline
skin_tone, hair_color, hair_texture, eye_color, epicanthic_foldstringThe render's bucketed values, on the same buckets as image_observed_distribution
dims_scored, hitsint or nullHow many dimensions were scored and how many were hits
promotedint1 if an editor copied this render into the catalog as a portrait
scored_at, created_atstringISO 8601

generator_pair_votes, one row per reader answer:

ColumnTypeDescription
row_idintSequential row number within this release
pair_idstringThe two render ids, sorted ascending and joined with a colon, so a pair has one id whichever side each was dealt
left_render_id, right_render_idintThe renders as dealt to this reader
winnerint0 right, 1 left, 2 both, 3 neither
winner_labelstringThe same as a word
left_is_aint1 when the lower render id was dealt on the left; the side is randomized on every showing and recorded, so a left-right preference can be measured
surfacestringinline_deck (the card inside the judging deck) or generator_page (the card on the people's comparison page)
rater_id, is_memberAs in reader_verdicts, same pseudonym rules and the same privacy note
created_atstringISO 8601

Quick start

python
from datasets import load_dataset

# Ethnic groups (1,779 rows, normalized metadata + phenotype profile)
ethnicities = load_dataset("EthnicErotic/phenotype-catalog", "ethnicities", split="train")

# Phenotype atlas (~21 reference categories)
atlas = load_dataset("EthnicErotic/phenotype-catalog", "atlas", split="train")

# Notable people (~23K Wikipedia people, joinable to ethnicities via ethnic_id)
people = load_dataset("EthnicErotic/phenotype-catalog", "notable_people", split="train")

# Per-image phenotype observations (~5.5K rows, joinable to people via example_id)
observations = load_dataset("EthnicErotic/phenotype-catalog", "image_observations", split="train")

# Controlled vocabularies — flat dimension table (196 rows)
vocab_dims = load_dataset("EthnicErotic/phenotype-catalog", "vocabulary_dimensions", split="train")

# Full nested schemas — load directly via the HF Hub file API
from huggingface_hub import hf_hub_download
import json
path = hf_hub_download("EthnicErotic/phenotype-catalog", "data/vocabularies.jsonl", repo_type="dataset")
vocabularies = [json.loads(line) for line in open(path)]

Example join — get all phenotype observations for one ethnic group:

python
import pandas as pd
e = ethnicities.to_pandas()
o = observations.to_pandas()
target = e[e.name == "Punjabis"]
joined = o[o.ethnic_id == target.id.iloc[0]]
print(joined[["person_name", "skin_tone", "hair_color", "eye_color", "confidence"]])

Companion editorial reference

The classification scales referenced in the vocabulary_dimensions.scale column are each documented in a peer-reviewed source publication. For readers who want to understand a scale before using it, a long-form companion article is published on the live catalog covering each. Direct mappings:

Scale referenced in `vocabulary_dimensions.scale`Companion article
fitzpatrick_skin_type + ITA° + Halls undertoneHow skin tone is classified — Fitzpatrick I-VI, ITA° colorimetric, Halls undertone, von Luschan (historical-only)
andre_walker (hair texture) + MC1R / TYRP1 pigmentationHair texture, color, and density across populations — Andre Walker scale, MC1R red-hair genetics, Solomon Islands TYRP1 blondism distinct from European blondism, hair density by population
HERC2/OCA2 iris-color + epicanthic-fold morphologyEye color and morphology — HERC2 rs12913832 single-mutation blue-eye origin, population eye-color distributions, palpebral fissure obliquity
heath_carter somatotype + BMI + WHRBody composition and somatotype — Heath-Carter framework, WHO Asia-Pacific BMI cutoffs, regional fat distribution
Adult stature (NCD-RisC)Adult height across world populations — NCD-RisC 2016 measured-height dataset, secular trend, self-report inflation caveat
Penile dimensions (Veale et al. 2014)Penile dimensions across world populations — Veale 2014 BJUI systematic review (n=15,521), individual-population clinical studies, sampling-bias caveats
mendieta_buttock_classification + breast volumetric studiesBreast morphology — Mendieta classification, Sano/Mugea/Sigurdson volumetric studies, lingerie-sales-data is NOT anthropometry
ISO 9407 Mondopoint + Brannock + foot anthropometryFoot length and shoe size — Mondopoint system, population mean foot lengths from Krauss 2008, cross-system conversion table
All scales (catalog)The peer-reviewed classification scales we use, and why — inventory of every scale documented in vocabulary_dimensions with its source publication and known limitations
Dataset-comprehension primerHow to read the phenotype catalog — what catalog modal/range notation means, what the catalog is NOT (predictive of individuals, racial classification, static, exhaustive)

These articles cite the same source publications listed in vocabulary_dimensions.scale_citation, plus additional population-specific studies. They are editorial reference content (CC BY 4.0 on the live site) — not part of the dataset's structured rows, so the dataset schema is unchanged. Use them when you need plain-language context for what a scale column value actually denotes.

Methodology

Ethnographic catalog

Human-curated. Entries are added and edited through an admin interface against a SQL Server backing store; the dataset here is a snapshot exported via Prisma directly from production. Region taxonomy follows a standard continent → sub-region → ethnicity hierarchy (Americas, Europe, Africa, Asia, Oceania at the continental level; 23 sub-regions). Entries with no name are excluded.

Phenotype profiles (ethnicities.phenotype_profile)

Drafted by Anthropic Claude (Opus / Sonnet) from public anthropological sources, sanitized and stored on the live catalog. Each profile is 300–450 words. The profile is descriptive prose, not a quantitative measurement.

Notable people (notable_people)

Scraped from each group's "List of {Ethnicity} people" Wikipedia article (where one exists), with image URLs sourced from each person's individual Wikipedia article (upload.wikimedia.org). 499 of the 1,779 groups have at least one entry; the rest do not. Entries are person-shape validated and pages that redirect to a non-list article are skipped.

Per-image phenotype observations (image_observations)

Each row is the output of a single vision-LLM call analyzing one Wikipedia public-domain photograph. The model used is the Anthropic Claude Sonnet 4.6 inference profile on AWS Bedrock (us.anthropic.claude-sonnet-4-6). The prompt asks for 14 structured fields plus a self-reported confidence score (0–1) and image-quality bucket. Output is parsed JSON, sanitized into NVARCHAR columns, and stored alongside a raw_json audit trail (not redistributed in this dataset).

Coverage: 5,478 of ~23K notable-people rows have an image_observations row — those for which (a) an image URL was discoverable from Wikipedia, (b) the image fetched successfully under upload.wikimedia.org's rate-limit, and (c) the model returned valid structured output. Average self-reported confidence is 0.67. Quality split: 42% high · 39% medium · 16% low · 3% very-low.

Aggregated observed distributions (ethnicities.image_observed_distribution)

A deterministic SQL/JS aggregator (no LLM) produces a per-group HTML summary card from each group's image_observations rows: sample size, source breakdown, quality split, average confidence, Fitzpatrick distribution, hair color/texture distribution, eye color distribution, epicanthic-fold percentages, and explicit caveats:

  • —Sample-size warning if N < 10 or N < 25
  • —Low-quality warning if < 30% of images are image_quality = high
  • —Low-confidence warning if average confidence < 0.55
  • —Source-bias note when source is 100% Wikipedia, since "Wikipedia notable people" skews male and public-life and is not population-representative.

The HTML version of this summary is rendered on each /ethnic/{slug} page; the dataset's image_observed_distribution column is the same content as plain text.

Positioning vs Wikipedia

This dataset is a structural complement to Wikipedia, not a replacement.

NeedUse
Long-form prose, citations, multilingual articlesWikipedia (follow wiki_url)
Normalized fields for programmatic access (homeland, language family, ISO codes, religion, region)This dataset (ethnicities)
Cross-group queries ("all Bantu-language groups in Central Africa")This dataset (ethnicities)
Phenotype reference (eyes, lips, nose, hair, skin, body morphology)This dataset (atlas)
Vision-grounded per-image phenotype observations at scaleThis dataset (image_observations) — Wikipedia has no equivalent
Notable people, ethnic-groupedThis dataset (notable_people)

Intended uses

  • —Anthropological / ethnographic reference — quick lookup of homeland, language, religion, sub-group affiliations across 1,700+ groups
  • —Computer-vision evaluation — phenotype-balanced image set drawn from public-domain Wikipedia photos with structured ground-truth labels (image_observations)
  • —AI model fine-tuning — multi-ethnic prompt grounding for image generators; Fitzpatrick-balanced training data; multilingual NLP
  • —Bias auditing — image_observations is source-balanced via the ethnic_id column, allowing per-group fairness analysis on downstream vision systems
  • —Educational use — coursework, research projects, geographic / cultural surveys
  • —Editorial and journalism — reference for stories touching ethnic diversity

Out-of-scope uses

  • —The dataset describes ethnic groups in aggregate. It does not describe individuals as ethnic representatives. Each row in image_observations is one observation of one photograph — not a claim about an individual's identity. Individual rows are not appropriate for individual profiling, identification, or surveillance applications.
  • —The dataset is descriptive, not prescriptive. Phenotypes within any ethnic group are diverse; the variations recorded are documented observations, not deterministic claims.
  • —The image_observations config should not be used as ground truth for human classification. The model's self-reported confidence and quality fields are guidance only.

Limitations & biases

  • —Wikipedia source skew. Notable-people rows come from Wikipedia's curatorial choices. Public-life prominence, gender, and Western-language coverage all bias which people end up listed. The aggregated image_observed_distribution column on ethnicities flags this explicitly when source is 100% Wikipedia.
  • —Sample-size variance. 209 of 1,779 groups have ≥3 image observations (the threshold below which image_observed_distribution is omitted). Below that, no aggregate is published; smaller groups should not be inferred from the per-image rows alone.
  • —Model self-report. image_observations.confidence is the model's own estimate, not external validation. Treat as a soft signal.
  • —Image-quality bucket bias. ~16% of images are image_quality = low and another 3% very_low. Filter on image_quality IN ('high','medium') for higher-fidelity subsets.
  • —Anglo-centric framing. Phenotype categories use English anthropological vocabulary. Some self-identifications and exonyms are not captured.
  • —Atlas granularity. The phenotype atlas is descriptive labels, not anthropometric standards.
  • —Editorial review. description and phenotype_profile are LLM-drafted from public sources. They are not peer-reviewed.

License

Creative Commons Attribution 4.0 International (CC BY 4.0)

You are free to use, share, and adapt this dataset for any purpose, including commercial — provided you give appropriate credit. See "Citation" below.

Image URLs in notable_people.image_url and image_observations.image_url reference Wikipedia / Wikimedia Commons content under their own licenses (typically CC-BY-SA, public domain, or various Creative Commons variants). Follow each reference_url to confirm the per-image license before redistributing the actual image data.

The portraits in reader_verdicts are NOT under this license

reader_verdicts.image_url points at portraits generated for, and owned by, Website Design Company, LLC. They are referenced so you can see what each judgement was actually about — a preference label is not verifiable, or much use for training, if the image behind it is invisible.

The CC BY 4.0 grant above covers this dataset: the rows, columns and labels. It does not grant any right in the images themselves. They are not placed in the public domain, not CC-licensed, and not offered for redistribution or for inclusion in derivative image collections. All rights in them are reserved.

Fetching them to inspect or evaluate the labels is expected and fine. Rehosting them, bundling them into another dataset, or redistributing them is not covered by this license — ask first: dmca@ethnicerotic.com.

Two practical notes. The images are adult content, served from a site with an age gate, and nothing in this dataset is intended to route around that. And these URLs are live production assets: portraits get regenerated and re-keyed, so treat any given URL as current-at-release rather than permanent, and use image_id as the stable join key.

The renders in generator_renders are NOT under this license either

When that config is published, generator_renders.image_url will point at images produced by third-party text-to-image generators on our prompts, held under each provider's output terms and owned by Website Design Company, LLC to the extent those terms assign ownership. The CC BY 4.0 grant covers the rows, columns, prompts and scores; it grants no right in the render images. The same practical notes apply as for the portraits above: fetch them to evaluate the labels, do not rehost or bundle them, and ask first at dmca@ethnicerotic.com.

Citation

The pipeline + methodology paper has a Zenodo DOI; please cite both records when using this dataset:

bibtex
@misc{phenotype_catalog_pipeline_2026,
  title         = {phenotype-catalog-pipeline: Wikipedia-sourced per-image phenotype observations across 239 ethnic groups},
  author        = {Jacoby, Jason},
  year          = {2026},
  publisher     = {Zenodo},
  version       = {v1.0.0},
  doi           = {10.5281/zenodo.20075616},
  url           = {https://doi.org/10.5281/zenodo.20075616}
}

@misc{ethnicerotic_phenotype_catalog_2026,
  title         = {PhenotypeCatalog: a public dataset of 5,668 per-image phenotype observations},
  author        = {Jacoby, Jason},
  year          = {2026},
  publisher     = {Hugging Face},
  url           = {https://huggingface.co/datasets/EthnicErotic/phenotype-catalog},
  note          = {Pipeline DOI: \url{https://doi.org/10.5281/zenodo.20075616}; Source: \url{https://ethnicerotic.com}}
}

<!-- begin: featured-groups (auto-generated by scripts/enrich-hf-dataset-readme.mjs) -->

Featured groups

These 30 ethnic groups have the deepest coverage in the catalog (CoverageScore ≥ 30). Each link goes to the live profile page with aggregated phenotype data, notable-people references, demographic context, and citation chain.

Ethnic groupHomelandCoverage
SoninkeMali85
TatarsTatarstan (Russia)85
UzbeksUzbekistan85
TuluvasKarnataka(India)84
IrishIreland (Republic of Ireland, United Kingdom)84
IranunMindanao (Philippines)83
MakassareseSouth Sulawesi (Indonesia)83
IcelandersIceland83
IgboIgboland (Nigeria)82
WelshWales (United Kingdom)82
IbanSarawak (Malaysia)80
BelarusiansBelarus80
Ga-AdangbeGreater Accra (Ghana)79
EstoniansEstonia, Setomaa79
JavaneseJava (Indonesia)79
MinangkabauMinangkabau Highlands (Indonesia)79
MandinkaMali, The Gambia, Guinea, Senegal79
TajiksAfghanistan, Tajikistan, Uzbekistan79
OssetiansSouth Ossetia, North Ossetia-Alania (Russia)78
Kadazan-DusunSabah (Malaysia)78
KikuyuKenya78
GarhwalisUttarakhand (India)78
SusuGuinea, Kambia (Sierra Leone)77
TigrayansEritrean Highlands (Eritrea), Tigrayia (Ethiopia)76
BidayuhSarawak (Malaysia)76
AustriansAustria, South Tyrol76
KambaUkambani (Kenya)76
IlocanoIlocos Region (Philippines)76
BashkirsBashkortostan (Russia)75
BasquesBasque Country (Spain, France)75

Browse all ~1,780 groups at ethnicerotic.com/world. Read the methodology at doi.org/10.5281/zenodo.20075616.

<!-- end: featured-groups -->

Contact

  • —Live catalog: https://ethnicerotic.com
  • —Issues / contributions: https://ethnicerotic.com/community

Updates

This dataset is refreshed from the live source as the catalog grows. See the manifest in the repo root for the last refresh timestamp, schema version, and row counts.