EthnicErotic/phenotype-catalog
Ethnic Erotic Phenotype Catalog A structured complement to Wikipedia for ethnographic data — 1,700+ ethnic groups indexed with normalized linguistic, geographic, cultural, and phenotype metadata, plus 23K+ notable-people references and 5K+ vision-grounded per-image phenotype observations. Curated from the live catalog at ethnicerotic.com and published as an open dataset for anthropological reference, AI training, and ethnographic research. What's in v6 Two columns… See the full description on the dataset page: https://huggingface.co/datasets/EthnicErotic/phenotype-catalog.
Ethnic Erotic Phenotype Catalog
A structured complement to Wikipedia for ethnographic data — 1,700+ ethnic groups indexed with normalized linguistic, geographic, cultural, and phenotype metadata, plus 23K+ notable-people references and 5K+ vision-grounded per-image phenotype observations.
Curated from the live catalog at ethnicerotic.com and published as an open dataset for anthropological reference, AI training, and ethnographic research.
What's in v6
Two columns this card has always documented had never carried a value, and one column carried two types. All three are repaired, and the vocabularies now ship a measured reliability record.
- `observability` was empty on all 196 rows since the config shipped in v3. The exporter read a key (
from_photo) that no dimension has ever had; the real key isfrom_photograph. This card described the column as populated the whole time. It is now populated: 120 high, 41 medium, 6 low, 29 not assessable. - `unit` was empty on all 196 rows for the same class of reason: it read
d.unit, a key that does not exist on any dimension in any vocabulary file. It is now derived from the dimension id for the numeric dimensions and left empty elsewhere, which is what this card always said it was. - `birth_year`, `death_year`, `coverage_score` and `confidence` were mixed-type columns in JSONL. They used an empty string for null, so the published
notable_people.jsonlcarried 3,279 integers and 9,815 strings inbirth_year. They now emit a real null. The CSVs were always correct and are unchanged. - `atlas.description` shipped raw editorial HTML, Bootstrap classes and all, while every other text column in the dataset is flattened to plain text. Now consistent.
- New in `vocabulary_dimensions`:
requires_unclothed,minimum_visible_extent, and four reliability columns (reliability_kappa,reliability_n,reliability_verdict,reliability_study). See the reliability section below.
No column was renamed or removed.
Vocabulary reliability, measured
The ten head-and-neck vocabularies now carry a per-dimension reliability record from a test-retest study run on 2026-08-31: 196 portraits across 29 ethnic groups, re-read with the identical prompt and compared against their stored codings, dimension by dimension.
- 92.9% agreement over dimensions both readings answered.
- 90.6% counting a dimension one reading declined as a miss. Both are given because the choice of denominator has reversed a conclusion on this corpus before.
- 73 of 84 rankable dimensions reach Cohen's kappa 0.60 or better.
- Three are marked `uninformative`: they return a single value for nearly every portrait, so the two readings agree without the dimension distinguishing anything. Ranking on raw agreement would have put them at the top.
This measures reproducibility, not validity. It is one model agreeing with itself, which says the instrument is stable and not that it is correct. No human rater has scored these images, and no inter-rater agreement exists for this vocabulary. Do not cite these numbers as validation, accuracy or ground truth.
Kappa is fragile wherever one value dominates. An earlier 78-portrait read put six of these dimensions at kappa at or below zero; at 196 they read between 0.43 and 0.66, and one read a perfect 1.000 at n=72 and 0.535 at n=176. Every dimension that moved that far has chance agreement above 85%. Those carry estimate_is_fragile in the vocabulary JSON. Read the band, not the digits.
Full per-dimension counts, including the marginals the kappa was computed from, are in the pipeline repository under reliability/.
What's in v4
This release is a coverage expansion and a data-quality pass over notable_people.
- Notable-people coverage nearly doubled — 13,094 → 23,034 rows, and the groups carrying at least one entry went 291 → 499. The largest rosters went to the big diaspora groups (Ukrainian, Tamil, Serbian, Slovak, Scottish), which is where the demand is.
- Person-shape validation. The original scrape had no guard against non-people entries, and a "List of X people" title on Wikipedia is very often a redirect to the main ethnic article — so whole-article harvests contributed references, external links and prose where people were expected. Entries are now validated for person shape (no event, place, ethnonym or year-leading names; the item must open with the linked person), and any page that resolves to something other than a real list page is skipped. 107 groups were re-scraped under the corrected rules.
- Cleaner `known_for` values. Life-date parentheticals are removed properly rather than leaving an orphaned bracket, so descriptions read
singerinstead of1930–1977), singer.
Known gaps, stated plainly: 190 image observations are not joinable to a notable_people row, because the person they describe was removed as a non-person entry during the quality pass. They are excluded from image_observations rather than published with a dangling example_id. Coverage of groups with at least one image observation is 233 (was 239 in v3) for the same reason.
What's in v3
This release adds two structured surfaces on top of v2:
- Controlled vocabularies — 22 per-anatomy-category JSON schemas covering 196 dimensions and 853 buckets across the full body (skin, head & face, eyes, lips & mouth, nose, ears, jaw & chin, head shape, head hair, body hair, neck, torso, breast, body shape, butt, arms, hands, legs, feet, plus three internal-only categories). Each dimension is grounded in a peer-reviewed scale (Fitzpatrick, Halls, Heath-Carter, Andre Walker, Manning 2D:4D, Hamilton-Norwood, Ludwig, Mendieta, Cavanagh-Rodgers, etc.) with its citation. The intent is for researchers to use this taxonomy as a re-implementable reference for vision-grounded phenotype analysis. New
vocabulary_dimensionsconfig (flat, ~196 rows) plus the full nested JSON schemas atdata/vocabularies.jsonl. - Coverage scores on every ethnic group —
coverage_score(0–100, blended across data depth + image-observation count + per-dimension fill rate + atlas-category breadth) andcoverage_breakdown(full per-component JSON). Surfaces SERP-side as the "Data Depth" card on each/ethnic/{slug}page; here it lets researchers filter for well-covered groups before training.
What's in v2
The previous release added three structured surfaces on top of the original ethnographic catalog:
- Synthesized phenotype profiles for 1,777 of 1,779 groups (
ethnicities.phenotype_profilecolumn) — editorial-anthropology prose covering skin tone (Fitzpatrick), hair, facial features, and build, drafted from public anthropological sources. - Notable people references — 13K Wikipedia-sourced people names linked back to their ethnic group via
ethnic_id, with image URLs and short "known for" descriptors. Newnotable_peopleconfig. - Per-image phenotype observations — 5.6K Wikipedia public-domain photos analyzed via vision LLM (
us.anthropic.claude-sonnet-4-6Bedrock inference profile) into 14 structured phenotype fields per image: skin Fitzpatrick + undertone, hair color/texture/pattern, eye color/shape (incl. epicanthic-fold detection), facial features, build, image quality, confidence (0–1), and per-row source URL. Newimage_observationsconfig.
The image-observations config is the headline addition of v2. It is, to our knowledge, the largest publicly-released dataset of structured per-image phenotype observations grounded in source-attributable photos. Its rows still carry the legacy 14 fields; re-running the analysis pipeline against v3's 22-category 196-dimension schema remains ahead of us, and is not part of v4.
Source
- Live catalog: https://ethnicerotic.com
- Browse by region: https://ethnicerotic.com/world
- Phenotype atlas: https://ethnicerotic.com/atlas
- Editorial reference articles: https://ethnicerotic.com/articles — companion long-form articles describing the peer-reviewed classification scales used in
vocabulary_dimensions(see "Companion editorial reference" below) - Reproducibility code (Apache 2.0): https://github.com/Agaveis/phenotype-catalog-pipeline
- Markdown corpus: https://ethnicerotic.com/llms-full.txt
- Sitemap: https://ethnicerotic.com/sitemap-index.xml
Each row's canonical_url field links back to the site's page for that entity, where you'll find the full editorial, image references, and community discussion.
Configurations
This dataset has six configs, each a separate config_name:
ethnicities (1,779 rows)
One row per ethnic group.
Row coverage is uneven. Thenotable_people,image_observations,coverage_score, andimage_observed_distributionassets cover the original ~484-group cohort. The country-composition, greenfield, andcorpus_onlyrows added since v2 carryphenotype_profileprose but few or no notable-people / per-image-observation rows.
atlas (~21 rows)
The phenotype atlas — morphological reference categories (eyes, lips, nose, hair, skin, body composition).
notable_people (~23K rows) — NEW in v2, expanded in v4
Wikipedia-sourced notable-people references, joined to ethnic groups. Sourced from each group's "List of {Ethnicity} people" article (where one exists). 499 of the 1,779 ethnic-group rows have at least one row here; the rest do not (small / obscure groups, plus the country-composition and corpus-only rows, for which no such Wikipedia list exists).
Every row is validated for person shape (see What's in v4). Groups whose "List of X people" title redirects to the main ethnic article are skipped rather than harvested, so references and prose from those articles do not appear here.
image_observations (~5.5K rows) — NEW in v2
Per-image phenotype observations. Each row is a single Wikipedia-sourced public-domain photo of a notable person, analyzed by a vision LLM into structured phenotype fields. Joinable to notable_people via example_id and to ethnicities via ethnic_id.
vocabulary_dimensions (~196 rows) — NEW in v3
A flat tabular view of the 22 controlled-vocabulary JSON schemas — one row per (category, dimension) pair. The full nested schemas are also published at data/vocabularies.jsonl (22 rows, full structure preserved with bucket definitions and citations).
The internal-only categories (vulva, penis, pubic-region) are included in the schemas for completeness of the canonical taxonomy but are flagged as not analysed from the photo corpus, which is composed of Wikipedia-sourced public-figure portraits where genital anatomy is not visible.
reader_verdicts — NEW in v5
The only crowd-contributed table in this dataset. These rows are one-tap judgements cast by visitors to the live site, answering a single question about a portrait: does this look like this group?
It is not the only human input. The ethnographic catalog (ethnicities) and the controlled vocabularies are human-curated through an editorial interface, as the Methodology section below describes. image_observations is vision-LLM output and notable_people is scraped from Wikipedia. What is new here is contributions from readers rather than from an editor.
This is a preference label set over synthetic images. Only verdicts on portraits served from our own generated catalog are exported; anything pointing at an externally-hosted image is dropped, because the catalog also references some third-party photographs and a crowd judgement about the appearance of an identifiable real person does not belong in an openly-licensed dataset. That makes this a different corpus from the Wikipedia-sourced portraits behind image_observations. Do not conflate the two.
Generator metadata predates only part of the catalog, so image_generator_recorded marks the rows where the model that produced the portrait was recorded. A 0 there means the provenance was not captured at generation time, not that the image is a photograph.
image_url gives you the portrait each judgement is about. It is licensed separately from this dataset — the labels are CC BY 4.0, the images are not. Read the License section before doing anything with them beyond looking.
Privacy. No raw voter key, no IP-derived value and no account identifier is published, ever. The site stores a signed cookie value per browser and a keyed hash of the client IP purely as a vote-stuffing signal; neither is even read by the export.
rater_id is a salted hash, regenerated with a fresh salt on every build. It is a pseudonym, not an anonymity guarantee, and we would rather say so than let you assume otherwise: because each release is a cumulative re-export and created_at is published at millisecond precision, rows can be matched across releases on (image_id, created_at), which recovers the mapping between one release's pseudonyms and the next. Re-salting is not a defence against that. It is a defence against rater_id becoming a durable public identifier in its own right.
Known bias. Judgements come from search visitors to an adult-content site, overwhelmingly mobile, and they are non-expert snap judgements of perceived likeness rather than anthropological assessments. The sample is self-selected toward people who chose to tap.
Context differs by surface, which is why surface is published: an ethnic_page rater has the group's phenotype profile, observed distribution and notable-people list on screen, while a country_carousel rater sees a portrait captioned with a demographic weight and nothing else. Those are different questions in the reader's head. Filter on `surface` before pooling.
generator_renders and generator_pair_votes (published from the 2026-09-21 refresh)
Two configs from the generator comparison layer, in which several text-to-image generators are given the same portrait prompt for the same people and their renders are read by the same vision model that produced image_observations, then compared by readers in pairs. They are written by the build only while the layer is switched on; the layer was switched on 2026-09-16 and both configs are declared in this card's configs: list, so the first weekly refresh after that date carries the files. A refresh from before that date does not have them.
The people is the fixed axis in both tables: rows sort by people then generator, no column ranks peoples, and the comparison is always across generators for one people.
generator_renders, one row per render:
generator_pair_votes, one row per reader answer:
Quick start
from datasets import load_dataset
# Ethnic groups (1,779 rows, normalized metadata + phenotype profile)
ethnicities = load_dataset("EthnicErotic/phenotype-catalog", "ethnicities", split="train")
# Phenotype atlas (~21 reference categories)
atlas = load_dataset("EthnicErotic/phenotype-catalog", "atlas", split="train")
# Notable people (~23K Wikipedia people, joinable to ethnicities via ethnic_id)
people = load_dataset("EthnicErotic/phenotype-catalog", "notable_people", split="train")
# Per-image phenotype observations (~5.5K rows, joinable to people via example_id)
observations = load_dataset("EthnicErotic/phenotype-catalog", "image_observations", split="train")
# Controlled vocabularies — flat dimension table (196 rows)
vocab_dims = load_dataset("EthnicErotic/phenotype-catalog", "vocabulary_dimensions", split="train")
# Full nested schemas — load directly via the HF Hub file API
from huggingface_hub import hf_hub_download
import json
path = hf_hub_download("EthnicErotic/phenotype-catalog", "data/vocabularies.jsonl", repo_type="dataset")
vocabularies = [json.loads(line) for line in open(path)]Example join — get all phenotype observations for one ethnic group:
import pandas as pd
e = ethnicities.to_pandas()
o = observations.to_pandas()
target = e[e.name == "Punjabis"]
joined = o[o.ethnic_id == target.id.iloc[0]]
print(joined[["person_name", "skin_tone", "hair_color", "eye_color", "confidence"]])Companion editorial reference
The classification scales referenced in the vocabulary_dimensions.scale column are each documented in a peer-reviewed source publication. For readers who want to understand a scale before using it, a long-form companion article is published on the live catalog covering each. Direct mappings:
These articles cite the same source publications listed in vocabulary_dimensions.scale_citation, plus additional population-specific studies. They are editorial reference content (CC BY 4.0 on the live site) — not part of the dataset's structured rows, so the dataset schema is unchanged. Use them when you need plain-language context for what a scale column value actually denotes.
Methodology
Ethnographic catalog
Human-curated. Entries are added and edited through an admin interface against a SQL Server backing store; the dataset here is a snapshot exported via Prisma directly from production. Region taxonomy follows a standard continent → sub-region → ethnicity hierarchy (Americas, Europe, Africa, Asia, Oceania at the continental level; 23 sub-regions). Entries with no name are excluded.
Phenotype profiles (ethnicities.phenotype_profile)
Drafted by Anthropic Claude (Opus / Sonnet) from public anthropological sources, sanitized and stored on the live catalog. Each profile is 300–450 words. The profile is descriptive prose, not a quantitative measurement.
Notable people (notable_people)
Scraped from each group's "List of {Ethnicity} people" Wikipedia article (where one exists), with image URLs sourced from each person's individual Wikipedia article (upload.wikimedia.org). 499 of the 1,779 groups have at least one entry; the rest do not. Entries are person-shape validated and pages that redirect to a non-list article are skipped.
Per-image phenotype observations (image_observations)
Each row is the output of a single vision-LLM call analyzing one Wikipedia public-domain photograph. The model used is the Anthropic Claude Sonnet 4.6 inference profile on AWS Bedrock (us.anthropic.claude-sonnet-4-6). The prompt asks for 14 structured fields plus a self-reported confidence score (0–1) and image-quality bucket. Output is parsed JSON, sanitized into NVARCHAR columns, and stored alongside a raw_json audit trail (not redistributed in this dataset).
Coverage: 5,478 of ~23K notable-people rows have an image_observations row — those for which (a) an image URL was discoverable from Wikipedia, (b) the image fetched successfully under upload.wikimedia.org's rate-limit, and (c) the model returned valid structured output. Average self-reported confidence is 0.67. Quality split: 42% high · 39% medium · 16% low · 3% very-low.
Aggregated observed distributions (ethnicities.image_observed_distribution)
A deterministic SQL/JS aggregator (no LLM) produces a per-group HTML summary card from each group's image_observations rows: sample size, source breakdown, quality split, average confidence, Fitzpatrick distribution, hair color/texture distribution, eye color distribution, epicanthic-fold percentages, and explicit caveats:
- Sample-size warning if N < 10 or N < 25
- Low-quality warning if < 30% of images are
image_quality = high - Low-confidence warning if average confidence < 0.55
- Source-bias note when source is 100% Wikipedia, since "Wikipedia notable people" skews male and public-life and is not population-representative.
The HTML version of this summary is rendered on each /ethnic/{slug} page; the dataset's image_observed_distribution column is the same content as plain text.
Positioning vs Wikipedia
This dataset is a structural complement to Wikipedia, not a replacement.
Intended uses
- Anthropological / ethnographic reference — quick lookup of homeland, language, religion, sub-group affiliations across 1,700+ groups
- Computer-vision evaluation — phenotype-balanced image set drawn from public-domain Wikipedia photos with structured ground-truth labels (
image_observations) - AI model fine-tuning — multi-ethnic prompt grounding for image generators; Fitzpatrick-balanced training data; multilingual NLP
- Bias auditing —
image_observationsis source-balanced via theethnic_idcolumn, allowing per-group fairness analysis on downstream vision systems - Educational use — coursework, research projects, geographic / cultural surveys
- Editorial and journalism — reference for stories touching ethnic diversity
Out-of-scope uses
- The dataset describes ethnic groups in aggregate. It does not describe individuals as ethnic representatives. Each row in
image_observationsis one observation of one photograph — not a claim about an individual's identity. Individual rows are not appropriate for individual profiling, identification, or surveillance applications. - The dataset is descriptive, not prescriptive. Phenotypes within any ethnic group are diverse; the variations recorded are documented observations, not deterministic claims.
- The
image_observationsconfig should not be used as ground truth for human classification. The model's self-reported confidence and quality fields are guidance only.
Limitations & biases
- Wikipedia source skew. Notable-people rows come from Wikipedia's curatorial choices. Public-life prominence, gender, and Western-language coverage all bias which people end up listed. The aggregated
image_observed_distributioncolumn onethnicitiesflags this explicitly when source is 100% Wikipedia. - Sample-size variance. 209 of 1,779 groups have ≥3 image observations (the threshold below which
image_observed_distributionis omitted). Below that, no aggregate is published; smaller groups should not be inferred from the per-image rows alone. - Model self-report.
image_observations.confidenceis the model's own estimate, not external validation. Treat as a soft signal. - Image-quality bucket bias. ~16% of images are
image_quality = lowand another 3%very_low. Filter onimage_quality IN ('high','medium')for higher-fidelity subsets. - Anglo-centric framing. Phenotype categories use English anthropological vocabulary. Some self-identifications and exonyms are not captured.
- Atlas granularity. The phenotype atlas is descriptive labels, not anthropometric standards.
- Editorial review.
descriptionandphenotype_profileare LLM-drafted from public sources. They are not peer-reviewed.
License
Creative Commons Attribution 4.0 International (CC BY 4.0)
You are free to use, share, and adapt this dataset for any purpose, including commercial — provided you give appropriate credit. See "Citation" below.
Image URLs in notable_people.image_url and image_observations.image_url reference Wikipedia / Wikimedia Commons content under their own licenses (typically CC-BY-SA, public domain, or various Creative Commons variants). Follow each reference_url to confirm the per-image license before redistributing the actual image data.
The portraits in reader_verdicts are NOT under this license
reader_verdicts.image_url points at portraits generated for, and owned by, Website Design Company, LLC. They are referenced so you can see what each judgement was actually about — a preference label is not verifiable, or much use for training, if the image behind it is invisible.
The CC BY 4.0 grant above covers this dataset: the rows, columns and labels. It does not grant any right in the images themselves. They are not placed in the public domain, not CC-licensed, and not offered for redistribution or for inclusion in derivative image collections. All rights in them are reserved.
Fetching them to inspect or evaluate the labels is expected and fine. Rehosting them, bundling them into another dataset, or redistributing them is not covered by this license — ask first: dmca@ethnicerotic.com.
Two practical notes. The images are adult content, served from a site with an age gate, and nothing in this dataset is intended to route around that. And these URLs are live production assets: portraits get regenerated and re-keyed, so treat any given URL as current-at-release rather than permanent, and use image_id as the stable join key.
The renders in generator_renders are NOT under this license either
When that config is published, generator_renders.image_url will point at images produced by third-party text-to-image generators on our prompts, held under each provider's output terms and owned by Website Design Company, LLC to the extent those terms assign ownership. The CC BY 4.0 grant covers the rows, columns, prompts and scores; it grants no right in the render images. The same practical notes apply as for the portraits above: fetch them to evaluate the labels, do not rehost or bundle them, and ask first at dmca@ethnicerotic.com.
Citation
The pipeline + methodology paper has a Zenodo DOI; please cite both records when using this dataset:
@misc{phenotype_catalog_pipeline_2026,
title = {phenotype-catalog-pipeline: Wikipedia-sourced per-image phenotype observations across 239 ethnic groups},
author = {Jacoby, Jason},
year = {2026},
publisher = {Zenodo},
version = {v1.0.0},
doi = {10.5281/zenodo.20075616},
url = {https://doi.org/10.5281/zenodo.20075616}
}
@misc{ethnicerotic_phenotype_catalog_2026,
title = {PhenotypeCatalog: a public dataset of 5,668 per-image phenotype observations},
author = {Jacoby, Jason},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/datasets/EthnicErotic/phenotype-catalog},
note = {Pipeline DOI: \url{https://doi.org/10.5281/zenodo.20075616}; Source: \url{https://ethnicerotic.com}}
}<!-- begin: featured-groups (auto-generated by scripts/enrich-hf-dataset-readme.mjs) -->
Featured groups
These 30 ethnic groups have the deepest coverage in the catalog (CoverageScore ≥ 30). Each link goes to the live profile page with aggregated phenotype data, notable-people references, demographic context, and citation chain.
Browse all ~1,780 groups at ethnicerotic.com/world. Read the methodology at doi.org/10.5281/zenodo.20075616.
<!-- end: featured-groups -->
Contact
- Live catalog: https://ethnicerotic.com
- Issues / contributions: https://ethnicerotic.com/community
Updates
This dataset is refreshed from the live source as the catalog grows. See the manifest in the repo root for the last refresh timestamp, schema version, and row counts.
