cometadata/2025-08-datacite-normalized-affiliation-string-distribution
DataCite Normalized Affiliation Distribution Summary normalized_distribution.json contains one JSON object per normalized affiliation string. It aggregates the total occurrence count, a ranked list of the raw affiliation strings that collapse into the normalized form, and the provider/client entities that asserted them. This dataset is derived from the August 2025 DataCite creator/contributor export. Structure { "normalized": "example university"… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/2025-08-datacite-normalized-affiliation-string-distribution.
DataCite Normalized Affiliation Distribution
Summary
normalized_distribution.json contains one JSON object per normalized affiliation string. It aggregates the total occurrence count, a ranked list of the raw affiliation strings that collapse into the normalized form, and the provider/client entities that asserted them. This dataset is derived from the August 2025 DataCite creator/contributor export.
Structure
{
"normalized": "example university",
"total_count": 314,
"affiliations": [
{"affiliation": "Example University", "occurrences": 300},
{"affiliation": "Example Univ.", "occurrences": 14}
],
"providers": {"unique_total": 5, "counts": {"tib.example": 200, "cdr.sample": 114}},
"clients": {"unique_total": 4, "counts": {"example.client": 314}}
}Fields
normalized(string): normalized ASCII/whitespace-stripped token.total_count(int): total occurrences across the snapshot.affiliations(array): ranked list of raw affiliation strings and their counts.providers/clients(object):unique_totalplus per-ID counts.
