Team Ai
Datasetpublic

cometadata/2025-08-datacite-normalized-affiliation-string-distribution

DataCite Normalized Affiliation Distribution Summary normalized_distribution.json contains one JSON object per normalized affiliation string. It aggregates the total occurrence count, a ranked list of the raw affiliation strings that collapse into the normalized form, and the provider/client entities that asserted them. This dataset is derived from the August 2025 DataCite creator/contributor export. Structure { "normalized": "example university"… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/2025-08-datacite-normalized-affiliation-string-distribution.

sourceHugging Facecc0-1.0updated 11mo agoView on Hugging Face
0likes26downloads
Dataset Card

DataCite Normalized Affiliation Distribution

Summary

normalized_distribution.json contains one JSON object per normalized affiliation string. It aggregates the total occurrence count, a ranked list of the raw affiliation strings that collapse into the normalized form, and the provider/client entities that asserted them. This dataset is derived from the August 2025 DataCite creator/contributor export.

Structure

json
{
  "normalized": "example university",
  "total_count": 314,
  "affiliations": [
    {"affiliation": "Example University", "occurrences": 300},
    {"affiliation": "Example Univ.", "occurrences": 14}
  ],
  "providers": {"unique_total": 5, "counts": {"tib.example": 200, "cdr.sample": 114}},
  "clients": {"unique_total": 4, "counts": {"example.client": 314}}
}

Fields

  • —normalized (string): normalized ASCII/whitespace-stripped token.
  • —total_count (int): total occurrences across the snapshot.
  • —affiliations (array): ranked list of raw affiliation strings and their counts.
  • —providers / clients (object): unique_total plus per-ID counts.