Team Ai
Datasetpublic

rustemgareev/russian-names

Russian Names with Popularity Scores Description This dataset contains over 12,000 multinational given names in Russia, including their popularity ranks and scores. The data is based on statistics published by the Unified State Register of Civil Status Records (EGR ZAGS) as of July 2025. Usage The dataset can be loaded using the Hugging Face datasets library. from datasets import load_dataset dataset = load_dataset("rustemgareev/russian-names"… See the full description on the dataset page: https://huggingface.co/datasets/rustemgareev/russian-names.

sourceHugging Facecc-by-sa-4.0updated 1y agoView on Hugging Face
0likes208downloads
Dataset Card

Russian Names with Popularity Scores

Description

This dataset contains over 12,000 multinational given names in Russia, including their popularity ranks and scores. The data is based on statistics published by the Unified State Register of Civil Status Records (EGR ZAGS) as of July 2025.

Usage

The dataset can be loaded using the Hugging Face datasets library.

python
from datasets import load_dataset

dataset = load_dataset("rustemgareev/russian-names", split='train')

# Print the first example
print(dataset[0])

Example Output:

json
{
  "name_cyrl": "Аарон",
  "name_latn": "Aaron",
  "gender": "m",
  "popularity_rank": 1389,
  "popularity_score": 0.8093,
  "rarity_rank": 5892,
  "rarity_score": 0.1907
}

Data Structure

Each record in the dataset contains the following fields:

  • —name_cyrl (string): The name in Cyrillic script.
  • —name_latn (string): The name in Latin script.
  • —gender (string): The gender associated with the name: m (male), f (female).
  • —popularity_rank (integer): The popularity rank of the name. A lower value indicates higher popularity.
  • —popularity_score (float): A normalized popularity rank from 0.0 to 1.0. A higher value indicates higher popularity.
  • —rarity_rank (integer): The rarity rank of the name. A lower value indicates higher rarity.
  • —rarity_score (float): A normalized rarity rank from 0.0 to 1.0. A higher value indicates higher rarity.

License

The dataset is distributed under the Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) license.