thealper2/xlm-roberta-base-text-geolocation
xlm-roberta-base-text-geolocation
FacebookAI/xlm-roberta-base fine-tuned end-to-end for multiclass classification of short social-media text into one of 123 geographic regions. Output is a region label with a probability distribution; the model does not predict coordinates.
Labels
yachay/text_coordinates_regions contains one JSON file per region (c_0.json … c_122.json) and no region names. Labels are these file ids (c_0 … c_122). For orientation, the table at the end lists the empirical medoid of each region's training coordinates (descriptive statistic, not an official region definition).
Data
- Source:
yachay/text_coordinates_regions@b9fa48181e3c93791d0613382c77aeb91214f324, 615,000 posts, 5,000 per region (balanced). Coordinates are place-level (≈12k distinct points). - Model input: text only. Coordinates are used for evaluation only.
- Preprocessing: strip, lower-case (corpus is already lower-cased), URLs →
HTTPURL. Mentions, hashtags, emoji and place names are kept.
Remaining: 600,803 posts (2.31% removed).
- Split: 480,496 / 60,103 / 60,204 (train / validation / test), stratified by region, seed 42.
- Leakage control: posts that are identical after removing URLs and collapsing whitespace form a group (4,776 multi-post groups) and are assigned to a single split. Exact-text and group overlap between splits: 0.
Training
Evaluation
Classification (60,204 test posts):
Geographic distance (test). Distance between the post's coordinate and the medoid of the predicted region. "Oracle" uses the medoid of the true region, i.e. the floor for a perfect classifier under this representation.
Usage
from transformers import pipeline
clf = pipeline("text-classification", model="thealper2/xlm-roberta-base-text-geolocation", top_k=5)
print(clf("just landed at jfk, the traffic in queens is already insane"))Apply the training preprocessing (lower-case, URLs → HTTPURL) for best results; settings are stored in preprocessing.json.
Limitations
- Region labels are unnamed dataset clusters; regions have unequal geographic extent.
- Coordinates in the source data are place-level, so distance metrics are coarse.
- Twitter data from 2021: topical, demographic and platform biases; performance on other domains or periods is not measured.
- Many posts carry no geographic signal (emoji-only, generic replies); confidence is low for these.
- Some posts contain explicit place names from app templates (check-ins, "just posted a photo @ …"), which are easy cases.
