Sunbird/sunflower_language_classification_v2
Sunflower Language ID v2
Language identification across 64 African languages, fine-tuned from google/t5-efficient-tiny with a replacement SentencePiece vocabulary trained on the target corpus.
Why v2
v1 used the stock T5 English-C4 vocabulary, which has no byte fallback and could not represent large parts of this language set — Amharic was 47% <unk>, Yoruba 20%. That damage falls hardest on short input, where there is no redundancy to absorb a lost token. v2 ships a 32k byte-fallback vocabulary trained on balanced text from all 64 languages, measured at 0.00% `<unk>`.
Accuracy by input length
Held-out split, 64 languages, ~16,000 sentences per bucket.
Usage — two things that must match training
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
REPO = "yigagilbert/sunflower_language_classification_v2"
# 1. use_fast=False. The fast conversion drops SentencePiece byte_fallback and
# reintroduces <unk> (0.47% on Amharic), which is the bug this model fixes.
tok = AutoTokenizer.from_pretrained(REPO, use_fast=False)
mdl = AutoModelForSequenceClassification.from_pretrained(REPO).eval()
# 2. Normalise exactly as training did: lowercase + collapse whitespace.
def normalize_text(s):
return " ".join(s.lower().split())
def predict(text):
enc = tok(normalize_text(text), return_tensors="pt", truncation=True, max_length=64)
with torch.no_grad():
probs = mdl(**enc).logits[0].softmax(-1)
top = probs.argmax().item()
return mdl.config.id2label[top], probs[top].item()
print(predict("Webale nyo okutuyamba")) # ('lug', 0.62)Skipping either step silently degrades short-text accuracy — passing raw-case text to a model trained on lowercase is what caused the original regression.
Supported languages
ISO 639-3 codes. "Short" is accuracy on 1-3 word input; chance is ~0.016.
Known limitations
Several pairs are genuinely close to inseparable on short input, and the confusions are symmetric because the languages overlap in reality:
For a single word, expect ~0.41 accuracy across 64 classes (chance is ~0.016). If your deployment only serves a known subset of languages, mask the logits to that subset — it recovers a large amount of accuracy for free.
Training
- Base:
google/t5-efficient-tiny, embeddings re-initialised for the new vocab - 60,000 steps, batch 64, lr 1e-3 cosine, bf16, best checkpoint by short-text (1-3 word) accuracy
- Augmentation: log-uniform random crops biased toward short spans, light character noise
- Logit adjustment (tau=1.0) with the class prior clipped to 50:1
- Languages with under 1,400 unique examples excluded as unlearnable
