Team Ai
Modelpublic

Sunbird/sunflower_language_classification_v2

sourceHugging Faceapache-2.0updated 26d agoView on Hugging Face
0likes548downloads
Model Card

Sunflower Language ID v2

Language identification across 64 African languages, fine-tuned from google/t5-efficient-tiny with a replacement SentencePiece vocabulary trained on the target corpus.

Why v2

v1 used the stock T5 English-C4 vocabulary, which has no byte fallback and could not represent large parts of this language set — Amharic was 47% <unk>, Yoruba 20%. That damage falls hardest on short input, where there is no redundancy to absorb a lost token. v2 ships a 32k byte-fallback vocabulary trained on balanced text from all 64 languages, measured at 0.00% `<unk>`.

Accuracy by input length

Held-out split, 64 languages, ~16,000 sentences per bucket.

Input lengthv1v2
1 word0.2010.405
2 words0.4770.597
3 words0.6140.701
5 words0.7490.805
8 words0.8320.864
full sentence0.9070.928

Usage — two things that must match training

python
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

REPO = "yigagilbert/sunflower_language_classification_v2"

# 1. use_fast=False. The fast conversion drops SentencePiece byte_fallback and
#    reintroduces <unk> (0.47% on Amharic), which is the bug this model fixes.
tok = AutoTokenizer.from_pretrained(REPO, use_fast=False)
mdl = AutoModelForSequenceClassification.from_pretrained(REPO).eval()

# 2. Normalise exactly as training did: lowercase + collapse whitespace.
def normalize_text(s):
    return " ".join(s.lower().split())

def predict(text):
    enc = tok(normalize_text(text), return_tensors="pt", truncation=True, max_length=64)
    with torch.no_grad():
        probs = mdl(**enc).logits[0].softmax(-1)
    top = probs.argmax().item()
    return mdl.config.id2label[top], probs[top].item()

print(predict("Webale nyo okutuyamba"))   # ('lug', 0.62)

Skipping either step silently degrades short-text accuracy — passing raw-case text to a model trained on lowercase is what caused the original regression.

Supported languages

ISO 639-3 codes. "Short" is accuracy on 1-3 word input; chance is ~0.016.

CodeLanguageShortCodeLanguageShort
achAcholi0.605lucAringa0.479
adhAdhola (Jopadhola)0.292lugLuganda0.692
afrAfrikaans0.866luoDholuo0.705
akaAkan0.505luyLuyia0.228
alzAlur0.453mhiMa'di0.268
amhAmharic0.940mlgMalagasy0.906
bamBambara0.476myxMasaaba0.727
bemBemba0.746nblSouthern Ndebele0.579
bfaBari0.468nujNyole0.338
cggChiga (Rukiga)0.193nyaNyanja (Chichewa)0.810
dagDagbani0.485nynNyankole0.582
dinDinka0.696nyoNyoro (Runyoro)0.203
engEnglish0.698ormOromo0.838
eweEwe0.809pcmNigerian Pidgin0.564
fraFrench0.655pokPokoot0.047
fulFulah0.511rubGungu0.439
gwrGwere0.256rucRuuli0.114
hauHausa0.709runRundi (Kirundi)0.627
iboIgbo0.819rwmAmba0.031
kabKabyle0.674snaShona0.813
kauKanuri0.321somSomali0.864
kdiKumam0.376sotSouthern Sotho0.556
kdjKaramojong0.432swaSwahili0.814
keoKakwa0.421teoTeso (Ateso)0.793
kikKikuyu0.796tljTalinga-Bwisi0.417
kinKinyarwanda0.615tsnTswana0.747
kooKonzo0.702ttjTooro (Rutooro)0.523
kpzKupsabiny0.667wolWolof0.669
lajLango0.252xhoXhosa0.635
lggLugbara0.593xogSoga (Lusoga)0.497
linLingala0.845yorYoruba0.885
lsmSaamia0.401zulZulu0.640

Known limitations

Several pairs are genuinely close to inseparable on short input, and the confusions are symmetric because the languages overlap in reality:

PairRateNote
cgg to nyn34.2%Chiga/Nyankole, often treated as one language
nyo to ttj30.4%Nyoro/Tooro, likewise
run to kin19.7%Kirundi/Kinyarwanda, mutually intelligible
nbl to zul18.6%Nguni cluster
pcm to eng14.0%Nigerian Pidgin vs English

For a single word, expect ~0.41 accuracy across 64 classes (chance is ~0.016). If your deployment only serves a known subset of languages, mask the logits to that subset — it recovers a large amount of accuracy for free.

Training

  • —Base: google/t5-efficient-tiny, embeddings re-initialised for the new vocab
  • —60,000 steps, batch 64, lr 1e-3 cosine, bf16, best checkpoint by short-text (1-3 word) accuracy
  • —Augmentation: log-uniform random crops biased toward short spans, light character noise
  • —Logit adjustment (tau=1.0) with the class prior clipped to 50:1
  • —Languages with under 1,400 unique examples excluded as unlearnable