script-langid/fasttext-leaf-surgery-11lang
0
fastText Continual Leaf Surgery: 11-Language Multilingual Rehearsal (Table 3)
This model represents our Continual Learning (Leaf Surgery) architecture trained with the 11-language balanced rehearsal buffer (Table 3 in the research paper):
"A Benchmark for Sinhala-Script Language Identification: Disambiguating Sinhala, Pali, and Sanskrit in Closed and Open World Contexts" University of Moratuwa, Department of Computer Science & Engineering
1. Architecture & Innovation
- Hierarchical Softmax Leaf Surgery: The Sinhala (
si) leaf node in the official Facebooklid.176.binHuffman tree was split into Sinhala and Pali branches. - Multilingual Rehearsal: Fine-tuned on the balanced 11-language Aya dataset buffer (3 target languages in Sinhala script + 8 global/anchor languages: English, Tamil, Hindi, Bengali, Arabic, French, German, and Sanskrit Devanagari) to preserve cross-lingual representations.
2. Benchmark Performance (11-Language Evaluation)
- WiLI-2018 (Macro-F1): 0.9811 (Accuracy: 98.00%)
- CommonLID (Macro-F1): 0.9486 (Accuracy: 96.82%)
- FLORES+ (Macro-F1): 0.9395 (Accuracy: 92.71%)
3. How to Load and Run Inference in Python
from data_pipeline.fasttext_continual.model import ContinualLID
# Loads weights.pt, config.json, and vocab.json automatically from Hugging Face
model = ContinualLID.from_pretrained("script-langid/fasttext-leaf-surgery-11lang")
# Inference example:
text = "නමෝ තස්ස භගවතෝ අරහතෝ සම්මා සම්බුද්ධස්ස"
prediction, score = model.predict(text)
print(f"Language: {prediction}, Score: {score:.4f}")
# Output: Language: pi, Score: 0.99...