Team Ai
Modelpublic

script-langid/fasttext-leaf-surgery-11lang

sourceHugging Facecc-by-nc-4.0updated 12h agoView on Hugging Face
0likes
Model Card

fastText Continual Leaf Surgery: 11-Language Multilingual Rehearsal (Table 3)

This model represents our Continual Learning (Leaf Surgery) architecture trained with the 11-language balanced rehearsal buffer (Table 3 in the research paper):

"A Benchmark for Sinhala-Script Language Identification: Disambiguating Sinhala, Pali, and Sanskrit in Closed and Open World Contexts" University of Moratuwa, Department of Computer Science & Engineering

1. Architecture & Innovation

  • —Hierarchical Softmax Leaf Surgery: The Sinhala (si) leaf node in the official Facebook lid.176.bin Huffman tree was split into Sinhala and Pali branches.
  • —Multilingual Rehearsal: Fine-tuned on the balanced 11-language Aya dataset buffer (3 target languages in Sinhala script + 8 global/anchor languages: English, Tamil, Hindi, Bengali, Arabic, French, German, and Sanskrit Devanagari) to preserve cross-lingual representations.

2. Benchmark Performance (11-Language Evaluation)

  • —WiLI-2018 (Macro-F1): 0.9811 (Accuracy: 98.00%)
  • —CommonLID (Macro-F1): 0.9486 (Accuracy: 96.82%)
  • —FLORES+ (Macro-F1): 0.9395 (Accuracy: 92.71%)

3. How to Load and Run Inference in Python

python
from data_pipeline.fasttext_continual.model import ContinualLID

# Loads weights.pt, config.json, and vocab.json automatically from Hugging Face
model = ContinualLID.from_pretrained("script-langid/fasttext-leaf-surgery-11lang")

# Inference example:
text = "නමෝ තස්ස භගවතෝ අරහතෝ සම්මා සම්බුද්ධස්ස"
prediction, score = model.predict(text)
print(f"Language: {prediction}, Score: {score:.4f}")
# Output: Language: pi, Score: 0.99...