clincolnoz/mmbert-small-ner
029
mmBERT-small fine-tuned for multilingual NER
Token classification (PER / ORG / LOC, IOB2) fine-tuned from `jhu-clsp/mmBERT-small` on a 32-language news NER dataset.
Training
- lr 8e-5, up to 10 epochs with early stopping (patience 3) on validation F1, warmup 0.1, weight decay 0.01, max length 256
- Training labels denoised by masking contradictory repeat mentions from the loss
- Uniform weight average of 3 seeds
Validation results
Leak-free, stratified validation split, seqeval micro scores at 256-token truncation.
Per class F1: LOC 0.7804, ORG 0.6953, PER 0.7876.
Usage
from transformers import pipeline
ner = pipeline("token-classification", model="clincolnoz/mmbert-small-ner", aggregation_strategy="simple")
ner("Angela Merkel met executives from Siemens in Munich.")