Team Ai
Modelpublic

clincolnoz/mmbert-small-ner

sourceHugging Faceupdated 18d agoView on Hugging Face
0likes29downloads
Model Card

mmBERT-small fine-tuned for multilingual NER

Token classification (PER / ORG / LOC, IOB2) fine-tuned from `jhu-clsp/mmBERT-small` on a 32-language news NER dataset.

Training

  • —lr 8e-5, up to 10 epochs with early stopping (patience 3) on validation F1, warmup 0.1, weight decay 0.01, max length 256
  • —Training labels denoised by masking contradictory repeat mentions from the loss
  • —Uniform weight average of 3 seeds

Validation results

Leak-free, stratified validation split, seqeval micro scores at 256-token truncation.

metricvalue
F10.7347
precision0.6662
recall0.8189

Per class F1: LOC 0.7804, ORG 0.6953, PER 0.7876.

Usage

python
from transformers import pipeline

ner = pipeline("token-classification", model="clincolnoz/mmbert-small-ner", aggregation_strategy="simple")
ner("Angela Merkel met executives from Siemens in Munich.")