Omared1/LFM2.5-Encoder-350M-PII-Detector
015
<div align="center">
LFM2.5-Encoder-350-PII-Detector
A full fine-tune of LFM2.5-Encoder-350M with a token-classification head, covering 40 PII types across 16 languages.
Entity types (40 PII types across 11 domains)
Benchmarks (18-locale-filtered, partial-F1, hybrid decode)
- Best on every benchmark except MAPA, whose idiosyncratic date-as-
date_of_birthlabeling convention penalises correctly-typed predictions. Only two external scores land higher anywhere — Piiranha-v1's 0.946 on ai4privacy and OpenMed's 0.918 on Nemotron — and both are in-distribution: each model trains on that exact corpus, as did our own encoder's pretraining on those same two. - Detection tier is the same model and the same predictions, scored with the type label ignored — did it find the PII span at all, which is the metric that matters for redaction. The gap to exact-type is the type-confusion rate, e.g. SPY 0.428 → 0.509.
Usage
⚠️ Loads custom code viatrust_remote_code=True(the model wraps atrust_remote_codeencoder).
Install the required packages:
pip install torch transformers huggingface_hubRun PII detection:
import importlib.util
import sys
from huggingface_hub import hf_hub_download
from transformers import AutoModelForTokenClassification, AutoTokenizer
model_id = "LiquidAI/LFM2.5-Encoder-350-PII-Detector"
helper_path = hf_hub_download(model_id, "pii_hybrid_decode.py")
hf_hub_download(model_id, "context_cued.py")
sys.path.insert(0, helper_path.rsplit("/", 1)[0])
spec = importlib.util.spec_from_file_location("pii_hybrid_decode", helper_path)
hd = importlib.util.module_from_spec(spec)
spec.loader.exec_module(hd)
tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForTokenClassification.from_pretrained(model_id, trust_remote_code=True).eval()
spans = hd.predict("Email Dr. Laura Schmidt at laura@charite.de.", tok, model)
print(spans)