LiquidAI/LFM2.5-Encoder-350M-PII-Detector
<div align="center"> <img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/2b08LKpev0DNEk6DlnWkY.png" alt="Liquid AI" style="width: 100%; max-width: 100%; height: auto; display: inline-block; margin-bottom: 0.5em; margin-top: 0.5em;" /> <div style="display: flex; justify-content: center; gap: 0.5em; margin-bottom: 1em;"> <a href="https://playground.liquid.ai/"><strong>Try LFM</strong></a> • <a href="https://docs.liquid.ai/lfm/getting-started/welcome"><strong>Docs</strong></a> • <a href="https://leap.liquid.ai/"><strong>LEAP</strong></a> • <a href="https://discord.com/invite/liquid-ai"><strong>Discord</strong></a> </div> </div>
LFM2.5-Encoder-350-PII-Detector
A full fine-tune of LFM2.5-Encoder-350M with a token-classification head, covering 40 PII types across 16 languages.
Find more details about our encoders in our blog post.
[!NOTE] 💻 Demos: Try this fine-tuned model running in a CPU-only Hugging Face space: [PII detection](https://huggingface.co/spaces/LiquidAI/pii-detection) — spot and remove 40 kinds of personal information across 16 languages.
Entity types (40 PII types across 11 domains)
Benchmarks (18-locale-filtered, partial-F1, hybrid decode)
- Best on every benchmark except MAPA, whose idiosyncratic date-as-
date_of_birthlabeling convention penalises correctly-typed predictions. Only two external scores land higher anywhere — Piiranha-v1's 0.946 on ai4privacy and OpenMed's 0.918 on Nemotron — and both are in-distribution: each model trains on that exact corpus, as did our own encoder's pretraining on those same two. - Detection tier is the same model and the same predictions, scored with the type label ignored — did it find the PII span at all, which is the metric that matters for redaction. The gap to exact-type is the type-confusion rate, e.g. SPY 0.428 → 0.509.
Usage
⚠️ Loads custom code viatrust_remote_code=True(the model wraps atrust_remote_codeencoder).
Install the required packages:
pip install torch transformers huggingface_hubRun PII detection:
import importlib.util
import sys
from huggingface_hub import hf_hub_download
from transformers import AutoModelForTokenClassification, AutoTokenizer
model_id = "LiquidAI/LFM2.5-Encoder-350-PII-Detector"
helper_path = hf_hub_download(model_id, "pii_hybrid_decode.py")
hf_hub_download(model_id, "context_cued.py")
sys.path.insert(0, helper_path.rsplit("/", 1)[0])
spec = importlib.util.spec_from_file_location("pii_hybrid_decode", helper_path)
hd = importlib.util.module_from_spec(spec)
spec.loader.exec_module(hd)
tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForTokenClassification.from_pretrained(model_id, trust_remote_code=True).eval()
spans = hd.predict("Email Dr. Laura Schmidt at laura@charite.de.", tok, model)
print(spans)📬 Contact
- Got questions or want to connect? Join our Discord community
- If you are interested in custom solutions with edge deployment, please contact our sales team.
Citation
@article{liquidAI2026Encoders,
author = {Liquid AI},
title = {LFM2.5-Encoders: Fast at Long Context, Even on CPU},
journal = {Liquid AI Blog},
year = {2026},
note = {www.liquid.ai/blog/lfm2-5-encoders},
}