Team Ai
Modelpublic

LiquidAI/LFM2.5-Encoder-350M-PII-Detector

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
45likes3.9kdownloads
Model Card

<div align="center"> <img src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/2b08LKpev0DNEk6DlnWkY.png" alt="Liquid AI" style="width: 100%; max-width: 100%; height: auto; display: inline-block; margin-bottom: 0.5em; margin-top: 0.5em;" /> <div style="display: flex; justify-content: center; gap: 0.5em; margin-bottom: 1em;"> <a href="https://playground.liquid.ai/"><strong>Try LFM</strong></a> • <a href="https://docs.liquid.ai/lfm/getting-started/welcome"><strong>Docs</strong></a> • <a href="https://leap.liquid.ai/"><strong>LEAP</strong></a> • <a href="https://discord.com/invite/liquid-ai"><strong>Discord</strong></a> </div> </div>

LFM2.5-Encoder-350-PII-Detector

A full fine-tune of LFM2.5-Encoder-350M with a token-classification head, covering 40 PII types across 16 languages.

Find more details about our encoders in our blog post.

[!NOTE] 💻 Demos: Try this fine-tuned model running in a CPU-only Hugging Face space: [PII detection](https://huggingface.co/spaces/LiquidAI/pii-detection) — spot and remove 40 kinds of personal information across 16 languages.

Entity types (40 PII types across 11 domains)

DomainTypes
Identityidentity.person_name, identity.ssn, identity.national_id, identity.passport, identity.drivers_license, identity.date_of_birth, identity.tax_id
Contactcontact.email, contact.phone, contact.address, contact.postal_code, contact.ip_address
Financialfinancial.credit_card, financial.iban, financial.bank_account, financial.swift_bic, financial.crypto_wallet, financial.amount
Credentialscredential.api_key, credential.password, credential.private_key, credential.jwt, credential.connection_string, developer.login_credentials
Onlineonline.username, online.url
Devicedevice.mac_address, device.imei, developer.device_id
Locationlocation.gps_coordinates
Healthcarehealthcare.medical_record, healthcare.condition, healthcare.medication, healthcare.health_plan_id
Organizationorg.company_name
Special-categoryspecial.religion, special.political, special.orientation, special.health_status
Legallegal.case_number

Benchmarks (18-locale-filtered, partial-F1, hybrid decode)

Benchmark**this model**detection tierSauerkrautLM GLiNERopenai/privacy-filterPiiranha-v1OpenMed privacy-filterregex + validators
SPY0.4280.5090.2800.2640.2320.2260.358
Gretel0.8800.8850.6630.4580.5530.7700.337
TAB0.8670.8880.6850.5430.2620.6720.000
ai4privacy0.7150.7740.4880.3940.9460.4320.195
Nemotron0.8550.8630.6390.5720.6580.9180.335
MAPA0.2360.2670.4160.2880.2280.1640.000

[image]

  • —Best on every benchmark except MAPA, whose idiosyncratic date-as-date_of_birth labeling convention penalises correctly-typed predictions. Only two external scores land higher anywhere — Piiranha-v1's 0.946 on ai4privacy and OpenMed's 0.918 on Nemotron — and both are in-distribution: each model trains on that exact corpus, as did our own encoder's pretraining on those same two.
  • —Detection tier is the same model and the same predictions, scored with the type label ignored — did it find the PII span at all, which is the metric that matters for redaction. The gap to exact-type is the type-confusion rate, e.g. SPY 0.428 → 0.509.

Usage

⚠️ Loads custom code via trust_remote_code=True (the model wraps a trust_remote_code encoder).

Install the required packages:

bash
pip install torch transformers huggingface_hub

Run PII detection:

python
import importlib.util
import sys

from huggingface_hub import hf_hub_download
from transformers import AutoModelForTokenClassification, AutoTokenizer

model_id = "LiquidAI/LFM2.5-Encoder-350-PII-Detector"

helper_path = hf_hub_download(model_id, "pii_hybrid_decode.py")
hf_hub_download(model_id, "context_cued.py")
sys.path.insert(0, helper_path.rsplit("/", 1)[0])

spec = importlib.util.spec_from_file_location("pii_hybrid_decode", helper_path)
hd = importlib.util.module_from_spec(spec)
spec.loader.exec_module(hd)

tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForTokenClassification.from_pretrained(model_id, trust_remote_code=True).eval()

spans = hd.predict("Email Dr. Laura Schmidt at laura@charite.de.", tok, model)
print(spans)

📬 Contact

Citation

bibtex
@article{liquidAI2026Encoders,
  author = {Liquid AI},
  title = {LFM2.5-Encoders: Fast at Long Context, Even on CPU},
  journal = {Liquid AI Blog},
  year = {2026},
  note = {www.liquid.ai/blog/lfm2-5-encoders},
}