Team Ai
Modelpublic

Omared1/LFM2.5-Encoder-350M-PII-Detector

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes15downloads
Model Card

<div align="center">

LFM2.5-Encoder-350-PII-Detector

A full fine-tune of LFM2.5-Encoder-350M with a token-classification head, covering 40 PII types across 16 languages.

Entity types (40 PII types across 11 domains)

DomainTypes
Identityidentity.person_name, identity.ssn, identity.national_id, identity.passport, identity.drivers_license, identity.date_of_birth, identity.tax_id
Contactcontact.email, contact.phone, contact.address, contact.postal_code, contact.ip_address
Financialfinancial.credit_card, financial.iban, financial.bank_account, financial.swift_bic, financial.crypto_wallet, financial.amount
Credentialscredential.api_key, credential.password, credential.private_key, credential.jwt, credential.connection_string, developer.login_credentials
Onlineonline.username, online.url
Devicedevice.mac_address, device.imei, developer.device_id
Locationlocation.gps_coordinates
Healthcarehealthcare.medical_record, healthcare.condition, healthcare.medication, healthcare.health_plan_id
Organizationorg.company_name
Special-categoryspecial.religion, special.political, special.orientation, special.health_status
Legallegal.case_number

Benchmarks (18-locale-filtered, partial-F1, hybrid decode)

Benchmark**this model**detection tierSauerkrautLM GLiNERopenai/privacy-filterPiiranha-v1OpenMed privacy-filterregex + validators
SPY0.4280.5090.2800.2640.2320.2260.358
Gretel0.8800.8850.6630.4580.5530.7700.337
TAB0.8670.8880.6850.5430.2620.6720.000
ai4privacy0.7150.7740.4880.3940.9460.4320.195
Nemotron0.8550.8630.6390.5720.6580.9180.335
MAPA0.2360.2670.4160.2880.2280.1640.000

[image]

  • —Best on every benchmark except MAPA, whose idiosyncratic date-as-date_of_birth labeling convention penalises correctly-typed predictions. Only two external scores land higher anywhere — Piiranha-v1's 0.946 on ai4privacy and OpenMed's 0.918 on Nemotron — and both are in-distribution: each model trains on that exact corpus, as did our own encoder's pretraining on those same two.
  • —Detection tier is the same model and the same predictions, scored with the type label ignored — did it find the PII span at all, which is the metric that matters for redaction. The gap to exact-type is the type-confusion rate, e.g. SPY 0.428 → 0.509.

Usage

⚠️ Loads custom code via trust_remote_code=True (the model wraps a trust_remote_code encoder).

Install the required packages:

bash
pip install torch transformers huggingface_hub

Run PII detection:

python
import importlib.util
import sys

from huggingface_hub import hf_hub_download
from transformers import AutoModelForTokenClassification, AutoTokenizer

model_id = "LiquidAI/LFM2.5-Encoder-350-PII-Detector"

helper_path = hf_hub_download(model_id, "pii_hybrid_decode.py")
hf_hub_download(model_id, "context_cued.py")
sys.path.insert(0, helper_path.rsplit("/", 1)[0])

spec = importlib.util.spec_from_file_location("pii_hybrid_decode", helper_path)
hd = importlib.util.module_from_spec(spec)
spec.loader.exec_module(hd)

tok = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForTokenClassification.from_pretrained(model_id, trust_remote_code=True).eval()

spans = hd.predict("Email Dr. Laura Schmidt at laura@charite.de.", tok, model)
print(spans)