Team Ai
Modelpublic

DataikuNLP/kiji-pii-model

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
1likes25downloads
Model Card

Kiji PII Detection Model

Token classification model for detecting Personally Identifiable Information (PII) in text. Fine-tuned from `microsoft/deberta-v3-small` and decoded with a CRF layer for valid BIO sequence prediction.

Model Summary

Base modelmicrosoft/deberta-v3-small
ArchitectureDeBERTa-v3 encoder + MLP token classifier + CRF
Parameters184M
Model size703 MB (SafeTensors)
Hidden size768
TaskPII token classification (53 BIO labels)
PII entity types26
DecoderCRF (Viterbi)
Max sequence length512 tokens

Architecture

Input (input_ids, attention_mask)
        │
  DeBERTa-v3 encoder (hidden_size=768)
        │
  Dropout → Linear(768 → 384) → GELU → Dropout
        │
  Linear(384 → 53)        [BIO emission scores]
        │
  CRF                                  [valid BIO transitions]
        │
  Predicted label sequence

The token classifier emits per-token BIO scores; a learned CRF layer enforces valid transitions (e.g., an I-EMAIL cannot follow a B-PHONENUMBER). The training loss is the CRF negative log-likelihood + 0.2×class-weighted token cross-entropy. At inference time, predictions are produced by Viterbi decoding.

Usage

The repository contains the encoder weights, MLP head, and CRF parameters in a single SafeTensors file. The architecture is custom (PIIDetectionModel) and is not loadable via AutoModelForTokenClassification — see model/src/model.py in the source repository for the head + CRF wiring.

python
from transformers import AutoTokenizer
from safetensors.torch import load_file

tokenizer = AutoTokenizer.from_pretrained("DataikuNLP/kiji-pii-model")
weights = load_file("model.safetensors")  # downloaded from this repo

text = "Contact John Smith at john.smith@example.com or call +1-555-123-4567."
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)
# See label_mappings.json for the BIO label set.

PII Labels (BIO tagging)

The model uses BIO tagging with 26 entity types:

LabelDescription
AGEAge
BUILDINGNUMBuilding number
CITYCity
COMPANYNAMECompany name
COUNTRYCountry
CREDITCARDNUMBERCredit Card Number
DATEOFBIRTHDate of birth
DRIVERLICENSENUMDriver's License Number
EMAILEmail
FIRSTNAMEFirst name
IBANIBAN
IDCARDNUMID Card Number
LICENSEPLATENUMLicense Plate Number
NATIONALIDNational ID
PASSPORTIDPassport ID
PASSWORDPassword
PHONENUMBERPhone number
SECURITYTOKENAPI Security Tokens
SSNSocial Security Number
STATEState
STREETStreet
SURNAMELast name
TAXNUMTax Number
URLURL
USERNAMEUsername
ZIPZip code

Each entity type has B- (beginning) and I- (inside) variants, plus O for non-PII tokens.

Training

Epochs30 (with early stopping)
Batch size128
Learning rate2e-05
Weight decay0.01
Warmup steps500
Precisionbf16 mixed precision
Early stoppingpatience=3, threshold=0.50%
LossCRF NLL + 0.2×class-weighted token cross-entropy
OptimizerAdamW
MetricWeighted F1 (token-level)

Training Data

Trained on the DataikuNLP/kiji-pii-training-data dataset — a synthetic multilingual PII dataset with entity annotations.

Limitations

  • —Trained on synthetically generated data — may not generalize perfectly to all real-world text
  • —Optimized for the 6 languages in the training data (English, German, French, Spanish, Dutch, Danish)
  • —Max sequence length is 512 tokens
  • —CRF transitions are learned from training data — rare BIO transitions may be underweighted