DataikuNLP/kiji-pii-model
Kiji PII Detection Model
Token classification model for detecting Personally Identifiable Information (PII) in text. Fine-tuned from `microsoft/deberta-v3-small` and decoded with a CRF layer for valid BIO sequence prediction.
Model Summary
Architecture
Input (input_ids, attention_mask)
│
DeBERTa-v3 encoder (hidden_size=768)
│
Dropout → Linear(768 → 384) → GELU → Dropout
│
Linear(384 → 53) [BIO emission scores]
│
CRF [valid BIO transitions]
│
Predicted label sequenceThe token classifier emits per-token BIO scores; a learned CRF layer enforces valid transitions (e.g., an I-EMAIL cannot follow a B-PHONENUMBER). The training loss is the CRF negative log-likelihood + 0.2×class-weighted token cross-entropy. At inference time, predictions are produced by Viterbi decoding.
Usage
The repository contains the encoder weights, MLP head, and CRF parameters in a single SafeTensors file. The architecture is custom (PIIDetectionModel) and is not loadable via AutoModelForTokenClassification — see model/src/model.py in the source repository for the head + CRF wiring.
from transformers import AutoTokenizer
from safetensors.torch import load_file
tokenizer = AutoTokenizer.from_pretrained("DataikuNLP/kiji-pii-model")
weights = load_file("model.safetensors") # downloaded from this repo
text = "Contact John Smith at john.smith@example.com or call +1-555-123-4567."
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)
# See label_mappings.json for the BIO label set.PII Labels (BIO tagging)
The model uses BIO tagging with 26 entity types:
Each entity type has B- (beginning) and I- (inside) variants, plus O for non-PII tokens.
Training
Training Data
Trained on the DataikuNLP/kiji-pii-training-data dataset — a synthetic multilingual PII dataset with entity annotations.
Limitations
- Trained on synthetically generated data — may not generalize perfectly to all real-world text
- Optimized for the 6 languages in the training data (English, German, French, Spanish, Dutch, Danish)
- Max sequence length is 512 tokens
- CRF transitions are learned from training data — rare BIO transitions may be underweighted
