vllm-sr/mmbert-safety-binary-merged
043
MLCommons AI Safety Classifier - Level 1 (Binary)
A LoRA-finetuned multilingual BERT model for binary content safety classification (safe/unsafe), following the MLCommons AI Safety Hazard Taxonomy.
Model Description
This is Level 1 of a hierarchical safety classification system:
- Level 1 (this model): Binary classification (safe vs unsafe)
- Level 2: 9-class hazard category classification
The model uses mmBERT (Multilingual ModernBERT) as the base, supporting 1800+ languages.
Training Results
Training Data
- Total samples: 20,000 (balanced)
- Safe: 10,000
- Unsafe: 10,000
- Sources:
- AEGIS AI Content Safety Dataset 2.0
- MLCommons AI Safety Synth
Model Architecture & Training
Base Model
- Model: jhu-clsp/mmBERT-base
- Architecture: ModernBERT (314M parameters)
LoRA Configuration
Training Hyperparameters
Hardware & Environment
Usage
from transformers import AutoTokenizer, AutoModelForSequenceClassification
from peft import PeftModel
# Load base model and tokenizer
base_model = "jhu-clsp/mmBERT-base"
tokenizer = AutoTokenizer.from_pretrained("vllm-sr/mmbert-safety-binary-merged")
model = AutoModelForSequenceClassification.from_pretrained(base_model, num_labels=2)
model = PeftModel.from_pretrained(model, "vllm-sr/mmbert-safety-binary-merged")
# Classify
text = "How do I make a cake?"
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)
outputs = model(**inputs)
prediction = outputs.logits.argmax(-1).item()
label = "safe" if prediction == 0 else "unsafe"
print(f"Classification: {label}")Label Mapping
{
"safe": 0,
"unsafe": 1
}Intended Use
This model is designed for:
- Content moderation pipelines
- LLM input/output safety filtering
- Jailbreak and prompt injection detection
- First-stage filtering before detailed hazard classification
Limitations
- Optimized for English but supports 1800+ languages via mmBERT
- Should be used as part of a broader safety system
- May require domain-specific fine-tuning for specialized applications
Citation
@misc{mlcommons-safety-classifier,
title={MLCommons AI Safety Classifier},
author={LLM Semantic Router Team},
year={2026},
publisher={Hugging Face},
url={https://huggingface.co/vllm-sr/mmbert-safety-binary-merged}
}License
Apache 2.0
