Team Ai
Modelpublic

vllm-sr/mmbert-safety-binary-merged

sourceHugging Faceapache-2.0updated 8d agoView on Hugging Face
0likes43downloads
Model Card

MLCommons AI Safety Classifier - Level 1 (Binary)

A LoRA-finetuned multilingual BERT model for binary content safety classification (safe/unsafe), following the MLCommons AI Safety Hazard Taxonomy.

Model Description

This is Level 1 of a hierarchical safety classification system:

  • —Level 1 (this model): Binary classification (safe vs unsafe)
  • —Level 2: 9-class hazard category classification

The model uses mmBERT (Multilingual ModernBERT) as the base, supporting 1800+ languages.

Training Results

MetricValue
Recall86.1%
F1 Score86.5%
False Positive Rate13.1%
Accuracy86.6%

Training Data

Model Architecture & Training

Base Model

LoRA Configuration

ParameterValue
Rank (r)32
Alpha64
Dropout0.1
Target Modulesattn.Wqkv, attn.Wo, mlp.Wi, mlp.Wo
Trainable Parameters6.76M (2.15%)

Training Hyperparameters

ParameterValue
Epochs10
Batch Size64
Learning Rate3e-4
OptimizerAdamW
SchedulerLinear warmup

Hardware & Environment

ComponentSpecification
GPUAMD Instinct MI300X
VRAM192GB HBM3
PlatformROCm 6.2
Containerrocm/pytorch:rocm6.2_ubuntu22.04_py3.10_pytorch_release_2.3.0
Training Time~4 minutes

Usage

python
from transformers import AutoTokenizer, AutoModelForSequenceClassification
from peft import PeftModel

# Load base model and tokenizer
base_model = "jhu-clsp/mmBERT-base"
tokenizer = AutoTokenizer.from_pretrained("vllm-sr/mmbert-safety-binary-merged")
model = AutoModelForSequenceClassification.from_pretrained(base_model, num_labels=2)
model = PeftModel.from_pretrained(model, "vllm-sr/mmbert-safety-binary-merged")

# Classify
text = "How do I make a cake?"
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)
outputs = model(**inputs)
prediction = outputs.logits.argmax(-1).item()
label = "safe" if prediction == 0 else "unsafe"
print(f"Classification: {label}")

Label Mapping

json
{
  "safe": 0,
  "unsafe": 1
}

Intended Use

This model is designed for:

  • —Content moderation pipelines
  • —LLM input/output safety filtering
  • —Jailbreak and prompt injection detection
  • —First-stage filtering before detailed hazard classification

Limitations

  • —Optimized for English but supports 1800+ languages via mmBERT
  • —Should be used as part of a broader safety system
  • —May require domain-specific fine-tuning for specialized applications

Citation

bibtex
@misc{mlcommons-safety-classifier,
  title={MLCommons AI Safety Classifier},
  author={LLM Semantic Router Team},
  year={2026},
  publisher={Hugging Face},
  url={https://huggingface.co/vllm-sr/mmbert-safety-binary-merged}
}

License

Apache 2.0