vllm-sr/mmbert32k-feedback-detector-merged
15.9k
mmBERT-32K Feedback Detector (Merged)
A 4-class user feedback classifier based on mmbert-32k-yarn. This is the merged version with LoRA weights integrated - no PEFT library required.
Model Description
This model classifies user messages into 4 feedback categories to help conversational AI systems understand user satisfaction:
Performance
Validation Results (2,985 samples):
Per-Class Performance:
Quick Start
from transformers import AutoModelForSequenceClassification, AutoTokenizer
import torch
# Load model and tokenizer
model = AutoModelForSequenceClassification.from_pretrained(
"vllm-sr/mmbert32k-feedback-detector-merged"
)
tokenizer = AutoTokenizer.from_pretrained(
"vllm-sr/mmbert32k-feedback-detector-merged"
)
model.eval()
# Label mapping
labels = ["SAT", "NEED_CLARIFICATION", "WRONG_ANSWER", "WANT_DIFFERENT"]
# Example inference
texts = [
"Thank you, that's exactly what I needed!",
"I don't understand, can you explain more?",
"That's incorrect, the answer should be different.",
"Can you give me another approach?",
]
for text in texts:
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
outputs = model(**inputs)
probs = torch.softmax(outputs.logits, dim=-1)
pred = outputs.logits.argmax(-1).item()
conf = probs[0][pred].item()
print(f"{labels[pred]:20} ({conf:.1%}) | {text}")Output:
SAT (81.2%) | Thank you, that's exactly what I needed!
NEED_CLARIFICATION (100.0%) | I don't understand, can you explain more?
WRONG_ANSWER (100.0%) | That's incorrect, the answer should be different.
WANT_DIFFERENT (100.0%) | Can you give me another approach?Batch Inference
# Efficient batch processing
texts = ["Your text 1", "Your text 2", "Your text 3"]
inputs = tokenizer(texts, return_tensors="pt", padding=True, truncation=True, max_length=512)
with torch.no_grad():
outputs = model(**inputs)
predictions = outputs.logits.argmax(-1).tolist()
feedback_types = [labels[p] for p in predictions]Training Details
This model was fine-tuned using LoRA (Low-Rank Adaptation) with the following configuration:
Training Data
- Dataset: vllm-sr/feedback-detector-dataset
- Training samples: 17,896 (balanced)
- Validation samples: 2,985
Hardware
- GPU: AMD Instinct MI300X
- Training Time: ~10 minutes
Model Architecture
- Architecture: ModernBERT (Sequence Classification)
- Parameters: ~321M (base) with merged LoRA weights
- Max Context: 32,768 tokens (YaRN-scaled RoPE)
- Hidden Size: 768
- Layers: 22
- Attention Heads: 12
Multilingual Support
Supports 1800+ languages via Glot500 tokenizer. Best performance on:
- English (primary)
- Chinese
- French
- Spanish
Use Cases
- Conversational AI: Detect user satisfaction in chatbots
- Customer Support: Route conversations based on feedback type
- Quality Monitoring: Track user satisfaction trends
- Dialog Systems: Trigger clarification or correction flows
Comparison: LoRA vs Merged
Citation
@misc{mmbert32k-feedback-detector,
title={mmBERT-32K Feedback Detector},
author={LLM Semantic Router Team},
year={2026},
publisher={Hugging Face},
url={https://huggingface.co/vllm-sr/mmbert32k-feedback-detector-merged}
}License
Apache 2.0
