Team Ai
Modelpublic

vllm-sr/mmbert32k-feedback-detector-merged

sourceHugging Faceapache-2.0updated 4d agoView on Hugging Face
1likes5.9kdownloads
Model Card

mmBERT-32K Feedback Detector (Merged)

A 4-class user feedback classifier based on mmbert-32k-yarn. This is the merged version with LoRA weights integrated - no PEFT library required.

Model Description

This model classifies user messages into 4 feedback categories to help conversational AI systems understand user satisfaction:

LabelIDDescription
SAT0User is satisfied with the response
NEED_CLARIFICATION1User needs more explanation or details
WRONG_ANSWER2User indicates the response was incorrect
WANT_DIFFERENT3User wants an alternative approach/answer

Performance

Validation Results (2,985 samples):

MetricValue
Accuracy98.83%
F1 (macro)98.24%
F1 (weighted)98.83%

Per-Class Performance:

ClassPrecisionRecallF1-ScoreSupport
SAT1.00001.00001.00001,491
NEED_CLARIFICATION0.99800.99800.9980498
WRONG_ANSWER0.96040.97390.9671498
WANT_DIFFERENT0.97150.95780.9646498

Quick Start

python
from transformers import AutoModelForSequenceClassification, AutoTokenizer
import torch

# Load model and tokenizer
model = AutoModelForSequenceClassification.from_pretrained(
    "vllm-sr/mmbert32k-feedback-detector-merged"
)
tokenizer = AutoTokenizer.from_pretrained(
    "vllm-sr/mmbert32k-feedback-detector-merged"
)
model.eval()

# Label mapping
labels = ["SAT", "NEED_CLARIFICATION", "WRONG_ANSWER", "WANT_DIFFERENT"]

# Example inference
texts = [
    "Thank you, that's exactly what I needed!",
    "I don't understand, can you explain more?",
    "That's incorrect, the answer should be different.",
    "Can you give me another approach?",
]

for text in texts:
    inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)
    with torch.no_grad():
        outputs = model(**inputs)
    
    probs = torch.softmax(outputs.logits, dim=-1)
    pred = outputs.logits.argmax(-1).item()
    conf = probs[0][pred].item()
    
    print(f"{labels[pred]:20} ({conf:.1%}) | {text}")

Output:

SAT                  (81.2%) | Thank you, that's exactly what I needed!
NEED_CLARIFICATION   (100.0%) | I don't understand, can you explain more?
WRONG_ANSWER         (100.0%) | That's incorrect, the answer should be different.
WANT_DIFFERENT       (100.0%) | Can you give me another approach?

Batch Inference

python
# Efficient batch processing
texts = ["Your text 1", "Your text 2", "Your text 3"]
inputs = tokenizer(texts, return_tensors="pt", padding=True, truncation=True, max_length=512)

with torch.no_grad():
    outputs = model(**inputs)

predictions = outputs.logits.argmax(-1).tolist()
feedback_types = [labels[p] for p in predictions]

Training Details

This model was fine-tuned using LoRA (Low-Rank Adaptation) with the following configuration:

ParameterValue
Base Modelvllm-sr/mmbert-32k-yarn
LoRA Rank64
LoRA Alpha128
Learning Rate2e-5
Batch Size16
Epochs10 (early stopped at ~5.4)
Precisionbf16

Training Data

Hardware

  • —GPU: AMD Instinct MI300X
  • —Training Time: ~10 minutes

Model Architecture

  • —Architecture: ModernBERT (Sequence Classification)
  • —Parameters: ~321M (base) with merged LoRA weights
  • —Max Context: 32,768 tokens (YaRN-scaled RoPE)
  • —Hidden Size: 768
  • —Layers: 22
  • —Attention Heads: 12

Multilingual Support

Supports 1800+ languages via Glot500 tokenizer. Best performance on:

  • —English (primary)
  • —Chinese
  • —French
  • —Spanish

Use Cases

  • —Conversational AI: Detect user satisfaction in chatbots
  • —Customer Support: Route conversations based on feedback type
  • —Quality Monitoring: Track user satisfaction trends
  • —Dialog Systems: Trigger clarification or correction flows

Comparison: LoRA vs Merged

VersionSizeRequires PEFTUse Case
LoRA~54MBYesFine-tuning, research
Merged (this)~615MBNoProduction, inference

Citation

bibtex
@misc{mmbert32k-feedback-detector,
  title={mmBERT-32K Feedback Detector},
  author={LLM Semantic Router Team},
  year={2026},
  publisher={Hugging Face},
  url={https://huggingface.co/vllm-sr/mmbert32k-feedback-detector-merged}
}

License

Apache 2.0