Team Ai
Modelpublic

vllm-sr/mmbert32k-feedback-detector-lora

sourceHugging Faceapache-2.0updated 9d agoView on Hugging Face
1likes48downloads
Model Card

mmBERT-32K Feedback Detector (LoRA)

A 4-class user feedback classifier fine-tuned from mmbert-32k-yarn using LoRA (Low-Rank Adaptation).

Model Description

This model classifies user messages into 4 feedback categories to help conversational AI systems understand user satisfaction and respond appropriately:

LabelIDDescription
SAT0User is satisfied with the response
NEED_CLARIFICATION1User needs more explanation or details
WRONG_ANSWER2User indicates the response was incorrect
WANT_DIFFERENT3User wants an alternative approach/answer

Performance

Validation Results (2,985 samples):

MetricValue
Accuracy98.83%
F1 (macro)98.24%
F1 (weighted)98.83%

Per-Class Performance:

ClassPrecisionRecallF1-ScoreSupport
SAT1.00001.00001.00001,491
NEED_CLARIFICATION0.99800.99800.9980498
WRONG_ANSWER0.96040.97390.9671498
WANT_DIFFERENT0.97150.95780.9646498

Usage

With PEFT (Recommended)

python
from transformers import AutoModelForSequenceClassification, AutoTokenizer
from peft import PeftModel

# Load base model
base_model = AutoModelForSequenceClassification.from_pretrained(
    "vllm-sr/mmbert-32k-yarn",
    num_labels=4
)
tokenizer = AutoTokenizer.from_pretrained("vllm-sr/mmbert-32k-yarn")

# Load LoRA adapter
model = PeftModel.from_pretrained(base_model, "vllm-sr/mmbert32k-feedback-detector-lora")
model.eval()

# Inference
labels = ["SAT", "NEED_CLARIFICATION", "WRONG_ANSWER", "WANT_DIFFERENT"]
text = "I don't understand your explanation, can you clarify?"

inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=512)
outputs = model(**inputs)
prediction = outputs.logits.argmax(-1).item()

print(f"Feedback: {labels[prediction]}")  # Output: NEED_CLARIFICATION

Using Merged Model (No PEFT required)

For easier deployment, use the merged version:

python
from transformers import AutoModelForSequenceClassification, AutoTokenizer

model = AutoModelForSequenceClassification.from_pretrained(
    "vllm-sr/mmbert32k-feedback-detector-merged"
)
tokenizer = AutoTokenizer.from_pretrained(
    "vllm-sr/mmbert32k-feedback-detector-merged"
)

Training Details

Hyperparameters

ParameterValue
Base Modelvllm-sr/mmbert-32k-yarn
LoRA Rank64
LoRA Alpha128
LoRA Dropout0.1
Target Modulesattn.Wqkv, attn.Wo, mlp.Wi, mlp.Wo
Learning Rate2e-5
Batch Size16
Epochs10 (early stopping at ~5.4)
Warmup Ratio0.1
Weight Decay0.01
Precisionbf16
OptimizerAdamW

Training Data

Trained on vllm-sr/feedback-detector-dataset:

  • —Training samples: 17,896 (balanced across 4 classes)
  • —Validation samples: 2,985

Hardware

  • —GPU: AMD Instinct MI300X (192GB HBM3)
  • —Training Time: ~10 minutes
  • —Framework: PyTorch 2.x with ROCm

Multilingual Support

The model inherits multilingual capabilities from mmbert-32k-yarn (Glot500 tokenizer supporting 1800+ languages). Best performance on:

  • —English (primary)
  • —Chinese (Simplified/Traditional)
  • —French
  • —Spanish

Limitations

  • —The SAT class has the strongest performance; some edge cases between WRONG_ANSWER and WANT_DIFFERENT may be ambiguous
  • —Phrases like "That's perfect, no more questions" may sometimes be misclassified
  • —Best suited for conversational AI feedback detection, not general sentiment analysis

Citation

bibtex
@misc{mmbert32k-feedback-detector,
  title={mmBERT-32K Feedback Detector},
  author={LLM Semantic Router Team},
  year={2026},
  publisher={Hugging Face},
  url={https://huggingface.co/vllm-sr/mmbert32k-feedback-detector-lora}
}

License

Apache 2.0

Framework Versions

  • —PEFT: 0.18.1
  • —Transformers: 4.48+
  • —PyTorch: 2.6+