Team Ai
Modelpublic

AbhishekG711/Qwen3-0.6B-Insight-Extractor

sourceHugging Faceapache-2.0updated 1d agoView on Hugging Face
0likes967downloads
Model Card

Model Summary

Qwen3-0.6B-Insight-Extractor is a LoRA fine-tune of [`Qwen/Qwen3-0.6B`](https://huggingface.co/Qwen/Qwen3-0.6B). It converts a raw customer-support ticket into a single structured JSON object across 9 fields, for automated ticket classification and downstream entity-extraction workflows.

Base modelQwen/Qwen3-0.6B
Fine-tuning methodLoRA Adapters via Unsloth + TRL SFTTrainer
TaskStructured JSON extraction from support-ticket text
Training data`AbhishekG711/processed_support_tickets` (train split)
Target repoAbhishekG711/Qwen3-0.6B-Insight-Extractor

Base Model Architecture (Qwen3-0.6B)

TypeDense, causal decoder-only Transformer
Parameters0.6B total (~0.44B non-embedding)
Layers28
AttentionGrouped-Query Attention — 16 query heads / 8 KV heads
Native context length32,768 tokens
Reasoning modesHybrid — supports both "thinking" (<think>...</think>) and "non-thinking" modes via enable_thinking
Base model licenseApache 2.0

Intended Use

Given a raw support ticket, produce one JSON object for automated triage, routing, sentiment/urgency dashboards, or downstream alerting — not a conversational assistant.

Fine-Tuning Training Details

HyperparameterValue
QuantizationNone (load_in_4bit=False)
LoRA rank (r)32
LoRA alpha64 (= 2 × r)
LoRA dropout0.0
LoRA target modulesq_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
LoRA biasnone
Max sequence length4096
Per-device train batch size4
Target global batch size32
Gradient accumulation steps8 (= 32 // 4)
Epochs1
Learning rate2e-4
LR schedulerlinear
Warmup steps10
Weight decay0.01
Optimizeradamw_8bit
Precisionbf16 if supported, else fp16 (auto-detected via is_bf16_supported())
Completion-only lossTrue — loss computed only on completion tokens, not the prompt
Sequence packingFalse
Eval strategyevery 25 steps
Save strategyevery 50 steps, save_total_limit=3
Load best model at endTrue (metric_for_best_model="eval_loss", greater_is_better=False)
Gradient checkpointing"unsloth" mode
Seed3407
Experiment trackerreport_to="none" — console/log file only

Evaluation

MetricValue
Validation loss (eval_loss, best checkpoint)0.233829

How to Prompt This Model

The model was fine-tuned against exactly this system/user structure — matching it as closely as possible at inference time will give the best results.

System prompt:

# ROLE
You are a deterministic extraction engine. You convert raw support emails into exactly one structured JSON object. You never converse, explain, or output anything except that JSON object.

User message template:

Analyze the following customer support email and extract structured insights.

<raw ticket text>

Expected output: One JSON object with exactly these 9 keys, in this order: is_actionable, summary, sentiment, category, intent, aspect, urgency, reported_cause, entities.

Inference example

python
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "AbhishekG711/Qwen3-0.6B-Insight-Extractor"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")

SYSTEM_PROMPT = """
# ROLE
You are a deterministic extraction engine. You convert raw support emails into exactly one structured JSON object. You never converse, explain, or output anything except that JSON object.
"""
USER_INSTRUCTION = "Analyze the following customer support email and extract structured insights.\n\n"

ticket_text = "My order #4471 arrived damaged and support hasn't replied in 3 days."

messages = [
    {"role": "system", "content": SYSTEM_PROMPT},
    {"role": "user", "content": USER_INSTRUCTION + ticket_text},
]

inputs = tokenizer.apply_chat_template(
    messages,
    add_generation_prompt=True,
    enable_thinking=False,
    return_tensors="pt",
).to(model.device)

output = model.generate(inputs, max_new_tokens=512)
print(tokenizer.decode(output[0][inputs.shape[-1]:], skip_special_tokens=True))