Team Ai
Modelpublic

AbhishekG711/Qwen3-4B-Instruct-2507-Insight-Extractor

sourceHugging Faceapache-2.0updated 6d agoView on Hugging Face
0likes502downloads
Model Card

Model Summary

Qwen3-4B-Instruct-2507-Insight-Extractor is a LoRA fine-tune of [`Qwen/Qwen3-4B-Instruct-2507`](https://huggingface.co/Qwen/Qwen3-4B-Instruct-2507). It converts a raw customer-support ticket into a single structured JSON object across 9 fields, for automated ticket classification and downstream entity-extraction workflows.

Base modelQwen/Qwen3-4B-Instruct-2507
Fine-tuning methodLoRA Adapters via Unsloth + TRL SFTTrainer
TaskStructured JSON extraction from support-ticket text
Training data`AbhishekG711/processed_support_tickets` (train split)
Target repoAbhishekG711/Qwen3-4B-Instruct-2507-Insight-Extractor

Base Model Architecture (Qwen3-4B-Instruct-2507)

TypeDense, causal decoder-only Transformer
Parameters4.0B total (~3.6B non-embedding)
Layers36
AttentionGrouped-Query Attention — 32 query heads / 8 KV heads
Native context length262,144 tokens
Reasoning modesNon-thinking only. Unlike the 0.6B and 1.7B checkpoints, the "-Instruct-2507" line was released specifically without a thinking mode — it never emits <think>...</think> blocks.
Base model licenseApache 2.0

Intended Use

Given a raw support ticket, produce one JSON object for automated triage, routing, sentiment/urgency dashboards, or downstream alerting — not a conversational assistant.

Fine-Tuning Training Details

HyperparameterValue
QuantizationNone (load_in_4bit=False)
LoRA rank (r)32
LoRA alpha64 (= 2 × r)
LoRA dropout0.0
LoRA target modulesq_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
LoRA biasnone
Max sequence length4096 (intentional, confirmed — see note above)
Per-device train batch size4
Target global batch size32
Gradient accumulation steps8 (= 32 // 4)
Epochs1
Learning rate2e-4
LR schedulerlinear
Warmup steps10
Weight decay0.01
Optimizeradamw_8bit
Precisionbf16 if supported, else fp16 (auto-detected)
Completion-only lossTrue
Sequence packingFalse
Eval strategyevery 25 steps
Save strategyevery 50 steps, save_total_limit=3
Load best model at endTrue (metric_for_best_model="eval_loss", greater_is_better=False)
Gradient checkpointing"unsloth" mode
Seed3407
Experiment trackerreport_to="none"

Evaluation

MetricValue
Validation loss (eval_loss, best checkpoint)0.222669

How to Prompt This Model

Same system/user structure as the rest of the Insight-Extractor family:

System prompt:

# ROLE
You are a deterministic extraction engine. You convert raw support emails into exactly one structured JSON object. You never converse, explain, or output anything except that JSON object.

User message template:

Analyze the following customer support email and extract structured insights.

<raw ticket text>

Expected output: one JSON object with exactly these 9 keys, in this order: is_actionable, summary, sentiment, category, intent, aspect, urgency, reported_cause, entities.

Inference example

python
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "AbhishekG711/Qwen3-4B-Instruct-2507-Insight-Extractor"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")

SYSTEM_PROMPT = """
# ROLE
You are a deterministic extraction engine. You convert raw support emails into exactly one structured JSON object. You never converse, explain, or output anything except that JSON object.
"""
USER_INSTRUCTION = "Analyze the following customer support email and extract structured insights.\n\n"

ticket_text = "My order #4471 arrived damaged and support hasn't replied in 3 days."

messages = [
    {"role": "system", "content": SYSTEM_PROMPT},
    {"role": "user", "content": USER_INSTRUCTION + ticket_text},
]

# See the architecture note above: this base model is non-thinking-only, so do not assume
# enable_thinking behaves the same way here as it does for the 0.6B/1.7B checkpoints.
inputs = tokenizer.apply_chat_template(
    messages,
    add_generation_prompt=True,
    return_tensors="pt",
).to(model.device)

output = model.generate(inputs, max_new_tokens=512)
print(tokenizer.decode(output[0][inputs.shape[-1]:], skip_special_tokens=True))