Team Ai
Modelpublic

ssingh20062000/security-copilot-qwen2.5-3b

sourceHugging Faceotherupdated 4d agoView on Hugging Face
0likes408downloads
Model Card

Security Co-pilot (security-copilot-qwen2.5-3b)

A 3B-parameter model fine-tuned to triage suspicious messages. Give it an email, URL, SMS message, or call transcript and it answers with a verdict: PHISHING, SCAM, SPAM, or SAFE. It follows the verdict with a one- or two-sentence reason that cites the evidence.

It is small enough to run on a laptop once quantized, so messages can be checked without sending them to an external API.

Results

Validation set, 400 held-out examples:

MetricValue
Accuracy95.5%
Macro F195.4%
Unparseable answers0
Input typeAccuracy
conversation100.0%
email97.2%
sms89.7%
url91.2%
LabelPrecisionRecallF1Support
PHISHING96.6%91.4%93.9%93
SCAM98.9%97.9%98.4%96
SPAM96.2%92.6%94.3%54
SAFE92.7%97.5%95.0%157

These are validation results. Evaluation on the separate test set, and a comparison with the untuned base model, are reported separately.

How to use

python
import json, torch
from huggingface_hub import hf_hub_download
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "ssingh20062000/security-copilot-qwen2.5-3b"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, dtype=torch.bfloat16, device_map="auto")
cfg = json.load(open(hf_hub_download(repo, "prompt_config.json")))

message = "Your bank KYC has expired. Update now at kyc-update-verify.in/login or your account will be blocked."
messages = [
    {"role": "system", "content": cfg["system_prompt"]},
    {"role": "user", "content": cfg["user_message_template"].format(input_type="sms", input=message)},
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt", add_special_tokens=False).to(model.device)
output = model.generate(**inputs, max_new_tokens=128, do_sample=False)
print(tokenizer.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))   # "Verdict: ..."  (then "Reason: ...")

Use the exact prompt in prompt_config.json, and format inputs the way the training data was formatted: emails start with From: / Subject: lines followed by a blank line and the body, and URLs are given without http:// or https://.

The LoRA adapter on its own is in the adapter/ folder, for use with peft on top of Qwen/Qwen2.5-3B-Instruct.

Training data

About 18,500 examples after cleaning, balanced across four input types:

  • —Emails: Nazario (phishing), Nigerian Fraud (advance-fee scams), CEAS 2008, SpamAssassin, Enron-Spam, and Ling-Spam (spam and legitimate mail)
  • —URLs: a public phishing-site URL dataset
  • —SMS and call transcripts: SMS spam plus translated scam-call and everyday conversation transcripts

Cleaning went beyond de-duplication. Collection artefacts that would let a model identify the source dataset instead of the threat were removed or randomised: collector mailbox addresses, anonymisation tokens, template names, currency words that differed by class, transcription marks, MIME boilerplate, and a length difference between classes. Near-duplicates were removed, and whole templates and URL domains were kept within a single split to prevent leakage.

Training procedure

SettingValue
MethodQLoRA (4-bit NF4 base, LoRA adapters on all linear layers)
LoRA r / alpha / dropout16 / 32 / 0.05
Learning rate0.0002 (cosine schedule)
Epochs1
Effective batch size16
Max sequence length1536 tokens
Training examples14,806
Hardware1x NVIDIA T4 (Kaggle), fp16

Loss was computed only on the answer tokens, not the prompt.

Limitations

  • —English only. Non-English messages were removed from the training data.
  • —SMS scams are labelled SPAM. The SMS source data doesn't separate scams from other spam, so scam texts (fake prizes, KYC updates) come out as SPAM rather than SCAM.
  • —Dated and skewed data. Most emails are from 2002 to 2008, apart from the phishing set (2015 to 2022). Legitimate emails come mostly from technical mailing lists, so genuine bank alerts, OTPs, and delivery notices may be flagged incorrectly.
  • —No live signals. It sees only the text: it cannot check whether a domain is newly registered, look at attachments, or verify a sender.
  • —Adversarial inputs. Messages written to fool classifiers can evade it.

Use it as a first-pass assistant alongside other controls, not as the only line of defence.

Licence

This model is a derivative of Qwen2.5-3B-Instruct and is released under the Qwen Research License, which restricts commercial use. Read the licence before using it.