ssingh20062000/security-copilot-qwen2.5-3b
Security Co-pilot (security-copilot-qwen2.5-3b)
A 3B-parameter model fine-tuned to triage suspicious messages. Give it an email, URL, SMS message, or call transcript and it answers with a verdict: PHISHING, SCAM, SPAM, or SAFE. It follows the verdict with a one- or two-sentence reason that cites the evidence.
It is small enough to run on a laptop once quantized, so messages can be checked without sending them to an external API.
Results
Validation set, 400 held-out examples:
These are validation results. Evaluation on the separate test set, and a comparison with the untuned base model, are reported separately.
How to use
import json, torch
from huggingface_hub import hf_hub_download
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "ssingh20062000/security-copilot-qwen2.5-3b"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, dtype=torch.bfloat16, device_map="auto")
cfg = json.load(open(hf_hub_download(repo, "prompt_config.json")))
message = "Your bank KYC has expired. Update now at kyc-update-verify.in/login or your account will be blocked."
messages = [
{"role": "system", "content": cfg["system_prompt"]},
{"role": "user", "content": cfg["user_message_template"].format(input_type="sms", input=message)},
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt", add_special_tokens=False).to(model.device)
output = model.generate(**inputs, max_new_tokens=128, do_sample=False)
print(tokenizer.decode(output[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True)) # "Verdict: ..." (then "Reason: ...")Use the exact prompt in prompt_config.json, and format inputs the way the training data was formatted: emails start with From: / Subject: lines followed by a blank line and the body, and URLs are given without http:// or https://.
The LoRA adapter on its own is in the adapter/ folder, for use with peft on top of Qwen/Qwen2.5-3B-Instruct.
Training data
About 18,500 examples after cleaning, balanced across four input types:
- Emails: Nazario (phishing), Nigerian Fraud (advance-fee scams), CEAS 2008, SpamAssassin, Enron-Spam, and Ling-Spam (spam and legitimate mail)
- URLs: a public phishing-site URL dataset
- SMS and call transcripts: SMS spam plus translated scam-call and everyday conversation transcripts
Cleaning went beyond de-duplication. Collection artefacts that would let a model identify the source dataset instead of the threat were removed or randomised: collector mailbox addresses, anonymisation tokens, template names, currency words that differed by class, transcription marks, MIME boilerplate, and a length difference between classes. Near-duplicates were removed, and whole templates and URL domains were kept within a single split to prevent leakage.
Training procedure
Loss was computed only on the answer tokens, not the prompt.
Limitations
- English only. Non-English messages were removed from the training data.
- SMS scams are labelled SPAM. The SMS source data doesn't separate scams from other spam, so scam texts (fake prizes, KYC updates) come out as
SPAMrather thanSCAM. - Dated and skewed data. Most emails are from 2002 to 2008, apart from the phishing set (2015 to 2022). Legitimate emails come mostly from technical mailing lists, so genuine bank alerts, OTPs, and delivery notices may be flagged incorrectly.
- No live signals. It sees only the text: it cannot check whether a domain is newly registered, look at attachments, or verify a sender.
- Adversarial inputs. Messages written to fool classifiers can evade it.
Use it as a first-pass assistant alongside other controls, not as the only line of defence.
Licence
This model is a derivative of Qwen2.5-3B-Instruct and is released under the Qwen Research License, which restricts commercial use. Read the licence before using it.
