Team Ai
Modelpublic

AventIQ-AI/Securebert-website-phishing-prediction

sourceHugging Faceupdated 2y agoView on Hugging Face
5likes9downloads
Model Card

๐Ÿ”’ SecureBERT Phishing Detection Model

This repository hosts a fine-tuned SecureBERT-based model optimized for phishing URL detection using a cybersecurity dataset. The model classifies URLs as either phishing (malicious) or safe (benign).


๐Ÿ“š Model Details

  • โ€”Model Architecture: SecureBERT (Based on BERT)
  • โ€”Task: Binary Classification (Phishing vs. Safe)
  • โ€”Dataset: shashwatwork/web-page-phishing-detection-dataset (11,431 URLs, 88 features)
  • โ€”Framework: PyTorch & Hugging Face Transformers
  • โ€”Input Data: URL strings & extracted numerical features
  • โ€”Number of Classes: 2 (Phishing, Safe)
  • โ€”Quantization: FP16 (for efficiency)

๐Ÿš€ Usage

Installation

bash
pip install torch transformers scikit-learn pandas

Loading the Model

python
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification

# Load the fine-tuned model and tokenizer
model_path = "./fine_tuned_SecureBERT"
tokenizer = AutoTokenizer.from_pretrained(model_path)
model = AutoModelForSequenceClassification.from_pretrained(model_path)
model.eval()  # Set model to evaluation mode

print("โœ… SecureBERT model loaded successfully and ready for inference!")

๐Ÿ” Perform Phishing Detection

python
def predict_url(url):
    # Tokenize input
    encoding = tokenizer(url, truncation=True, padding=True, max_length=512, return_tensors="pt")
    
    # Perform inference
    with torch.no_grad():
        output = model(**encoding)
    
    # Get predicted class
    predicted_class = torch.argmax(output.logits, dim=1).item()
    
    # Map label
    label = "Phishing" if predicted_class == 1 else "Safe"
    return label

# Example usage
custom_url = "http://example.com/free-gift"
prediction = predict_url(custom_url)
print(f"Predicted label: {prediction}")

๐Ÿ“Š Evaluation Results

After fine-tuning, the model was evaluated on a test set, achieving the following performance:

**Metric****Score**
Accuracy97.2%
Precision96.8%
Recall97.5%
F1-Score97.1%
Inference SpeedFast (Optimized with FP16)

๐Ÿ› ๏ธ Fine-Tuning Details

Dataset

The model was trained on a shashwatwork/web-page-phishing-detection-dataset consisting of 11,431 URLs labeled as either phishing or safe. Features include URL characteristics, domain properties, and additional metadata.

Training Configuration

  • โ€”Number of epochs: 5
  • โ€”Batch size: 16
  • โ€”Optimizer: AdamW
  • โ€”Learning rate: 2e-5
  • โ€”Loss Function: Cross-Entropy
  • โ€”Evaluation Strategy: Validation at each epoch

Quantization

The model was quantized using FP16 precision, reducing latency and memory usage while maintaining high accuracy.


โš ๏ธ Limitations

  • โ€”Evasion Techniques: Attackers constantly evolve phishing techniques, which may reduce model effectiveness.
  • โ€”Dataset Bias: The model was trained on a specific dataset; new phishing tactics may require retraining.
  • โ€”False Positives: Some legitimate but unusual URLs might be classified as phishing.

โœ… Use this fine-tuned SecureBERT model for accurate and efficient phishing detection! ๐Ÿ”’๐Ÿš€