Team Ai
Modelpublic

codechrl/bert-base-cybersecurity

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes14downloads
README.md75 linesDownload Raw Back to root
1---2language:3- en4- id5tags:6- bert7- text-classification8- token-classification9- cybersecurity10- fill-mask11- named-entity-recognition12base_model: google-bert/bert-base-cased13library_name: transformers14---15 16# bert-base-cybersecurity17 18## 1. Model Details19 20**Model description**  21"bert-base-cybersecurity" is a transformer model adapted for cybersecurity text classification tasks (e.g., threat detection, incident reports, malicious vs benign content).22 23- Model type: fine-tuned lightweight BERT variant  24- Languages: English & Indonesia25- Finetuned from: `bert-base-cased`26- Status: **Early version** — trained on **0.00%** of planned data.27 28**Model sources**  29- Base model: [google-bert/bert-base-cased](https://huggingface.co/google-bert/bert-base-cased)30- Data: Cybersecurity Data31 32## 2. Uses33 34### Direct use  35You can use this model to classify cybersecurity-related text — for example, whether a given message, report or log entry indicates malicious intent, abnormal behaviour, or threat presence.36 37### Downstream use  38- Embedding extraction for clustering or anomaly detection in security logs.  39- As part of a pipeline for phishing detection, malicious email filtering, incident triage.  40- As a feature extractor feeding a downstream system (e.g., alert-generation, SOC dashboard).41 42### Out-of-scope use  43- Not meant for high-stakes automated blocking decisions without human review.  44- Not optimized for languages other than English and Indonesian.  45- Not tested for non-cybersecurity domains or out-of-distribution data.46 47## 3. Bias, Risks, and Limitations 48 49Because the model is based on a small subset (0.00%) of planned data, performance is preliminary and may degrade on unseen or specialized domains (industrial control, IoT logs, foreign language).50 51- Inherits any biases present in the base model (`google-bert/bert-base-cased`) and in the fine-tuning data — e.g., over-representation of certain threat types, vendor or tooling-specific vocabulary.52- Should not be used as sole authority for incident decisions; only as an aid to human analysts.53 54## 4. How to Get Started with the Model  55 56```python57from transformers import AutoTokenizer, AutoModelForSequenceClassification58 59tokenizer = AutoTokenizer.from_pretrained("codechrl/bert-base-cybersecurity")60model = AutoModelForSequenceClassification.from_pretrained("codechrl/bert-base-cybersecurity")61 62inputs = tokenizer("The server logged an unusual outbound connection to 123.123.123.123", 63                   return_tensors="pt", truncation=True, padding=True)64outputs = model(**inputs)65logits = outputs.logits66predicted_class = logits.argmax(dim=-1).item()67```68 69## 5. Training Details70 71- **Trained records**: 1 / 237,628 (0.00%)72- **Learning rate**: 5e-0573- **Epochs**: 374- **Batch size**: 175- **Max sequence length**: 512