codechrl/bert-base-cybersecurity
014
1---2language:3- en4- id5tags:6- bert7- text-classification8- token-classification9- cybersecurity10- fill-mask11- named-entity-recognition12base_model: google-bert/bert-base-cased13library_name: transformers14---15 16# bert-base-cybersecurity17 18## 1. Model Details19 20**Model description** 21"bert-base-cybersecurity" is a transformer model adapted for cybersecurity text classification tasks (e.g., threat detection, incident reports, malicious vs benign content).22 23- Model type: fine-tuned lightweight BERT variant 24- Languages: English & Indonesia25- Finetuned from: `bert-base-cased`26- Status: **Early version** — trained on **0.00%** of planned data.27 28**Model sources** 29- Base model: [google-bert/bert-base-cased](https://huggingface.co/google-bert/bert-base-cased)30- Data: Cybersecurity Data31 32## 2. Uses33 34### Direct use 35You can use this model to classify cybersecurity-related text — for example, whether a given message, report or log entry indicates malicious intent, abnormal behaviour, or threat presence.36 37### Downstream use 38- Embedding extraction for clustering or anomaly detection in security logs. 39- As part of a pipeline for phishing detection, malicious email filtering, incident triage. 40- As a feature extractor feeding a downstream system (e.g., alert-generation, SOC dashboard).41 42### Out-of-scope use 43- Not meant for high-stakes automated blocking decisions without human review. 44- Not optimized for languages other than English and Indonesian. 45- Not tested for non-cybersecurity domains or out-of-distribution data.46 47## 3. Bias, Risks, and Limitations 48 49Because the model is based on a small subset (0.00%) of planned data, performance is preliminary and may degrade on unseen or specialized domains (industrial control, IoT logs, foreign language).50 51- Inherits any biases present in the base model (`google-bert/bert-base-cased`) and in the fine-tuning data — e.g., over-representation of certain threat types, vendor or tooling-specific vocabulary.52- Should not be used as sole authority for incident decisions; only as an aid to human analysts.53 54## 4. How to Get Started with the Model 55 56```python57from transformers import AutoTokenizer, AutoModelForSequenceClassification58 59tokenizer = AutoTokenizer.from_pretrained("codechrl/bert-base-cybersecurity")60model = AutoModelForSequenceClassification.from_pretrained("codechrl/bert-base-cybersecurity")61 62inputs = tokenizer("The server logged an unusual outbound connection to 123.123.123.123", 63 return_tensors="pt", truncation=True, padding=True)64outputs = model(**inputs)65logits = outputs.logits66predicted_class = logits.argmax(dim=-1).item()67```68 69## 5. Training Details70 71- **Trained records**: 1 / 237,628 (0.00%)72- **Learning rate**: 5e-0573- **Epochs**: 374- **Batch size**: 175- **Max sequence length**: 512