attack-vector/SecureModernBERT-NER
21.6k
1---2language: en3library_name: transformers4pipeline_tag: token-classification5tags:6- ner7- token-classification8- cybersecurity9- threat-intelligence10- secureBert11license: mit12metrics:13- accuracy14base_model:15- answerdotai/ModernBERT-large16---17 18# Model Overview19 20**SecureModernBERT-NER** represents a new generation of cybersecurity-focused language models — combining the **state-of-the-art architecture of ModernBERT** with one of the **largest and most diverse CTI-labelled NER corpora ever built**. 21 22Unlike conventional NER systems, SecureModernBERT-NER recognises **22 finely-grained, security-specific entity types**, covering the full spectrum of cyber-threat intelligence — from `THREAT-ACTOR` and `MALWARE` to `CVE`, `IPV4`, `DOMAIN`, and `REGISTRY-KEYS`. 23 24Trained on more than **half a million manually curated spans** sourced from real-world threat reports, vulnerability advisories, and incident analyses, it achieves an exceptional balance of **accuracy, generalisation, and contextual depth**. 25 26This model is designed to **parse complex security narratives with human-level precision**, extracting both contextual metadata (e.g., `ORG`, `PRODUCT`, `PLATFORM`) and highly technical indicators (e.g., `HASHES`, `URLS`, `NETWORK ADDRESSES`) — all within a single unified framework. 27 28SecureModernBERT-NER sets a new standard for **automated CTI entity recognition**, enabling the next wave of **threat-intelligence automation, enrichment, and analytics**. 29 30## Quick Start31 32```python33from transformers import pipeline34 35model_id = "attack-vector/SecureModernBERT-NER"36 37pipe = pipeline(38 task="token-classification",39 model=model_id,40 tokenizer=model_id,41 aggregation_strategy="first",42)43 44text = "TrickBot connects to hxxp://185.222.202.55 to exfiltrate data from Windows hosts."45predictions = pipe(text)46for pred in predictions:47 print(pred)48```49 50Sample output:51 52```53{'entity_group': 'MALWARE', 'score': np.float32(0.9615546), 'word': 'TrickBot', 'start': 0, 'end': 8}54{'entity_group': 'URL', 'score': np.float32(0.9905957), 'word': ' hxxp://185.222.202.55', 'start': 20, 'end': 42}55{'entity_group': 'PLATFORM', 'score': np.float32(0.92317337), 'word': ' Windows', 'start': 66, 'end': 74}56```57 58## Intended Use & Limitations59 60- **Use cases:** automated tagging of CTI reports, IOC extraction pipelines, knowledge-base enrichment, security-focused RAG systems.61- **Languages:** English (model was trained and evaluated on English sources only).62- **Input format:** free-form prose or long-form CTI articles; maximum sequence length 128 tokens during training.63- **Limitations:** noisy or ambiguous extractions may occur, especially with rare entity types (`IPV6`, `EMAIL`) and obfuscated strings. The model does not normalise entities (e.g., deobfuscating `hxxp`) nor validate indicator authenticity. Always pair with downstream validation and human review.64 65## Training Data66 67- **Size:** 502,726 labelled text spans before filtering; 22 distinct entity classes in BIO format.68- **Label distribution (spans):** `ORG` (approx. 198k), `PRODUCT` (approx. 79k), `MALWARE` (approx. 67k), `PLATFORM` (approx. 57k), `THREAT-ACTOR` (approx. 49k), `SERVICE` (approx. 46k), `CVE` (approx. 41k), `LOC` (approx. 38k), `SECTOR` (approx. 34k), `TOOL` (approx. 29k), plus indicator types such as `URL`, `IPV4`, `SHA256`, `MD5`, and `REGISTRY-KEYS`.69- **Pre-processing:** JSONL articles were tokenised and converted to BIO tags; spans in conflict were resolved manually and via automated heuristics before upload.70 71## Label Mapping72 73| Label | Description | Example mention |74|-------|-------------|-----------------|75| URL | Web address or obfuscated link used in campaigns. | `hxxp://185.222.202.55` |76| ORG | Organisations such as companies, CERTs, or research groups. | `Microsoft Threat Intelligence` |77| SERVICE | Online or cloud services referenced in attacks. | `Google Ads` |78| SECTOR | Industry sectors or verticals targeted. | `critical infrastructure` |79| FILEPATH | File system paths observed in malware samples. | `C:\Windows\System32\svchost.exe` |80| DOMAIN | Fully qualified domains or subdomains. | `malicious-domain[.]com` |81| PLATFORM | Operating systems or computing platforms. | `Windows Server` |82| THREAT-ACTOR | Named adversary groups or aliases. | `LockBit` |83| PRODUCT | Commercial or open-source software products. | `VMware ESXi` |84| MALWARE | Malware families, strains, or toolkits. | `TrickBot` |85| LOC | Countries, cities, or regions. | `United States` |86| CVE | CVE identifiers for vulnerabilities. | `CVE-2023-23397` |87| TOOL | Legitimate or dual-use tools leveraged in incidents. | `Cobalt Strike` |88| IPV4 | IPv4 addresses. | `185.222.202.55` |89| MITRE-TACTIC | MITRE ATT&CK tactic categories. | `Credential Access` |90| MD5 | MD5 cryptographic hashes. | `d41d8cd98f00b204e9800998ecf8427e` |91| CAMPAIGN | Named operations or campaigns. | `Operation Cronos` |92| SHA1 | SHA-1 hashes. | `da39a3ee5e6b4b0d3255bfef95601890afd80709` |93| SHA256 | SHA-256 hashes. | `9e107d9d372bb6826bd81d3542a419d6...` |94| EMAIL | Email addresses. | `alerts@example.com` |95| IPV6 | IPv6 addresses. | `2001:0db8:85a3:0000:0000:8a2e:0370:7334` |96| REGISTRY-KEYS | Windows registry keys or paths. | `HKLM\Software\Microsoft\Windows\CurrentVersion\Run` |97 98## Training Procedure99 100- **Base model:** [`answerdotai/ModernBERT-large`](https://huggingface.co/answerdotai/ModernBERT-large).101- **Hardware:** single Nvidia L40S instance (8 vCPU / 62 GB RAM / 48 GB VRAM).102- **Optimisation setup:** mixed precision `fp16`, optimiser `adamw_torch`, cosine learning-rate scheduler, gradient accumulation `1`.103- **Key hyperparameters:** learning rate `5e-5`, batch size `128`, epochs `5`, maximum sequence length `128`.104 105| Parameter | Value |106|-----------|-------|107| Mixed precision | `fp16` |108| Batch size | `128` |109| Learning rate | `5e-5` |110| Optimiser | `adamw_torch` |111| Scheduler | `cosine` |112| Epochs | `5` |113| Gradient accumulation | `1` |114| Max sequence length | `128` |115 116## Evaluation117 118AutoTrain reports the following micro-averaged metrics on its validation split (seqeval entity scoring):119 120| Metric | Score |121|------------|--------|122| Precision | 0.8468 |123| Recall | 0.8484 |124| F1 | 0.8476 |125| Accuracy | 0.9589 |126 127An independent re-evaluation against a consolidated CTI set (same taxonomy as this model) produced the label-level accuracy breakdown below. These scores are macro-averaged across labels and therefore are not numerically comparable to the micro metrics above, but they provide insight into class balance and span quality.128 129| Label | Used | Accuracy |130|-------|------|----------|131| CAMPAIGN | 1,817 | 0.7980 |132| CVE | 28,293 | 0.9995 |133| DOMAIN | 12,182 | 0.8878 |134| EMAIL | 731 | 0.8495 |135| FILEPATH | 13,889 | 0.7957 |136| IPV4 | 1,164 | 0.9631 |137| IPV6 | 563 | 0.7425 |138| LOC | 7,915 | 0.9557 |139| MALWARE | 10,405 | 0.9087 |140| MD5 | 389 | 0.9100 |141| MITRE-TACTIC | 2,181 | 0.7093 |142| ORG | 36,324 | 0.9301 |143| PLATFORM | 8,036 | 0.8977 |144| PRODUCT | 18,720 | 0.8432 |145| REGISTRY-KEYS | 1,589 | 0.8490 |146| SECTOR | 6,453 | 0.8309 |147| SERVICE | 8,533 | 0.8179 |148| SHA1 | 222 | 0.9189 |149| SHA256 | 2,146 | 0.9874 |150| THREAT-ACTOR | 9,532 | 0.9418 |151| TOOL | 4,874 | 0.7895 |152| URL | 7,470 | 0.9801 |153 154- **Macro accuracy:** 0.8776155 156Because micro vs macro averaging and dataset composition differ, expect numerical gaps between the two evaluations even though both describe the same checkpoint.157 158These metrics were computed with the `seqeval` micro-average at the entity level.159 160## External Benchmarks161 162The following tables report detailed results on a shared CTI validation set. **Do not compare the per-label values across models directly:** each checkpoint uses a different taxonomy or remapping strategy, so accuracy percentages can be misleading when labels are aligned or collapsed differently. Use the per-model tables to understand performance within a single schema, and interpret macro-accuracy scores with caution.163 164 165### CyberPeace-Institute/SecureBERT-NER166 167| Label | Used | Accuracy |168|-------|------|----------|169| ACT | 3,945 | 0.1706 |170| APT | 9,518 | 0.5331 |171| DOM | 10,694 | 0.0196 |172| EMAIL | 731 | 0.0000 |173| FILE | 31,864 | 0.0747 |174| IP | 1,251 | 0.0088 |175| LOC | 7,895 | 0.8711 |176| MAL | 10,341 | 0.6076 |177| MD5 | 354 | 0.8672 |178| O | 16,275 | 0.4700 |179| OS | 7,974 | 0.6598 |180| SECTEAM | 36,083 | 0.3509 |181| SHA1 | 191 | 0.0209 |182| SHA2 | 1,647 | 0.9709 |183| TOOL | 4,816 | 0.4043 |184| URL | 6,997 | 0.0795 |185| VULID | 27,586 | 0.3849 |186 187- **Macro accuracy:** 0.3820188 189### PranavaKailash/CyNER-2.0-DeBERTa-v3-base190 191| Label | Used | Accuracy |192|-------|------|----------|193| Indicator | 35,936 | 0.7878 |194| Location | 7,895 | 0.0113 |195| Malware | 12,125 | 0.7800 |196| O | 2,896 | 0.7652 |197| Organization | 42,537 | 0.6556 |198| System | 35,063 | 0.7259 |199| TOOL | 4,820 | 0.0000 |200| Threat Group | 9,522 | 0.0000 |201| Vulnerability | 27,673 | 0.1876 |202 203- **Macro accuracy:** 0.4348204 205### cisco-ai/SecureBERT2.0-NER206 207| Label | Used | Accuracy |208|-------|------|----------|209| Indicator | 35,789 | 0.8854 |210| Malware | 16,926 | 0.6204 |211| O | 10,786 | 0.6813 |212| Organization | 51,993 | 0.5579 |213| System | 34,955 | 0.6600 |214| Vulnerability | 27,525 | 0.2552 |215 216- **Macro accuracy:** 0.6100217 218 219## Responsible Use220 221- Confirm entity detections before acting on indicators (e.g., automated blocking).222- Combine with enrichment and scoring systems to filter false positives.223- Monitor for drift if applying to new domains (e.g., non-English sources, informal channels).224- Respect licensing and confidentiality of any proprietary CTI sources used for inference.225 226 227## Support & Connect228 229* ❤️ **Like the repo** if you found it useful230* ☕ **Support me:** Say thanks by buying me a coffee! [https://buymeacoffee.com/juanmcristobal](https://buymeacoffee.com/juanmcristobal)231* 💼 **Open to work:** [https://www.linkedin.com/in/jmcristobal/](https://www.linkedin.com/in/jmcristobal/)232 233If you use SecureModernBERT-NER in a project, feel free to share it in the Discussions/Issues — I love seeing real-world use cases.234 235## Citation236 237If you find this model useful, please cite the repository and the base model:238 239```240@software{securemodernbert_ner_2025,241 author = {Juan Manuel Cristóbal Moreno},242 title = {SecureModernBERT-NER: Cyber Threat Intelligence Named Entity Recogniser},243 year = {2025},244 publisher = {Hugging Face},245 url = {https://huggingface.co/attack-vector/SecureModernBERT-NER}246}247```248 249## Contact250 251Questions or feedback? Open an issue on the Hugging Face model repository or reach out at [`@juanmcristobal`](https://huggingface.co/juanmcristobal).