CIRCL/vulnerability-attack-technique-biencoder
0127
1---2license: gpl-3.03base_model: FacebookAI/roberta-base4datasets:5 - CIRCL/vulnerability-attack-techniques6language:7 - en8tags:9 - security10 - vulnerability11 - mitre-attack12 - cve13 - bi-encoder14 - sentence-similarity15library_name: transformers16---17 18# vulnerability-attack-technique-biencoder19 20A label-semantics bi-encoder that suggests MITRE ATT&CK (Enterprise)21techniques for a CVE by scoring the vulnerability description against the22**official ATT&CK technique descriptions** in a shared embedding space.23Unlike the companion classification head24([`CIRCL/vulnerability-attack-technique-classification-roberta-base`](https://huggingface.co/CIRCL/vulnerability-attack-technique-classification-roberta-base)),25it can rank *any* technique that has an official description — the label26is text, not a learned output row.27 28One shared `roberta-base` encoder embeds both the CVE text29(title + description) and each technique's STIX name+description30(citation markup stripped, 256 tokens), mean-pooled and L2-normalized;31the score is a learned affine over the cosine. Trained on the curated32gold set [`CIRCL/vulnerability-attack-techniques`](https://huggingface.co/datasets/CIRCL/vulnerability-attack-techniques)33(~1,200 CVEs, CTID methodology) with per-label-weighted BCE over a3453-parent-technique vocabulary, with VulnTrain35(`vulntrain-train-attack-biencoder`).36 37## When to use which model38 39- **Classification head**: best top-5 ranking on the trained vocabulary40 (recall@5 0.667 ± 0.015 across five seeds).41- **This bi-encoder**: slightly lower recall@5 (0.643 ± 0.019) but the42 largest consistent rare-technique gain measured on this task43 (macro-F1 0.212 ± 0.011 vs 0.176 ± 0.016, +21% relative), and44 open-vocabulary ranking over all 222 active parent techniques45 (recall@5 0.515 ± 0.020, 2.3× a generic zero-shot sentence embedder).46 47Caveat measured in the accompanying paper: zero-shot ranking of48techniques *absent from training* does **not** benefit from this49fine-tuning — in a five-fold label-holdout evaluation the fine-tuned50encoder ranked held-out techniques below a generic MiniLM embedder.51Rankings for techniques outside the 53-technique training vocabulary52should be treated as no better than generic semantic similarity.53 54## Usage55 56The repository ships `technique_texts.json` (the exact technique texts57used at training time) and the scoring calibration in58`config.biencoder`:59 60```python61import json, torch62from huggingface_hub import hf_hub_download63from transformers import AutoModel, AutoTokenizer64 65model_id = "CIRCL/vulnerability-attack-technique-biencoder"66tokenizer = AutoTokenizer.from_pretrained(model_id)67encoder = AutoModel.from_pretrained(model_id).eval()68cfg = encoder.config.biencoder69texts = json.load(open(hf_hub_download(model_id, "technique_texts.json")))70 71def embed(batch, max_length=512):72 enc = tokenizer(batch, padding=True, truncation=True,73 max_length=max_length, return_tensors="pt")74 hidden = encoder(**enc).last_hidden_state75 mask = enc["attention_mask"].unsqueeze(-1)76 pooled = (hidden * mask).sum(1) / mask.sum(1)77 return torch.nn.functional.normalize(pooled, dim=-1)78 79techniques = sorted(texts)80with torch.no_grad():81 technique_emb = embed([texts[t] for t in techniques],82 cfg["technique_max_length"])83 cve_emb = embed(["Improper neutralization of special elements used "84 "in an OS command in the web management interface..."])85scores = cfg["logit_scale"] * (cve_emb @ technique_emb.T) + cfg["logit_bias"]86for idx in scores[0].topk(5).indices:87 print(techniques[idx], float(scores[0][idx]))88```89 90Evaluation and stratified breakdowns are reproducible with91`vulntrain-validate-attack-classification --method biencoder --model92CIRCL/vulnerability-attack-technique-biencoder` (add `--candidates full`93for open-vocabulary ranking over all active parent techniques).94 95## Intended use and limitations96 97The model generates **candidate techniques for analyst review**, not98authoritative mappings. Technique-to-CVE mapping involves analyst99judgment; the training labels inherit the CTID methodology's100subjectivity, and the gold set over-represents exploited and enriched101CVEs. English descriptions only; parent-level techniques only.102 103## References104 105- Bonhomme, C., & Dulaunoy, A. (2026). *Mapping CVEs to MITRE ATT&CK106 Techniques: A Curated Gold-Set Classifier and the Limits of107 LLM-Assisted Label Expansion.* [arXiv:2607.25572](https://arxiv.org/abs/2607.25572)108- Bonhomme, C., & Dulaunoy, A. (2026). *Beyond the Description:109 Structured Metadata and Label Semantics for CVE-to-ATT&CK Mapping.*110 (follow-up paper, in preparation — source of all numbers above)111- Trained with [VulnTrain](https://github.com/vulnerability-lookup/VulnTrain)112 as part of the [Vulnerability-Lookup](https://vulnerability.circl.lu) project.113 