Team Ai
Modelpublic

AXONVERTEX-AI-RESEARCH/cve-cvss-vector-modernbert-base

sourceHugging Faceapache-2.0updated 5d agoView on Hugging Face
0likes14downloads
Model Card

CVE → CVSS v3.1 vector predictor (ModernBERT-base)

Given an English vulnerability description, this model predicts all eight CVSS v3.1 base metrics (AV, AC, PR, UI, S, C, I, A) and, from them, the base score and severity. It returns per-metric probabilities, a severity distribution and a list of uncertain metrics, so a calling system (or an agent) knows when to trust the estimate and when to escalate.

  • —Architecture: fine-tuned ModernBERT-base encoder (≈150M parameters), masked mean pooling, one linear head per CVSS metric.
  • —Training data: ≈107,000 CVE Records from the CVE List published before July 2025.
  • —Evaluation: 57,611 CVEs published January–October 2026, all after the training period (temporal split).
  • —Output: always a valid CVSS v3.1 base vector, scored with the official v3.1 formula.
  • —Use it for: fast triage and a first estimate. It is not an authoritative score: CVSS is assigned by the CNA, NVD or your own analysts.

Built by AXONVERTEX AI Research. The pipeline-integration guide is in `PIPELINE_INTEGRATION.md`.

Results at a glance

Test set: CVEs published 2026-01-01 to early October 2026, excluding Oracle (n = 55,131; see Evaluation for why). 95% bootstrap intervals in brackets.

WhatResult
Exact match, all 8 metrics31.3% [30.9, 31.7]
Base-score mean absolute error (expected score)1.02 [1.02, 1.03]
Base score within ±1.059.0%
Severity accuracy (most likely severity)63.2% [62.8, 63.6]
Mean per-metric accuracy (8 metrics)82.6%

For comparison on the same rows: a model that knows only which CNA published the CVE gets 14.3% exact match; TF-IDF + logistic regression gets 29.9%.

Quick start

bash
pip install -U torch transformers huggingface_hub safetensors numpy
python
import os, sys
from huggingface_hub import hf_hub_download

REPO = "AXONVERTEX-AI-RESEARCH/cve-cvss-vector-modernbert-base"
sys.path.insert(0, os.path.dirname(hf_hub_download(REPO, "cvss_predictor.py")))
from cvss_predictor import CVSSPredictor

predictor = CVSSPredictor(REPO)          # CPU or GPU, picked automatically
r = predictor.predict(
    "SQL injection in the login form of Acme Portal 2.3 allows unauthenticated remote attackers "
    "to execute arbitrary SQL commands via the username parameter."
)
print(r["vector"], r["base_score"], r["most_likely_severity"], r["uncertain_metrics"])

Command line: python cvss_predictor.py "description one" "description two". HTTP API: uvicorn serve:app --port 8000, then POST /predict with {"descriptions": [...]} (see serve.py).

Example predictions from the trained model:

Description (abridged)Predicted vectorScoreSeverity
SQL injection in a login form, unauthenticated remote attackersAV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H9.8CRITICAL
Stored XSS in a comment editor, authenticated usersAV:N/AC:L/PR:L/UI:R/S:C/C:L/I:L/A:N5.4MEDIUM
Race condition in a device driver, local low-privileged user, use-after-freeAV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H7.0HIGH

Output contract

predict(text) returns one JSON-serialisable dict per description:

FieldTypeMeaning
vectorstringMost probable valid CVSS v3.1 base vector, e.g. CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H. Vectors with no impact (C:N/I:N/A:N) are never returned.
base_scorefloatCVSS v3.1 base score of vector.
severitystringSeverity of vector: LOW, MEDIUM, HIGH or CRITICAL.
vector_probabilityfloatModel probability of that exact vector among all 2,592 valid vectors. Often low (many vectors are close); compare across findings rather than reading it as a percentage chance.
most_likely_severitystringSeverity with the highest total probability across all vectors. Use this for triage queues; it is more accurate than severity (63.2% vs 61.7%).
severity_probabilitiesdictProbability of LOW / MEDIUM / HIGH / CRITICAL.
expected_scorefloatProbability-weighted base score; the lowest-error point estimate (MAE 1.02).
metricsdictPer metric: name, predicted value, value_name and probabilities over its values.
uncertain_metricslistMetrics whose top probability is below 0.7 (configurable). Candidates for review or for an LLM that can see the code.

score_vector(vector) scores a vector proposed elsewhere (e.g. by an LLM) with the same formula, for like-for-like comparison.

The probabilities come from a softmax per metric; their calibration has not been measured. The 0.7 threshold is a heuristic.

Use as a tool in an agent pipeline

Function-calling definition (works with OpenAI-style and most open-model tool formats):

json
{
  "name": "predict_cvss",
  "description": "Estimate the CVSS v3.1 base vector, base score and severity of a vulnerability from an English, CVE-style description. Returns per-metric probabilities and the metrics the model is unsure about. The result is an estimate for triage, not an official score.",
  "parameters": {
    "type": "object",
    "properties": {
      "descriptions": {
        "type": "array",
        "items": {"type": "string"},
        "description": "One description per vulnerability: affected component, vulnerability type, who can exploit it (remote/local, authentication needed), whether user interaction is needed, and the impact."
      }
    },
    "required": ["descriptions"]
  }
}

Rules for an agent using this tool:

  1. 1.Describe the vulnerability the way a CVE description does. Do not include a CVSS score or vector in the text.
  2. 2.Report most_likely_severity and expected_score for prioritisation, and vector for detail. Label them as model estimates.
  3. 3.If uncertain_metrics is non-empty, decide those metrics from evidence (code, configuration, exploit preconditions), keep the confident ones, and rescore with score_vector.
  4. 4.Never lower a finding's priority on this model's output alone when there is evidence of active exploitation.
  5. 5.Send HIGH and CRITICAL findings to human review.

`PIPELINE_INTEGRATION.md` covers the full integration with a three-model secure-coding pipeline (Antares-1B, Foundation-Sec-8B-Reasoning, Gemma 12B): reference flow, prompts, escalation logic, deployment and how to evaluate it.

Training data and preprocessing

  • —Source: CVE List, cvelistV5 repository snapshot of October 2026, CVE JSON 5.x records with state PUBLISHED.
  • —Text: the CNA's English description, whitespace collapsed. Descriptions shorter than 30 characters were dropped.
  • —Label: the CVSS v3.x base vector. The CNA's vector is used first, with an ADP vector (mainly CISA's vulnrichment container) as fallback. v3.1 is preferred over v3.0. Temporal and environmental metrics are ignored. NVD's own scores are not in the CVE List and were not used.
  • —Leakage removal: some CNAs write the answer into the description, e.g. Oracle (CVSS 3.1 Base Score 4.9 ... CVSS Vector: (CVSS:3.1/...)), Concrete CMS and dotCMS. Scores, vectors and CVSS-calculator links are stripped from all descriptions before training, and the loader applies the same cleaning to inputs.
  • —Deduplication: exact-duplicate descriptions removed; the earliest record is kept.
  • —Temporal split by datePublished:
SplitPublishedSize
Trainbefore 2025-07-01≈107,000
Validation2025-07-01 to 2025-12-31used for checkpoint selection
Test2026-01-01 to early October 202657,611

Test labels come from CNAs (46,441), CISA-ADP (10,962) and Red Hat (208).

Training procedure

SettingValue
Base modelanswerdotai/ModernBERT-base
Objectivemean of 8 cross-entropy losses, square-root inverse-frequency class weights
Max length512 tokens
Batch size / learning rate64 / 5e-5, AdamW, weight decay 0.01, linear schedule with 6% warm-up
Epochs4; checkpoint with the best validation mean macro-F1 kept (epoch 3)
Precision / hardwarebf16 autocast, one NVIDIA A100; ≈2.5 minutes per epoch

Validation at the selected epoch: mean macro-F1 0.797, exact match 43.1%. Validation plateaued after the first epoch (0.789); results/training_history.csv has the full curve.

The training notebook is in `training/CVE_finetune_A100.ipynb`.

Evaluation

Why Oracle is excluded from the headline numbers

Oracle's descriptions are generated from the CVSS vector itself: "Easily exploitable" means AC:L, "high privileged attacker" means PR:H, "scope change" means S:C, "complete DOS" means A:H. The model reaches 99.0% exact match on Oracle CVEs even after the explicit vector strings are removed. That is template decoding, not vulnerability understanding, so headline figures exclude Oracle (2,450 test CVEs). Results with Oracle are shown alongside.

Decoding

Choosing each metric independently can produce a no-impact vector (score 0), which almost never occurs in real CVEs; the independent decoder did so 680 times on the test set. The released loader uses joint decoding over all 2,592 valid vectors instead.

Test subsetDecodingExact vectorScore MAEWithin ±1.0Severity accSeverity macro-F1
Excl. Oracle (n = 55,131)per-metric argmax31.2%1.14158.4%61.2%0.524
joint, impact > 0 (vector, severity)31.3%1.08458.8%61.7%0.525
marginal (expected_score, most_likely_severity)—1.02459.0%63.2%0.520
All (n = 57,581)joint, impact > 034.2%1.03860.5%63.4%0.547
marginal—0.98460.8%64.8%0.545

Severity macro-F1 is over LOW/MEDIUM/HIGH/CRITICAL; the 30 test CVEs whose true vector has no impact are excluded from score and severity metrics.

Baselines (same test rows, excluding Oracle, same decoding)

MethodExact vectorMean metric macro-F1Score MAESeverity accSeverity macro-F1
Most frequent training vector5.8%0.3032.8511.8%0.053
CNA prior: the assigner's most frequent vector (no text)14.3%0.4991.6742.7%0.297
TF-IDF (word 1–2-grams) + logistic regression29.9%0.7211.1259.6%0.489
This model31.3%0.7421.0861.7%0.525

The description text more than doubles exact match over knowing the assigner alone. A linear model captures most of the signal; the encoder's clearest gains are on the rarer classes (severity macro-F1 +3.7 points).

Per metric (all 57,611 test CVEs, per-metric argmax)

MetricAccuracyMacro-F1Majority-class accuracyWeakest class
AV Attack Vector90.9%0.71680.1% (N)Adjacent (F1 0.49)
AC Attack Complexity86.0%0.69785.8% (L)High (recall 0.44)
PR Privileges Required78.3%0.71555.3% (N)High (F1 0.56)
UI User Interaction90.6%0.86576.7% (N)Required (recall 0.76)
S Scope87.2%0.78779.9% (U)Changed (recall 0.60)
C Confidentiality77.9%0.75951.0% (H)Low (recall 0.63)
I Integrity77.6%0.76942.5% (H)Low (recall 0.68)
A Availability77.7%0.72143.1% (H)Low (recall 0.45)

Attack Complexity barely beats always answering "Low": descriptions rarely state the conditions that make an attack complex. With per-metric decoding, CRITICAL is recovered 49% of the time (mostly predicted HIGH, usually one metric away) and LOW 33% of the time.

By CNA (15 largest in the test set, per-metric argmax)

CNATest CVEsExact vectorScore MAE
GitHub_M7,38916.1%1.49
VulnCheck6,67828.4%1.14
VulDB5,46334.4%1.07
Patchstack4,11430.1%1.16
mitre3,29232.2%1.29
Wordfence3,24167.4%0.57
Linux3,05340.3%0.71
Chrome2,78242.1%1.01
oracle2,45099.0%0.01
microsoft1,74345.7%0.67
WPScan1,51031.0%1.18
ibm1,15027.0%1.28
redhat1,06216.6%1.33
apache82826.8%1.43
apple75522.4%1.10

The model learns each CNA's scoring conventions as well as the vulnerability. Consistent scorers are predicted well, such as Wordfence and the Linux kernel. Advisories scored by many different maintainers (GitHub_M) are predicted poorly. CNA-scored and CISA-ADP-scored labels are equally predictable (34.3% vs 33.5% exact).

Limitations and appropriate use

  • —An estimate, not a score. Exact match is 31%, and 59% of scores land within ±1.0. Use it to sort and pre-fill, not to publish.
  • —Scoring conventions differ between CNAs, and human scorers disagree. Part of the remaining error is label inconsistency rather than model error.
  • —It reads text, not code. Its quality depends on how well the description states attack vector, privileges, user interaction and impact. In a code-scanning pipeline, have an LLM write the finding in CVE style first (see the integration guide).
  • —Weak spots: Attack Complexity, Scope, the Low value of C/I/A, and the LOW and CRITICAL severity bands.
  • —English only; CVSS v3.1 only (not v4.0). Trained on CVEs published up to June 2025; vulnerability language drifts, so retrain periodically.
  • —Probabilities are uncalibrated. Use them for ranking and for spotting uncertainty.
  • —Not for dependency scanning. For "is my library version affected", use OSV.dev or the GitHub Advisory Database.

Files

FileContents
model.safetensors, config.json, tokenizer*.jsonFine-tuned ModernBERT encoder and tokenizer (AutoModel.from_pretrained loads the encoder alone)
heads.safetensorsThe eight linear heads
cvss_config.jsonLabel order per head, max length, pooling, decoding and training settings
cvss_predictor.pyLoader, joint decoding, CVSS 3.1 calculator, CLI
serve.pyFastAPI server: /predict, /score, /health
PIPELINE_INTEGRATION.mdHow to use the model in an agentic secure-coding pipeline
results/Test metrics, decoding and baseline comparisons, confidence intervals, training history, per-CVE test predictions
training/CVE_finetune_A100.ipynbData parsing, training and evaluation notebook
CVE_TERMS_OF_USE.mdNotice for the CVE content used

License and data notice

Model weights and code: Apache 2.0. The training data and results/test_predictions.parquet come from the CVE List. CVE® is a registered trademark of The MITRE Corporation; CVE content is used under the CVE Terms of Use. See `CVE_TERMS_OF_USE.md`.

Citation

bibtex
@misc{axonvertex2026cvssvector,
  title  = {CVE to CVSS v3.1 Vector Prediction with a Fine-Tuned ModernBERT Encoder},
  author = {Dasgupta, Krishnendu},
  year   = {2026},
  publisher = {AXONVERTEX AI Research},
  howpublished = {\url{https://huggingface.co/AXONVERTEX-AI-RESEARCH/cve-cvss-vector-modernbert-base}}
}