Team Ai
Datasetpublic

Mizhhall/cpp-memsafe-v6

Qwen2.5-Coder-3B — C/C++ Memory-Safety Detector (SFT v6) PiSSA-SFT fine-tune of Qwen/Qwen2.5-Coder-3B-Instruct that performs a structured analysis of a C/C++ fragment and concludes VULNERABLE, NOT_VULNERABLE, or INSUFFICIENT_CONTEXT with a CWE id. Two defining properties: Calibrated abstention — it commits a verdict only when the deciding size/bound/lifetime is establishable from the visible fragment, otherwise it abstains instead of guessing (fixes the verdict-collapse of… See the full description on the dataset page: https://huggingface.co/datasets/Mizhhall/cpp-memsafe-v6.

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes8downloads
Dataset Card

Qwen2.5-Coder-3B — C/C++ Memory-Safety Detector (SFT v6)

PiSSA-SFT fine-tune of Qwen/Qwen2.5-Coder-3B-Instruct that performs a structured analysis of a C/C++ fragment and concludes VULNERABLE, NOT_VULNERABLE, or INSUFFICIENT_CONTEXT with a CWE id.

Two defining properties:

  • —Calibrated abstention — it commits a verdict only when the deciding size/bound/lifetime is establishable from the visible fragment, otherwise it abstains instead of guessing (fixes the verdict-collapse of earlier versions: everything labelled NOT_VULNERABLE on real code).
  • —Actionable abstention (v6) — when it abstains it emits a [MISSING] line naming the exact external symbol whose value decides safety, plus a conditional verdict ("safe IFF the capacity of dst covers len"). This turns "I can't tell" into a finding a human can act on and a ready-made retrieval query for a future agent/RAG layer — no RAG built yet.

Code + docs live here (GitHub). Weights and the full dataset are on HuggingFace:

  • —Model: <https://huggingface.co/Mizhhall/qwen2.5-coder-3b-memsafe-v6>
  • —Dataset: <https://huggingface.co/datasets/Mizhhall/cpp-memsafe-v6>
  • —(prior v5, detection-only: <https://huggingface.co/Mizhhall/qwen2.5-coder-3b-memsafe-v5>)

Results

Real-world (PrimeVul/DiverseVul buffer-family), shipped .up6 contract. Full tables: metrics/RESULTS_V6.txt (v5: metrics/RESULTS.txt).

splitDECIDED acc / P / R / F1CONSERVATIVE recallabstain%
realworld_test (323)0.578 / 0.67 / 0.44 / 0.5330.8574
realworld_benchmark (375)0.536 / 0.60 / 0.35 / 0.4440.8374
validation (499)0.862 / 0.88 / 0.88 / 0.8790.8818

Actionable abstention (the v6 addition): of the rows it abstains on, ~99% emit [MISSING], 60–87% name the specific missing in-code symbol (faithfully), and 62–91% give the conditional verdict. Inference specificity exceeds the deterministic training labels (56%) — the model generalized the pattern.

Two scoring modes (data_cleaning/score3way.py): DECIDED = acc/P/R/F1 over rows it commits a verdict on; CONSERVATIVE = abstain-on-a-true-vuln counts as a risk-flag (the security triage metric). Run in CONSERVATIVE mode for triage: ~85% of true vulns flagged, ~5% hard false-positives on safe code.

Output contract

text
[SINK] [DESTINATION] [SOURCE] [CONSTRAINTS] [SCRATCHPAD] [MATH] [MISSING] [DECISION] [CONCLUSION]

Run inference under the `.up6` contract (scripts/make_up6.py, data_cleaning/*.up6.jsonl). The system prompt declares the 3-way verdict, the [DECISION] bridge, and the [MISSING] rule (name the external symbol when abstaining; none otherwise; [DECISION] is conditional when abstaining). Parse the verdict from [CONCLUSION], the retrieval target from [MISSING].

Dataset (SFT v6)

data/train/sft_train_v6.jsonl — 4177 rows (V=1282, N=1290, I=1605; abstention 38.4%, decide VULN:SAFE 0.99:1). Built from the v5 set by enriching every INSUFFICIENT row with a deterministic, CWE-aware [MISSING] + conditional [DECISION]:

  • —bounds defects (CWE-119/120/125/787/…): [MISSING] = the declared capacity of the destination buffer; conditional = "capacity covers the access → NODEFECT, else DEFECTPRESENT".
  • —lifetime defects (CWE-416/415/590 UAF/double-free): [MISSING] = whether the pointer is still live; conditional = "not freed → NO_DEFECT, else use-after-free".

Only identifiers that appear in the code are named (faithful by construction). The verdicts are unchanged from v5 — v6 only enriches the abstention output. Provenance of the underlying rows (blind 32B-teacher mine + grounding gate + de-risk probe) is unchanged; see metrics/derisk_probe_result.json.

Held-out eval: data_cleaning/{rw_test,rw_bench,val}.bof.jsonl (PrimeVul/DiverseVul, never in train) → upgraded to .up6 and scored 3-way.

Data sources: NIST Juliet C/C++ 1.3 (<https://samate.nist.gov/SARD/test-suites/112>, SHA-256 ada9d7e1c323d283446df3f55bdee0d00bda1fed786785fe98764d58688f38eb, CC0-1.0); PrimeVul (paired before/after); DiverseVul.

Reproduce

bash
# v5 base pipeline (contrastive mine -> grounded abstention -> assemble -> SFT)
python scripts/build_contrastive_pool.py
python data_cleaning/regen_pipeline.py --model Qwen/Qwen2.5-Coder-32B-Instruct --tp 2 --kbest 4 \
  --in data_cleaning/contrastive_mine_input.jsonl --out data_cleaning/contrastive_mined.jsonl
python scripts/build_v5_contrastive.py
python scripts/assemble_v5.py
# v6: enrich abstentions + retrain + eval
python scripts/build_v6_conditional.py      # v5 train -> v6 ([MISSING] + conditional, CWE-aware)
python scripts/make_up6.py                  # .up5 eval splits -> .up6
bash   scripts/run_v6_chain.sh              # PiSSA SFT v6 + eval (.up6) + score + [MISSING] audit

Roadmap (not built)

v6's [MISSING] field is the interface for the next step: an agentic-RAG loop where INSUFFICIENT_CONTEXT triggers retrieval of the named symbol's definition, then a re-decide. The recommended first experiment is an oracle probe (give the model the real missing definition, measure abstention→correct-decision flip-rate) before building any retrieval corpus.

Limitations

  • —Fragment-level information limit: abstains ~74–82% on real fragments where the deciding bound/lifetime is in a caller/callee/global — honest, not a bug. [MISSING] names what's needed; higher coverage needs inter-procedural context or RAG.
  • —DECIDED recall on real is modest (0.35–0.44) by design; use CONSERVATIVE mode for triage.
  • —The confidence= token is uninformative — do not threshold on it.
  • —Possible synthetic-benchmark overlap (Juliet) — treat Juliet numbers as optimistic.

Install

bash
pip install -r requirements.txt

Tuned for a single 48 GB GPU: BF16, PiSSA rank 16, seq 3072. Env vars in train_qwen_pissa.py override defaults.