Mizhhall/cpp-memsafe-v6
Qwen2.5-Coder-3B — C/C++ Memory-Safety Detector (SFT v6) PiSSA-SFT fine-tune of Qwen/Qwen2.5-Coder-3B-Instruct that performs a structured analysis of a C/C++ fragment and concludes VULNERABLE, NOT_VULNERABLE, or INSUFFICIENT_CONTEXT with a CWE id. Two defining properties: Calibrated abstention — it commits a verdict only when the deciding size/bound/lifetime is establishable from the visible fragment, otherwise it abstains instead of guessing (fixes the verdict-collapse of… See the full description on the dataset page: https://huggingface.co/datasets/Mizhhall/cpp-memsafe-v6.
Qwen2.5-Coder-3B — C/C++ Memory-Safety Detector (SFT v6)
PiSSA-SFT fine-tune of Qwen/Qwen2.5-Coder-3B-Instruct that performs a structured analysis of a C/C++ fragment and concludes VULNERABLE, NOT_VULNERABLE, or INSUFFICIENT_CONTEXT with a CWE id.
Two defining properties:
- Calibrated abstention — it commits a verdict only when the deciding size/bound/lifetime is establishable from the visible fragment, otherwise it abstains instead of guessing (fixes the verdict-collapse of earlier versions: everything labelled NOT_VULNERABLE on real code).
- Actionable abstention (v6) — when it abstains it emits a
[MISSING]line naming the exact external symbol whose value decides safety, plus a conditional verdict ("safe IFF the capacity ofdstcoverslen"). This turns "I can't tell" into a finding a human can act on and a ready-made retrieval query for a future agent/RAG layer — no RAG built yet.
Code + docs live here (GitHub). Weights and the full dataset are on HuggingFace:
- Model: <https://huggingface.co/Mizhhall/qwen2.5-coder-3b-memsafe-v6>
- Dataset: <https://huggingface.co/datasets/Mizhhall/cpp-memsafe-v6>
- (prior v5, detection-only: <https://huggingface.co/Mizhhall/qwen2.5-coder-3b-memsafe-v5>)
Results
Real-world (PrimeVul/DiverseVul buffer-family), shipped .up6 contract. Full tables: metrics/RESULTS_V6.txt (v5: metrics/RESULTS.txt).
Actionable abstention (the v6 addition): of the rows it abstains on, ~99% emit [MISSING], 60–87% name the specific missing in-code symbol (faithfully), and 62–91% give the conditional verdict. Inference specificity exceeds the deterministic training labels (56%) — the model generalized the pattern.
Two scoring modes (data_cleaning/score3way.py): DECIDED = acc/P/R/F1 over rows it commits a verdict on; CONSERVATIVE = abstain-on-a-true-vuln counts as a risk-flag (the security triage metric). Run in CONSERVATIVE mode for triage: ~85% of true vulns flagged, ~5% hard false-positives on safe code.
Output contract
[SINK] [DESTINATION] [SOURCE] [CONSTRAINTS] [SCRATCHPAD] [MATH] [MISSING] [DECISION] [CONCLUSION]Run inference under the `.up6` contract (scripts/make_up6.py, data_cleaning/*.up6.jsonl). The system prompt declares the 3-way verdict, the [DECISION] bridge, and the [MISSING] rule (name the external symbol when abstaining; none otherwise; [DECISION] is conditional when abstaining). Parse the verdict from [CONCLUSION], the retrieval target from [MISSING].
Dataset (SFT v6)
data/train/sft_train_v6.jsonl — 4177 rows (V=1282, N=1290, I=1605; abstention 38.4%, decide VULN:SAFE 0.99:1). Built from the v5 set by enriching every INSUFFICIENT row with a deterministic, CWE-aware [MISSING] + conditional [DECISION]:
- bounds defects (CWE-119/120/125/787/…):
[MISSING]= the declared capacity of the destination buffer; conditional = "capacity covers the access → NODEFECT, else DEFECTPRESENT". - lifetime defects (CWE-416/415/590 UAF/double-free):
[MISSING]= whether the pointer is still live; conditional = "not freed → NO_DEFECT, else use-after-free".
Only identifiers that appear in the code are named (faithful by construction). The verdicts are unchanged from v5 — v6 only enriches the abstention output. Provenance of the underlying rows (blind 32B-teacher mine + grounding gate + de-risk probe) is unchanged; see metrics/derisk_probe_result.json.
Held-out eval: data_cleaning/{rw_test,rw_bench,val}.bof.jsonl (PrimeVul/DiverseVul, never in train) → upgraded to .up6 and scored 3-way.
Data sources: NIST Juliet C/C++ 1.3 (<https://samate.nist.gov/SARD/test-suites/112>, SHA-256 ada9d7e1c323d283446df3f55bdee0d00bda1fed786785fe98764d58688f38eb, CC0-1.0); PrimeVul (paired before/after); DiverseVul.
Reproduce
# v5 base pipeline (contrastive mine -> grounded abstention -> assemble -> SFT)
python scripts/build_contrastive_pool.py
python data_cleaning/regen_pipeline.py --model Qwen/Qwen2.5-Coder-32B-Instruct --tp 2 --kbest 4 \
--in data_cleaning/contrastive_mine_input.jsonl --out data_cleaning/contrastive_mined.jsonl
python scripts/build_v5_contrastive.py
python scripts/assemble_v5.py
# v6: enrich abstentions + retrain + eval
python scripts/build_v6_conditional.py # v5 train -> v6 ([MISSING] + conditional, CWE-aware)
python scripts/make_up6.py # .up5 eval splits -> .up6
bash scripts/run_v6_chain.sh # PiSSA SFT v6 + eval (.up6) + score + [MISSING] auditRoadmap (not built)
v6's [MISSING] field is the interface for the next step: an agentic-RAG loop where INSUFFICIENT_CONTEXT triggers retrieval of the named symbol's definition, then a re-decide. The recommended first experiment is an oracle probe (give the model the real missing definition, measure abstention→correct-decision flip-rate) before building any retrieval corpus.
Limitations
- Fragment-level information limit: abstains ~74–82% on real fragments where the deciding bound/lifetime is in a caller/callee/global — honest, not a bug.
[MISSING]names what's needed; higher coverage needs inter-procedural context or RAG. - DECIDED recall on real is modest (0.35–0.44) by design; use CONSERVATIVE mode for triage.
- The
confidence=token is uninformative — do not threshold on it. - Possible synthetic-benchmark overlap (Juliet) — treat Juliet numbers as optimistic.
Install
pip install -r requirements.txtTuned for a single 48 GB GPU: BF16, PiSSA rank 16, seq 3072. Env vars in train_qwen_pissa.py override defaults.
