Mizhhall/cpp-memsafe-v5
Qwen2.5-Coder-3B — C/C++ Memory-Safety Detector (SFT v5) PiSSA-SFT fine-tune of Qwen/Qwen2.5-Coder-3B-Instruct that performs a 7-step structured analysis of a C/C++ fragment and concludes VULNERABLE, NOT_VULNERABLE, or INSUFFICIENT_CONTEXT with a CWE id. Its defining property is calibrated abstention: it commits a verdict only when the deciding size/bound is establishable from the visible fragment, otherwise it abstains instead of guessing. This fixes the verdict-collapse of… See the full description on the dataset page: https://huggingface.co/datasets/Mizhhall/cpp-memsafe-v5.
Qwen2.5-Coder-3B — C/C++ Memory-Safety Detector (SFT v5)
PiSSA-SFT fine-tune of Qwen/Qwen2.5-Coder-3B-Instruct that performs a 7-step structured analysis of a C/C++ fragment and concludes VULNERABLE, NOT_VULNERABLE, or INSUFFICIENT_CONTEXT with a CWE id.
Its defining property is calibrated abstention: it commits a verdict only when the deciding size/bound is establishable from the visible fragment, otherwise it abstains instead of guessing. This fixes the verdict-collapse of earlier versions (everything labelled NOT_VULNERABLE on real code).
Code + data + docs live here (GitHub). Model weights (~5.8 GB) → HuggingFace model repo (push_model_hf.py). Training dataset → HuggingFace dataset repo (push_hf.py).
Results
Real-world (PrimeVul/DiverseVul buffer-family), shipped .up5 contract. Raw per-split tables in metrics/RESULTS.txt.
vs prior v4: DECIDED recall 0.05, CONSERVATIVE recall 0.49/0.53. Two scoring modes (data_cleaning/score3way.py): DECIDED = acc/P/R/F1 over rows it commits a verdict on (reliability when it acts); CONSERVATIVE = abstain-on-a-true -vuln counts as a risk-flag (the security triage metric).
Intended use
- Security triage of C/C++ functions for memory-safety defects (buffer overflow, OOB read/write, UAF, double-free; CWE-119/120/125/787/416/415/...).
- Best run in CONSERVATIVE mode: treat both VULNERABLE and INSUFFICIENT_CONTEXT as "needs human review" — flags ~86% of true vulnerabilities with ~5% hard false-positives on safe real-world code.
- A prioritiser, not a replacement for a sound static analyzer or human review.
Inference contract (REQUIRED)
[SINK] [DESTINATION] [SOURCE] [CONSTRAINTS] [SCRATCHPAD] [MATH] [DECISION] [CONCLUSION]Run inference under the `.up5` contract (scripts/convert_to_up5.py, data_cleaning/*.up5.jsonl). The system prompt must (1) permit the 3-way verdict and (2) declare the [DECISION] rule — a one-line [DECISION]: mapping [MATH]→verdict, emitted before [CONCLUSION]. Surfacing [DECISION] lifts real-world precision (0.52→0.63) at no recall cost. Parse the verdict from [CONCLUSION]. The plain .up contract (no DECISION_RULE) also works but with lower decided-precision.
Dataset (SFT v5 — latest)
data/train/sft_train_v5.jsonl — 4177 rows, 3-way. Validation data/validation/sft_v5_val.jsonl (220).
Composition gates (enforced by scripts/assemble_v5.py): abstention fraction 0.384 (≥0.30), decide VULN:SAFE 0.99:1 (≤1.15), [DECISION] in 100% of rows.
Provenance & gates. Decided-real and grounded-INSUFFICIENT rows come from a blind 32B teacher (Qwen/Qwen2.5-Coder-32B-Instruct, data_cleaning/regen_pipeline.py) run with no label in the prompt, kept only if they pass four gates: (1) format, (2) faithfulness (every cited number/sink exists in the code), (3) verdict == ground truth (rejection sampling), (4) identifier-overlap grounding (a decided row must cite a changed diff identifier — decided for the right reason). Contrastive source: PrimeVul paired before/after buffer-family functions (511 pairs). A pre-mine de-risk probe confirmed symmetric pre-fix-VULN / post-fix-SAFE decidability (1.13×, PASS — metrics/derisk_probe_result.json) so the bucket does not skew VULNERABLE. Deterministic INSUFFICIENT rows (scripts/build_abstention_deterministic.py) are faithful-by-construction (abstention is about the ABSENCE of an in-fragment bound, verifiable from code).
Held-out eval. data_cleaning/{rw_test,rw_bench,val}.bof.jsonl (PrimeVul/DiverseVul, never in train) → upgraded to .up / .up5 and scored 3-way.
Data sources. NIST Juliet C/C++ 1.3 (<https://samate.nist.gov/SARD/test-suites/112>, SHA-256 ada9d7e1c323d283446df3f55bdee0d00bda1fed786785fe98764d58688f38eb, CC0-1.0); PrimeVul (paired before/after); DiverseVul.
Lineage: the earlier binary corpus (sft_train_v2+dpo_*, marker-based, no abstention) has been removed fromdata/— it remains in earlier git history and the previous HuggingFace dataset release. v5 dropped the marker (a shortcut the model learned instead of detection) and added INSUFFICIENT_CONTEXT.
Reproduce
python scripts/build_contrastive_pool.py # 1. PrimeVul before/after pairs (511)
python scripts/derisk_probe.py # 2. symmetry gate, GPU (optional; got 1.13 PASS)
python data_cleaning/regen_pipeline.py \ # 3. blind 32B teacher mine (verdict==gt + faithful)
--model Qwen/Qwen2.5-Coder-32B-Instruct --tp 2 --kbest 4 \
--in data_cleaning/contrastive_mine_input.jsonl --out data_cleaning/contrastive_mined.jsonl
python scripts/build_v5_contrastive.py # 4. grounding gate + pair-completeness + [DECISION]
python scripts/assemble_v5.py # 5. assemble + GATES (abstain>=30%, VULN:SAFE<=1.15)
CUDA_VISIBLE_DEVICES=0 RUN_SFT=1 RUN_DPO=0 SFT_FILENAME=sft_train_v5.jsonl \
OUTPUT_DIR=./qwen_output_3b_v5 MAX_SEQ_LEN=3072 python train_qwen_pissa.py # 6. PiSSA SFT
python scripts/make_up5.py # 7a. build the shipped .up5 eval splits
bash scripts/run_v5_dec_eval.sh # 7b. eval (.up5) + 3-way score + calibrationPending from server
SSH dropped during prep — pull these final artifacts with sh pull_from_server.sh <password>, then upload the model with python push_model_hf.py:
data/train/sft_train_v5.jsonl+sft_v5_val.jsonldata_cleaning/{contrastive_pool,contrastive_mined,v5_contrastive,v5_insufficient_mined,abstain_full}.jsonl,oldfamily_sysprompts.json,{rw_test,rw_bench,val}.{up,up5}.jsonlqwen_output_3b_v5/merged/(~5.8 GB weights → HuggingFace, not GitHub)
Limitations
- Fragment-level information limit: abstains ~72–80% on real fragments where the deciding bound is in a caller/callee/global — honest, not a bug. Higher coverage needs inter-procedural context or RAG.
- DECIDED recall on real is modest (0.33–0.46) by design; use CONSERVATIVE mode for triage.
- The
confidence=token is uninformative — do not threshold on it. - Possible synthetic-benchmark overlap (Juliet) — treat Juliet numbers as optimistic.
Install / train
pip install -r requirements.txtTuned for a single 48 GB GPU: BF16, PiSSA rank 16, seq 3072. Env vars in train_qwen_pissa.py override defaults.
