Team Ai
Datasetpublic

Mizhhall/cpp-memsafe-v5

Qwen2.5-Coder-3B — C/C++ Memory-Safety Detector (SFT v5) PiSSA-SFT fine-tune of Qwen/Qwen2.5-Coder-3B-Instruct that performs a 7-step structured analysis of a C/C++ fragment and concludes VULNERABLE, NOT_VULNERABLE, or INSUFFICIENT_CONTEXT with a CWE id. Its defining property is calibrated abstention: it commits a verdict only when the deciding size/bound is establishable from the visible fragment, otherwise it abstains instead of guessing. This fixes the verdict-collapse of… See the full description on the dataset page: https://huggingface.co/datasets/Mizhhall/cpp-memsafe-v5.

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes11downloads
Dataset Card

Qwen2.5-Coder-3B — C/C++ Memory-Safety Detector (SFT v5)

PiSSA-SFT fine-tune of Qwen/Qwen2.5-Coder-3B-Instruct that performs a 7-step structured analysis of a C/C++ fragment and concludes VULNERABLE, NOT_VULNERABLE, or INSUFFICIENT_CONTEXT with a CWE id.

Its defining property is calibrated abstention: it commits a verdict only when the deciding size/bound is establishable from the visible fragment, otherwise it abstains instead of guessing. This fixes the verdict-collapse of earlier versions (everything labelled NOT_VULNERABLE on real code).

Code + data + docs live here (GitHub). Model weights (~5.8 GB) → HuggingFace model repo (push_model_hf.py). Training dataset → HuggingFace dataset repo (push_hf.py).

Results

Real-world (PrimeVul/DiverseVul buffer-family), shipped .up5 contract. Raw per-split tables in metrics/RESULTS.txt.

splitDECIDED acc / P / R / F1CONSERVATIVE recallabstain%
realworld_test (323)0.508 / 0.63 / 0.33 / 0.4360.8680
realworld_benchmark (375)0.598 / 0.68 / 0.46 / 0.5450.8776
validation (499)0.796 / 0.82 / 0.82 / 0.8210.8422
Juliet held-out (556)0.696 / 0.78 / 0.62 / 0.6920.8872

vs prior v4: DECIDED recall 0.05, CONSERVATIVE recall 0.49/0.53. Two scoring modes (data_cleaning/score3way.py): DECIDED = acc/P/R/F1 over rows it commits a verdict on (reliability when it acts); CONSERVATIVE = abstain-on-a-true -vuln counts as a risk-flag (the security triage metric).

Intended use

  • —Security triage of C/C++ functions for memory-safety defects (buffer overflow, OOB read/write, UAF, double-free; CWE-119/120/125/787/416/415/...).
  • —Best run in CONSERVATIVE mode: treat both VULNERABLE and INSUFFICIENT_CONTEXT as "needs human review" — flags ~86% of true vulnerabilities with ~5% hard false-positives on safe real-world code.
  • —A prioritiser, not a replacement for a sound static analyzer or human review.

Inference contract (REQUIRED)

text
[SINK] [DESTINATION] [SOURCE] [CONSTRAINTS] [SCRATCHPAD] [MATH] [DECISION] [CONCLUSION]

Run inference under the `.up5` contract (scripts/convert_to_up5.py, data_cleaning/*.up5.jsonl). The system prompt must (1) permit the 3-way verdict and (2) declare the [DECISION] rule — a one-line [DECISION]: mapping [MATH]→verdict, emitted before [CONCLUSION]. Surfacing [DECISION] lifts real-world precision (0.52→0.63) at no recall cost. Parse the verdict from [CONCLUSION]. The plain .up contract (no DECISION_RULE) also works but with lower decided-precision.

Dataset (SFT v5 — latest)

data/train/sft_train_v5.jsonl — 4177 rows, 3-way. Validation data/validation/sft_v5_val.jsonl (220).

LabelRows
VULNERABLE (1)1282
NOT_VULNERABLE (0)1290
INSUFFICIENT_CONTEXT1605

Composition gates (enforced by scripts/assemble_v5.py): abstention fraction 0.384 (≥0.30), decide VULN:SAFE 0.99:1 (≤1.15), [DECISION] in 100% of rows.

Source bucketRowsRole
contrastive_mined_v5172blind-teacher decided real (82V/90N), marker-free, grounded
tier1_vuln/safe_v5600 / 600marker-free real decided — decorrelates "marker⇒decide"
synth_vuln/safe_v5600 / 600synthetic anchor (marker-bearing)
real_insufficient_mined_v5705grounded INSUFFICIENT from real near-twin pairs
real_insufficient_det_v5900deterministic, faithful-by-construction INSUFFICIENT

Provenance & gates. Decided-real and grounded-INSUFFICIENT rows come from a blind 32B teacher (Qwen/Qwen2.5-Coder-32B-Instruct, data_cleaning/regen_pipeline.py) run with no label in the prompt, kept only if they pass four gates: (1) format, (2) faithfulness (every cited number/sink exists in the code), (3) verdict == ground truth (rejection sampling), (4) identifier-overlap grounding (a decided row must cite a changed diff identifier — decided for the right reason). Contrastive source: PrimeVul paired before/after buffer-family functions (511 pairs). A pre-mine de-risk probe confirmed symmetric pre-fix-VULN / post-fix-SAFE decidability (1.13×, PASS — metrics/derisk_probe_result.json) so the bucket does not skew VULNERABLE. Deterministic INSUFFICIENT rows (scripts/build_abstention_deterministic.py) are faithful-by-construction (abstention is about the ABSENCE of an in-fragment bound, verifiable from code).

Held-out eval. data_cleaning/{rw_test,rw_bench,val}.bof.jsonl (PrimeVul/DiverseVul, never in train) → upgraded to .up / .up5 and scored 3-way.

Data sources. NIST Juliet C/C++ 1.3 (<https://samate.nist.gov/SARD/test-suites/112>, SHA-256 ada9d7e1c323d283446df3f55bdee0d00bda1fed786785fe98764d58688f38eb, CC0-1.0); PrimeVul (paired before/after); DiverseVul.

Lineage: the earlier binary corpus (sft_train_v2 + dpo_*, marker-based, no abstention) has been removed from data/ — it remains in earlier git history and the previous HuggingFace dataset release. v5 dropped the marker (a shortcut the model learned instead of detection) and added INSUFFICIENT_CONTEXT.

Reproduce

bash
python scripts/build_contrastive_pool.py     # 1. PrimeVul before/after pairs (511)
python scripts/derisk_probe.py               # 2. symmetry gate, GPU (optional; got 1.13 PASS)
python data_cleaning/regen_pipeline.py \      # 3. blind 32B teacher mine (verdict==gt + faithful)
  --model Qwen/Qwen2.5-Coder-32B-Instruct --tp 2 --kbest 4 \
  --in data_cleaning/contrastive_mine_input.jsonl --out data_cleaning/contrastive_mined.jsonl
python scripts/build_v5_contrastive.py       # 4. grounding gate + pair-completeness + [DECISION]
python scripts/assemble_v5.py                # 5. assemble + GATES (abstain>=30%, VULN:SAFE<=1.15)
CUDA_VISIBLE_DEVICES=0 RUN_SFT=1 RUN_DPO=0 SFT_FILENAME=sft_train_v5.jsonl \
  OUTPUT_DIR=./qwen_output_3b_v5 MAX_SEQ_LEN=3072 python train_qwen_pissa.py   # 6. PiSSA SFT
python scripts/make_up5.py                   # 7a. build the shipped .up5 eval splits
bash scripts/run_v5_dec_eval.sh              # 7b. eval (.up5) + 3-way score + calibration

Pending from server

SSH dropped during prep — pull these final artifacts with sh pull_from_server.sh <password>, then upload the model with python push_model_hf.py:

  • —data/train/sft_train_v5.jsonl + sft_v5_val.jsonl
  • —data_cleaning/{contrastive_pool,contrastive_mined,v5_contrastive,v5_insufficient_mined,abstain_full}.jsonl, oldfamily_sysprompts.json, {rw_test,rw_bench,val}.{up,up5}.jsonl
  • —qwen_output_3b_v5/merged/ (~5.8 GB weights → HuggingFace, not GitHub)

Limitations

  • —Fragment-level information limit: abstains ~72–80% on real fragments where the deciding bound is in a caller/callee/global — honest, not a bug. Higher coverage needs inter-procedural context or RAG.
  • —DECIDED recall on real is modest (0.33–0.46) by design; use CONSERVATIVE mode for triage.
  • —The confidence= token is uninformative — do not threshold on it.
  • —Possible synthetic-benchmark overlap (Juliet) — treat Juliet numbers as optimistic.

Install / train

bash
pip install -r requirements.txt

Tuned for a single 48 GB GPU: BF16, PiSSA rank 16, seq 3072. Env vars in train_qwen_pissa.py override defaults.