aialchemist-dev/verdict-shell-safety
Verdict parse-v6 — shell-command safety classifier (parse-then-classify)
Model description
A 151M-parameter ModernBERT encoder fine-tuned for binary triage of shell commands: BLOCK (dangerous) or PASS (safe). This is the sixth and final iteration of the parse-then-classify approach: every command is first parsed with bashlex into a flat serialized structure (CMD/ARG/PIPE/ REDIRECT_\/SUBST markers, UNPARSED/EMPTY fail-closed fallbacks), and the classifier runs on the parsed form*, not the raw string.
- Base model:
heman10x/rlcd-modernbert-151m(Apache 2.0), fresh initialization — NOT any earlier fine-tuned checkpoint. Clean lineage. Reproducibility gap: the training run did not record a pinned base revision hash; the base is whatever the trainer pulled. - Architecture: ModernBERT 151M encoder + mean pooling over unpadded tokens
- binary head
Dropout(0.1) -> Linear(768, 1)(dropout is inert at inference). - Format: ONNX graph. Runs under onnxruntime on macOS, Windows, and Linux with no PyTorch dependency.
model.onnx: 596,691,143 bytes, SHA-25601cf7a82f811aec46d5c264b0d21f9ccc2ad65b3ee4943714884e3aca061e81d(verified byonnx.checker; SHA independently verified against delivery).- Decision: sigmoid(logit), fixed threshold 0.50; prob ≥ 0.50 → BLOCK, prob < 0.50 → PASS.
Intended use
Designed as the reflex layer of an agent framework: a fast local triage step that runs before a proposed shell command executes, emitting BLOCK or PASS. It serves as a degraded-but-available fallback when a hosted classifier is unreachable.
Not intended as a standalone authorization gate. A PASS prediction is a triage signal, not a safety proof. Not evaluated for non-shell text, non-English commands, or obfuscation beyond the training distribution.
Training data
- 4,670 examples/epoch: 4,314 distillation base examples (soft targets from the parse-v2/v3/v4 fp32 teacher trio, independently regenerated and verified exact, max deviation 0), 28 docker-destructive anchors at 2× weight, 100 correction anchors at 3× weight (65 soft emphasis on families the v5 distillation dropped, 35 hard overrides where the teacher trio itself was wrong: bare single-file
rm→ dangerous,curl -o /dev/null→ harmless). - Frozen 588-example validation set; 75-command family probe set for checkpoint selection.
- Ten smoke-test commands excluded from training and validation.
Training procedure
Recipe per the training brief; compliance verified by independent validation against training_log.json and family_probes.jsonl:
- Stage 1 (epoch 1): classifier-head warmup, encoder frozen, lr 1e-3.
- Stage 2 (epochs 2–12): full fine-tuning — encoder lr 2e-5, classifier lr 1e-4, AdamW, weight decay 0.01, gradient clipping max_norm 1.0, batch size 32, cosine decay with 10% warmup.
- Objective: 0.7 × BCEWithLogits(student logit, teacher soft target) + 0.3 × BCEWithLogits(student logit, hard label), T = 1.0.
- Family-aware checkpoint selection (the v5 root-cause fix): 75 family probes evaluated per epoch; checkpoint maximizing probe pass rate first, validation accuracy second. Epoch 9 selected (74/75 probes, 97.96% val accuracy); epochs 10–12 non-improving, early stopping (patience 3) honored. Plain validation accuracy was explicitly banned as the sole selection criterion.
Evaluation
One independent held-out run, executed post-training against the delivered ONNX with no tuning against the held-out sets:
The 104-case eval set ships with this repo as eval_core_104.jsonl (command + expected action per case) so the core-gate numbers above can be reproduced independently.
Evaluation notes. The variant headline above uses the strictly held-out subset. Two additional variant strings appeared in the 75-command checkpoint-selection probe set (a prompt leak, not a training leak); the model scores 145/145 including them, and selection was verified unaffected — epoch 9 is selected with or without those two probes, so the shipped weights are identical either way. The 104 core set was run once, after training, with no tuning against it.
Inference contract
The input format is load-bearing. Reproduce it exactly; near-matches silently degrade accuracy.
- Parser: the shipped
parser.py(bashlex 0.18, SHA-25611cfd7bbe48c2dfebef54da67e051fdd8721b640e41190f05831b34e27275f3b). Raw command → serialized AST with CMD/ARG/PIPE/REDIRECT_\*/SUBST markers. - Fail-closed: if parsing fails (
UNPARSED: ...) or input is empty, do NOT run the network — return score 1.0, BLOCK. - Tokenizer: the shipped
tokenizer.json(vocab 50,370, SHA-2568bb449eb0c037aae44115b65905bb339b8f3f74eb37067c19127feb3c0755723), max_length 128. - Model: single ONNX forward pass → scalar logit → sigmoid → threshold 0.50.
pip install onnxruntime transformers bashlex numpyimport numpy as np
import onnxruntime as ort
from transformers import AutoTokenizer
from parser import parse_and_serialize
class VerdictPredictor:
def __init__(self, model_path="model.onnx", tokenizer_dir="."):
so = ort.SessionOptions()
so.graph_optimization_level = ort.GraphOptimizationLevel.ORT_ENABLE_ALL
self.session = ort.InferenceSession(model_path, so, providers=["CPUExecutionProvider"])
self.tokenizer = AutoTokenizer.from_pretrained(tokenizer_dir)
self.threshold = 0.50
def predict(self, raw_command: str) -> dict:
parsed = parse_and_serialize(raw_command)
if parsed.startswith("UNPARSED:") or parsed == "EMPTY":
return {"command": raw_command, "parsed": parsed, "score": 1.0,
"decision": "BLOCK", "fail_closed": True}
enc = self.tokenizer(parsed, return_tensors="np", truncation=True, max_length=128)
logit = float(self.session.run(None, {
"input_ids": enc["input_ids"].astype(np.int64),
"attention_mask": enc["attention_mask"].astype(np.int64)})[0][0])
prob = 1.0 / (1.0 + np.exp(-logit))
return {"command": raw_command, "parsed": parsed, "score": float(round(prob, 6)),
"decision": "BLOCK" if prob >= self.threshold else "PASS",
"fail_closed": False}Latency
Measured by independent validation on the validation VM, single-threaded CPU, over all 588 validation examples:
A 450MB merged INT8 ensemble (single graph, SHA-256 47cce16e1801ca3844a86a69d889052de652be862ba4d3d7855366a640996f62) scores 141/145 with the 104 gate intact — shipped as a companion file alongside the fp32 model for deployments where footprint matters.
Limitations
- Narrow distribution. Commands outside the trained families, novel obfuscation (encoding, concatenation, env-var indirection), or unfamiliar shells are the expected failure modes.
- Residual credential-sabotage miss:
usermod -L telemetry_daemonscores 0.2157 (PASS) while human-account forms (passwd -d auditor,passwd -l sysoperator,usermod -L dbadmin) block at ≥ 0.54. Daemon account locks are attenuated. - Fail-closed by design: anything the parser cannot handle blocks with score 1.0. Conservative on novel input — false BLOCKs, not false PASSes.
- Pure learned model. Zero regexes, heuristics, or blocklists. The decision surface is entirely learned; drift outside the training distribution is the failure mode to watch.
- Single 104-case evaluation run. No adversarial red-teaming, no calibration measurement.
- Does not replace sandboxing, allowlists, or a hosted safety classifier.
- Base-model revision was not pinned by the training run (see Model description).
Provenance
- Teachers: parse-v2/v3/v4 fp32 (
9c0c72c7…,7df49719…,9b67dd5e…). - Soft targets: SHA-256
d0a619c011935ab3ce5611fc755a8dedaf3e7919172589d17d0f2a28454355fe(4,314 items, independently regenerated — max deviation 0). - v6 anchors: SHA-256
8519bc2e7110313aebf82568820ce3cdbc48521fc1fbaa135e79128e588643bb(100 items, verbatim-clean of 104 + 145 + 28 docker anchors). - Full delivery report and validation: Drive
verdict-finetune/parse-v6/(REPORT.md,VALIDATION_PARSE_V6.md).
Citation and contact
Author: Eric Maddox (aialchemist-dev)
