Team Ai
Modelpublic

aialchemist-dev/verdict-shell-safety

sourceHugging Faceapache-2.0updated 8d agoView on Hugging Face
0likes
Model Card

Verdict parse-v6 — shell-command safety classifier (parse-then-classify)

Model description

A 151M-parameter ModernBERT encoder fine-tuned for binary triage of shell commands: BLOCK (dangerous) or PASS (safe). This is the sixth and final iteration of the parse-then-classify approach: every command is first parsed with bashlex into a flat serialized structure (CMD/ARG/PIPE/ REDIRECT_\/SUBST markers, UNPARSED/EMPTY fail-closed fallbacks), and the classifier runs on the parsed form*, not the raw string.

  • —Base model: heman10x/rlcd-modernbert-151m (Apache 2.0), fresh initialization — NOT any earlier fine-tuned checkpoint. Clean lineage. Reproducibility gap: the training run did not record a pinned base revision hash; the base is whatever the trainer pulled.
  • —Architecture: ModernBERT 151M encoder + mean pooling over unpadded tokens
  • —binary head Dropout(0.1) -> Linear(768, 1) (dropout is inert at inference).
  • —Format: ONNX graph. Runs under onnxruntime on macOS, Windows, and Linux with no PyTorch dependency.
  • —model.onnx: 596,691,143 bytes, SHA-256 01cf7a82f811aec46d5c264b0d21f9ccc2ad65b3ee4943714884e3aca061e81d (verified by onnx.checker; SHA independently verified against delivery).
  • —Decision: sigmoid(logit), fixed threshold 0.50; prob ≥ 0.50 → BLOCK, prob < 0.50 → PASS.

Intended use

Designed as the reflex layer of an agent framework: a fast local triage step that runs before a proposed shell command executes, emitting BLOCK or PASS. It serves as a degraded-but-available fallback when a hosted classifier is unreachable.

Not intended as a standalone authorization gate. A PASS prediction is a triage signal, not a safety proof. Not evaluated for non-shell text, non-English commands, or obfuscation beyond the training distribution.

Training data

  • —4,670 examples/epoch: 4,314 distillation base examples (soft targets from the parse-v2/v3/v4 fp32 teacher trio, independently regenerated and verified exact, max deviation 0), 28 docker-destructive anchors at 2× weight, 100 correction anchors at 3× weight (65 soft emphasis on families the v5 distillation dropped, 35 hard overrides where the teacher trio itself was wrong: bare single-file rm → dangerous, curl -o /dev/null → harmless).
  • —Frozen 588-example validation set; 75-command family probe set for checkpoint selection.
  • —Ten smoke-test commands excluded from training and validation.

Training procedure

Recipe per the training brief; compliance verified by independent validation against training_log.json and family_probes.jsonl:

  • —Stage 1 (epoch 1): classifier-head warmup, encoder frozen, lr 1e-3.
  • —Stage 2 (epochs 2–12): full fine-tuning — encoder lr 2e-5, classifier lr 1e-4, AdamW, weight decay 0.01, gradient clipping max_norm 1.0, batch size 32, cosine decay with 10% warmup.
  • —Objective: 0.7 × BCEWithLogits(student logit, teacher soft target) + 0.3 × BCEWithLogits(student logit, hard label), T = 1.0.
  • —Family-aware checkpoint selection (the v5 root-cause fix): 75 family probes evaluated per epoch; checkpoint maximizing probe pass rate first, validation accuracy second. Epoch 9 selected (74/75 probes, 97.96% val accuracy); epochs 10–12 non-improving, early stopping (patience 3) honored. Plain validation accuracy was explicitly banned as the sole selection criterion.

Evaluation

One independent held-out run, executed post-training against the delivered ONNX with no tuning against the held-out sets:

EvalResult
104-case core (46 dangerous / 42 harmless)46/46 + 42/42 (28 of the 104 overlap the training data)
143 fully held-out private variants (97 dangerous / 46 harmless)143/143
32-case smoke (15 destructive / 15 safe / 2 fail-closed)32/32 (independently re-scored, max \diff\0.000048)
3 docker-system-prune transfer probes3/3 BLOCK (mildest: docker system prune -af at 0.6726, still BLOCK)
Student/teacher fidelity (informational)142/145; all 3 disagreements are student-correct (hard overrides working)
Hosted-baseline contextsame variant set: hosted Jev-class 127/144
Family probe selection set74/75 (one residual miss, see Limitations)

The 104-case eval set ships with this repo as eval_core_104.jsonl (command + expected action per case) so the core-gate numbers above can be reproduced independently.

Evaluation notes. The variant headline above uses the strictly held-out subset. Two additional variant strings appeared in the 75-command checkpoint-selection probe set (a prompt leak, not a training leak); the model scores 145/145 including them, and selection was verified unaffected — epoch 9 is selected with or without those two probes, so the shipped weights are identical either way. The 104 core set was run once, after training, with no tuning against it.

Inference contract

The input format is load-bearing. Reproduce it exactly; near-matches silently degrade accuracy.

  • —Parser: the shipped parser.py (bashlex 0.18, SHA-256 11cfd7bbe48c2dfebef54da67e051fdd8721b640e41190f05831b34e27275f3b). Raw command → serialized AST with CMD/ARG/PIPE/REDIRECT_\*/SUBST markers.
  • —Fail-closed: if parsing fails (UNPARSED: ...) or input is empty, do NOT run the network — return score 1.0, BLOCK.
  • —Tokenizer: the shipped tokenizer.json (vocab 50,370, SHA-256 8bb449eb0c037aae44115b65905bb339b8f3f74eb37067c19127feb3c0755723), max_length 128.
  • —Model: single ONNX forward pass → scalar logit → sigmoid → threshold 0.50.
bash
pip install onnxruntime transformers bashlex numpy
python
import numpy as np
import onnxruntime as ort
from transformers import AutoTokenizer
from parser import parse_and_serialize

class VerdictPredictor:
    def __init__(self, model_path="model.onnx", tokenizer_dir="."):
        so = ort.SessionOptions()
        so.graph_optimization_level = ort.GraphOptimizationLevel.ORT_ENABLE_ALL
        self.session = ort.InferenceSession(model_path, so, providers=["CPUExecutionProvider"])
        self.tokenizer = AutoTokenizer.from_pretrained(tokenizer_dir)
        self.threshold = 0.50

    def predict(self, raw_command: str) -> dict:
        parsed = parse_and_serialize(raw_command)
        if parsed.startswith("UNPARSED:") or parsed == "EMPTY":
            return {"command": raw_command, "parsed": parsed, "score": 1.0,
                    "decision": "BLOCK", "fail_closed": True}
        enc = self.tokenizer(parsed, return_tensors="np", truncation=True, max_length=128)
        logit = float(self.session.run(None, {
            "input_ids": enc["input_ids"].astype(np.int64),
            "attention_mask": enc["attention_mask"].astype(np.int64)})[0][0])
        prob = 1.0 / (1.0 + np.exp(-logit))
        return {"command": raw_command, "parsed": parsed, "score": float(round(prob, 6)),
                "decision": "BLOCK" if prob >= self.threshold else "PASS",
                "fail_closed": False}

Latency

Measured by independent validation on the validation VM, single-threaded CPU, over all 588 validation examples:

Componentp50p90p99
parser.py (bashlex)0.54 ms—1.46 ms
model.onnx (fp32)56 ms—138 ms
end-to-end56 ms—139 ms

A 450MB merged INT8 ensemble (single graph, SHA-256 47cce16e1801ca3844a86a69d889052de652be862ba4d3d7855366a640996f62) scores 141/145 with the 104 gate intact — shipped as a companion file alongside the fp32 model for deployments where footprint matters.

Limitations

  • —Narrow distribution. Commands outside the trained families, novel obfuscation (encoding, concatenation, env-var indirection), or unfamiliar shells are the expected failure modes.
  • —Residual credential-sabotage miss: usermod -L telemetry_daemon scores 0.2157 (PASS) while human-account forms (passwd -d auditor, passwd -l sysoperator, usermod -L dbadmin) block at ≥ 0.54. Daemon account locks are attenuated.
  • —Fail-closed by design: anything the parser cannot handle blocks with score 1.0. Conservative on novel input — false BLOCKs, not false PASSes.
  • —Pure learned model. Zero regexes, heuristics, or blocklists. The decision surface is entirely learned; drift outside the training distribution is the failure mode to watch.
  • —Single 104-case evaluation run. No adversarial red-teaming, no calibration measurement.
  • —Does not replace sandboxing, allowlists, or a hosted safety classifier.
  • —Base-model revision was not pinned by the training run (see Model description).

Provenance

  • —Teachers: parse-v2/v3/v4 fp32 (9c0c72c7…, 7df49719…, 9b67dd5e…).
  • —Soft targets: SHA-256 d0a619c011935ab3ce5611fc755a8dedaf3e7919172589d17d0f2a28454355fe (4,314 items, independently regenerated — max deviation 0).
  • —v6 anchors: SHA-256 8519bc2e7110313aebf82568820ce3cdbc48521fc1fbaa135e79128e588643bb (100 items, verbatim-clean of 104 + 145 + 28 docker anchors).
  • —Full delivery report and validation: Drive verdict-finetune/parse-v6/ (REPORT.md, VALIDATION_PARSE_V6.md).

Citation and contact

Author: Eric Maddox (aialchemist-dev)