genos-security/genos
Genos: command-line threat triage models
Genos classifies a single shell or script command line (Linux, macOS, Windows, PowerShell) in two tiers. This repository holds the runtime artifacts used by the genos_api service.
command ─► (Base64/obfuscation decode) ─► Tier 1 gatekeeper ─► Benign / Context_Dependent / Malicious
│
└─► Tier 2: family specialist (11 tactic families)
+ behavior encoder (attack stage + action tags)Benign commands skip Tier 2 by default. Every output is a model estimate, not a verdict to act on automatically.
Files
The CodeBERT backbone (microsoft/codebert-base) is not included; download it separately (see Usage).
Models
Tier 1: gatekeeper (gatekeeper.pt)
- Architecture: shared CodeBERT encoder with decision-decomposition heads (
verdict_logits,non_benign_logit,malicious_given_non_benign_logit,ordinal_risk_logit). - Classes:
Benign(0),Malicious(1),Context_Dependent(2).Context_Dependentmarks commands that cannot be judged benign or malicious without environmental context. - Training: 5 epochs (best epoch 5), max length 256, learning rate 1e-5, weight decay 0.01, label smoothing 0.05, effective batch size 256, balanced sampling, seed 42, plus a 250-row benign operational patch.
- Reported metrics (from
config/gatekeeper_meta.jsonin the source repo):
Test per class: Benign P 0.968 / R 0.957 / F1 0.962; Malicious P 0.944 / R 0.970 / F1 0.957; Context_Dependent P 0.880 / R 0.871 / F1 0.876.
These numbers come from the original split used when this checkpoint was trained. The project's later dataset audit found conflicting and development-exposed rows in earlier splits, so these metrics should be treated as optimistic and not as evidence of production accuracy.
Tier 2: family specialist (family_specialist_tfidf.joblib)
- Architecture: TF-IDF features (character
char_wb2-5-grams, 250k features; word 1-2-grams with a shell-aware token pattern, 100k features) feeding one linear SVM per family, with per-family sigmoid calibration fit on train only using template-group-isolated folds. - Output: multi-label scores over 11 families: Execution, Persistence, Privilege Escalation, Defense Evasion, Credential Access, Discovery, Lateral Movement, Command-and-Control / Payload Retrieval, Exfiltration, Impact, Benign Admin. Decision threshold 0.5. No MITRE technique IDs are predicted.
- Data: 37,846 train / 4,740 validation / 4,740 test commands, split 80/10/10 by parser residual-template group (seed 42) so normalized commands and templates never cross splits. The split audit found zero command overlap between splits.
- Speed: about 2.9 ms median (3.1 ms p95) per command on CPU, excluding HTTP and the other models.
Test-set results (threshold 0.5):
Macro F1 is 0.570 on test (0.604 on validation). The top-1 prediction matches at least one true family 93.7% of the time (top-2: 96.0%). The weakest families (Lateral Movement, Exfiltration, Impact) have very few test examples, so their scores are noisy. Precision is generally higher than recall: the model tends to miss attack families rather than invent them, and its most common error is predicting Benign Admin for a Defense Evasion or Discovery command.
Tier 2: behavior encoder (behavior_encoder.pt)
- Architecture: CodeBERT encoder with a stage classifier and a multi-label action head, max length 256.
- Stages (13): C2 / Remote Access, Collection / Staging, Context Required, Credential Access, Defense Evasion, Discovery / Recon, Execution, Exfiltration, Impact, Lateral Movement, Payload Retrieval, Persistence, Privilege Escalation.
- Action tags (9): archivedata, downloadremoteresource, executeinlinecode, executeinterpreter, extractarchive, remoteexecution, useencodedpayload, useobfuscation, usesignedproxybinary.
- Reported test metrics: stage accuracy 0.565, stage macro F1 0.500, action micro F1 0.936.
Action labels are derived from the same parser and rule features used for input, so the high action F1 shows the model reproduces those rules. It is not independent evidence of behavioral understanding. Stage accuracy of 0.565 over 13 classes is modest.
Usage
Download the files into the application's models/ directory:
hf download genos-security/genos --local-dir models
hf download microsoft/codebert-base --local-dir models/codebert-baseThen point the service at them:
GENOS_CODEBERT_PATH=models/codebert-base
GENOS_HF_LOCAL_ONLY=1
GENOS_SPECIALIST_MODE=family
GENOS_FAMILY_SPECIALIST_PATH=models/family_specialist_tfidf.joblibThe .pt files expect the genos_api code, label maps in config/, and the CodeBERT tokenizer. They are not drop-in transformers checkpoints.
Intended use
- Triage and prioritization of suspicious command lines in defensive tooling, research, and education.
- Use as one signal alongside other telemetry, with a human in the loop.
Out-of-scope use
- Automatically blocking, killing, or attributing activity from a score alone.
- Judging full scripts, binaries, or multi-command sessions. Inputs are single command lines truncated to 256 tokens.
- Offensive use, such as tuning commands to evade the models.
- Claims of production accuracy.
Limitations and known issues
- Weak labels. Training labels come from source-derived tactic families and parser rules, not human review. Independent annotation is not complete.
- Local evaluation only. Results are from template-grouped splits of the same data distribution. Generalization to new environments, attacker tooling, or novel obfuscation has not been measured.
- Cross-component overlap. Some test commands for one component appear in another component's training data, so the tiers are not evaluated end to end on untouched data.
- Uncalibrated gatekeeper. Scores are estimates, not probabilities, unless a validation-fitted calibration artifact is supplied.
- Known false positive. The gatekeeper has scored the benign command
pwdas Malicious (about 52%, uncalibrated), and the behavior model labeled it Persistence. Expect other false positives on very short commands. - Low-support families. Lateral Movement, Exfiltration, and Impact are recalled poorly (recall 0.11-0.38).
- English, command-line text only.
- Training data traces. The TF-IDF vocabulary is built from training commands and may contain fragments of them.
Security note
.pt and .joblib files are Python pickles. Loading one runs code embedded in the file. Only load these artifacts from this repository's verified revisions, and check file hashes before deployment. The specialist records its own SHA-256 in family_specialist_tfidf.json (checkpoint_sha256).
Citation and license
No license has been declared for these weights yet. Contact the maintainers before redistribution or commercial use. The base model, `microsoft/codebert-base`, is released under the MIT license.
