cantina-security/apex-flash-1
apex-flash-1
apex-flash-1 is Cantina Security's first open-weights security model, developed in partnership with Yeta (@yetalabs on X). It is a reinforcement learning post-train of GLM-5.3-Flash for focused investigations: reading code, using tools, pursuing an exploit, and verifying its effect in a running target.
Held-out case evaluation
We evaluated 60 tasks from 20 held-out vulnerability cases. Each case has guided whitebox, focused whitebox, and focused blackbox views. Targets ran in isolated environments, and verifiers checked the final target state. The table reports adjudicated first-draw pass@1.
Capabilities
The checkpoint retains the base model's image-text-to-text architecture. Our reported evaluation covers text-based security tasks; we have not evaluated image or video performance.
Training and use
Training used production-like software and protocol environments with the Codex agent harness. The model is intended as a focused worker under a larger agent's direction. We recommend the Codex harness for this checkpoint.
What comes next
We are building harder, multi-step investigations as the model improves, broadening the data mix, and training across agent harnesses. We plan to publish more held-out and public benchmark results as they are validated. Explore Apex to see how this work supports production security.
Related links
- Apex Flash release post
- apex-flash-1-abliterated, an experimental derivative with modified refusal behavior. The evaluation above applies to the standard checkpoint.
