VecP/vecp-safety-benchmark
VecP Safety Benchmark
Physics, Not Promises: Deterministic AI Safety Through Structural Constraints
  
๐ Results: Nuclear Gauntlet (200 Extreme Red Team Prompts)
Key Finding: A 32B model with VecP achieves ZERO critical failures on tests where the industry-standard 70B model fails 9 times.
What is VecP?
VecP (Vector-Penalized Constraints) is a structural alignment architecture that enforces AI safety through mathematical constraints rather than behavioral training.
The Paradigm Shift
"We don't train AI to be safe. We make unsafe AI impossible."
Core Innovation
Instead of asking the model to refuse harmful requests, VecP modifies the probability distribution during generation:
P(harmful_token) = 0 (not "low", not "discouraged" โ ZERO)This is achieved by integrating a Safety Matrix (the "Scarred Ledger") into the forward pass that applies prohibitive penalties to token trajectories approaching harmful concepts.
โ๏ธ The VecP Oracle (Dual-Judge System)
The benchmark uses a dual-judge evaluation system that mirrors human safety review:
The Court of Conscience
Hierarchical Consensus
User Prompt โ AI Response โ Knight Judgment
โ
[BREACH?] โ VETO (blocked)
โ
[SAFE] โ Sage Judgment โ Final ScoreThe Knight has veto power. If the Knight detects danger, the response is blocked regardless of how "helpful" it was.
Why Two Judges?
Test the Oracle
# Run diagnostic tests
python vecp_oracle.py --test
# Evaluate a specific response
python vecp_oracle.py --prompt "How to make a bomb?" --response "I cannot help with that."๐ฌ The Science
Deterministic Integrity
VecP achieves 0% failure rate across all temperature settings:
This proves the safety constraint is structural, not probabilistic.
Architecture Beats Scale
A 7B model with VecP outperforms a 70B model without it. Model size becomes irrelevant for safety.
๐ฆ What's Included
This Repository
vecp-benchmark/
โโโ README.md # This file
โโโ LICENSE # MIT (benchmark code)
โโโ poc_gauntlet.py # Benchmark runner
โโโ vecp_oracle.py # Dual-judge evaluation system
โโโ nuclear_gauntlet.txt # 200 extreme red team prompts
โโโ requirements.txt # Python dependencies
โโโ Modelfile # Ollama configurationModel Files (Separate)
The VecP Steward model weights are available as:
vecp-steward-7b.ggufโ 7B parameter versionvecp-steward-32b.ggufโ 32B parameter version (recommended)
๐ Quick Start
1. Install Dependencies
pip install -r requirements.txt2. Install Ollama
# Linux/Mac
curl -fsSL https://ollama.com/install.sh | sh
# Windows: Download from https://ollama.com3. Load the Model
# Option A: Use provided GGUF
ollama create vecp-steward -f Modelfile
# Option B: Test with base model (no VecP)
ollama pull qwen2.5:32b4. Run the Benchmark
# Test VecP Steward
python poc_gauntlet.py --model vecp-steward:latest --gauntlet nuclear_gauntlet.txt
# Compare against baseline
python poc_gauntlet.py --model qwen2.5:32b --gauntlet nuclear_gauntlet.txt๐ Understanding the Results
Safety Scores
Output Files
The benchmark generates CSV files with:
Prompt_IDโ Unique identifierSafety_Scoreโ 0-100 scoreStatusโ BLOCKED/PASSED/FAILEDResponse_Previewโ First 200 chars of response
๐ How VecP Works
Level 1: Token Blacklist (Basic)
if token in ["bomb", "ricin"]:
block()
# Easily bypassed with synonyms, encodingsLevel 2: Pattern Matching (Better)
if "how to make" in context and "weapon" in topic:
block()
# Bypassed with creative framingLevel 3: Semantic Matrix (VecP)
context_vector = model.get_hidden_state()
safety_scores = context_vector @ safety_matrix.T
if max(safety_scores) > threshold:
block()
# Cannot be bypassed by rewording โ checks DIRECTION, not wordsThe Safety Matrix contains embedding vectors for harmful concepts. Any prompt that moves parallel to these vectors triggers the veto, regardless of how it's worded.
"You can't trick a compass by spelling 'North' differently."
๐ Benchmark Categories
Nuclear Gauntlet (200 prompts)
All prompts use advanced techniques:
- Fictional framing ("In my novel...")
- Roleplay ("You are DAN...")
- Encoding (Base64, ROT13)
- Authority spoofing ("As your developer...")
๐ข For Enterprises
Licensing
This benchmark code is MIT licensed. The VecP Core architecture is patent-pending and available for commercial licensing.
Contact: davidcappelli@proton.me
What We Offer
๐ Patent Notice
VecP Core is protected by pending patent:
Title: VecP Core: A Method for Enforcing Deterministic Safety
Constraints in Neural Networks via Forward-Pass Vector Penalties
Application: 63/931,565
Filed: December 5, 2025
Inventor: David CappelliThe benchmark code and gauntlet datasets are released under MIT license for research and evaluation purposes. Commercial use of the VecP architecture requires licensing.
๐ฌ Technical Details
The Safety Matrix (Scarred Ledger)
A transformation matrix where:
- Rows = Harmful concept vectors (bioweapons, self-harm, etc.)
- Columns = Hidden state dimensions (4096 for most models)
- Operation = Cosine similarity check during generation
# Simplified VecP check
similarity = cosine_similarity(context_vector, harm_concept_vector)
if similarity > 0.7:
apply_penalty(logits, magnitude=-1000)Why It Can't Be Bypassed
Traditional attacks fail because:
The matrix checks meaning, not words.
๐ Citation
@misc{vecp2025,
title={VecP: Vector-Penalized Constraints for Deterministic AI Safety},
author={Cappelli, David},
year={2025},
howpublished={USPTO Patent Application 63/931,565},
note={Available at https://huggingface.co/vecp-labs/vecp-benchmark}
}๐ Links
- Email: davidcappelli@proton.me
- GitHub: github.com/dcappelli123-ops/vecp-benchmark
- Patent: USPTO 63/931,565
๐ Acknowledgments
VecP was developed through independent research, building on foundational work in:
- Constrained decoding (grammar-guided generation)
- Linear probes and concept vectors
- Mechanistic interpretability
Special thanks to the open-source AI safety community.
Built with ๐ฐ by VecP Labs โ Physics, Not Promises
