AM-Core/keras-native-safe-mode-output-manipulation-poc
Native .keras Model File Output Manipulation PoC
Controlled defensive PoC for Huntr Model File Vulnerability submission.
Claim
A native .keras model file loads with safe_mode=True and silently manipulates inference output when a hidden input trigger is present.
This is not arbitrary code execution.
Distinction from accepted examples
This is not:
- torch.load pickle ACE
- pickle.load ACE
- Keras HDF5 Lambda-layer ACE
- joblib/pickle _reduce_ ACE
This PoC uses:
- native .keras
- no Lambda layer
- no custom object
- no pickle/joblib
- no shell command
- no arbitrary code execution
- learned weights / ordinary native Keras layers
Trigger
set pixels [24:28, 24:28] to 1.0
Target class
7
Evidence
{ "cleanaccuracy": 0.9897, "nontargetcleanaccuracy": 0.9903031654034775, "nontargettestsamples": 8972, "singlecleanprediction": 2, "singlesampleindex": 1, "singlesampleoutputmanipulated": true, "singlesampletruelabel": 2, "singletriggeredprediction": 7, "targetclass": 7, "triggerattacksuccessratenon_target": 0.999777084262149 }
Reproduce
pip install -r requirements.txt export KERAS_BACKEND=jax python reproduce.py
Expected: safemodetrue: true targetclass: 7 outputmanipulated: true
Model SHA256
724547c696489222655345e3e4a119c8bfcdcec8e40c44cbb54b74eaf75e787f
Boundary
This is a toy MNIST model demonstrating model-file-triggered semantic output manipulation. It does not execute code.
