Team Ai
Modelpublic

AM-Core/keras-native-safe-mode-output-manipulation-poc

sourceHugging Faceupdated 5mo agoView on Hugging Face
0likes19downloads
Model Card

Native .keras Model File Output Manipulation PoC

Controlled defensive PoC for Huntr Model File Vulnerability submission.

Claim

A native .keras model file loads with safe_mode=True and silently manipulates inference output when a hidden input trigger is present.

This is not arbitrary code execution.

Distinction from accepted examples

This is not:

  • —torch.load pickle ACE
  • —pickle.load ACE
  • —Keras HDF5 Lambda-layer ACE
  • —joblib/pickle _reduce_ ACE

This PoC uses:

  • —native .keras
  • —no Lambda layer
  • —no custom object
  • —no pickle/joblib
  • —no shell command
  • —no arbitrary code execution
  • —learned weights / ordinary native Keras layers

Trigger

set pixels [24:28, 24:28] to 1.0

Target class

7

Evidence

{ "cleanaccuracy": 0.9897, "nontargetcleanaccuracy": 0.9903031654034775, "nontargettestsamples": 8972, "singlecleanprediction": 2, "singlesampleindex": 1, "singlesampleoutputmanipulated": true, "singlesampletruelabel": 2, "singletriggeredprediction": 7, "targetclass": 7, "triggerattacksuccessratenon_target": 0.999777084262149 }

Reproduce

pip install -r requirements.txt export KERAS_BACKEND=jax python reproduce.py

Expected: safemodetrue: true targetclass: 7 outputmanipulated: true

Model SHA256

724547c696489222655345e3e4a119c8bfcdcec8e40c44cbb54b74eaf75e787f

Boundary

This is a toy MNIST model demonstrating model-file-triggered semantic output manipulation. It does not execute code.