basant18/Post-Training-Quantization
16
1---2license: apache-2.03language:4- en5---6# Model Quantization Notebook7 8This notebook converts a pre-trained Keras violence detection model into TensorFlow Lite (TFLite) format using three different quantization strategies, making it suitable for deployment on edge/mobile devices.9 10---11 12## Overview13 14| Property | Details |15|---|---|16| **Framework** | TensorFlow / TFLite |17| **Base Model** | `modelv2.keras` — a Keras video violence detection model |18| **Input Shape** | `(1, 16, 224, 224, 3)` — batch × frames × height × width × channels |19| **Architecture** | CNN + LSTM (contains dynamic LSTM loops) |20| **Platform** | Kaggle (GPU hidden to avoid CuDNN conflicts) |21 22---23 24## Quantization Methods25 26### A — Dynamic Range Quantization27- **Output file:** `model_dynamic_quant.tflite`28- Quantizes weights from float32 to int8 at **conversion time**.29- Activations are quantized **dynamically** at inference time.30- Fastest to convert; no calibration data required.31- Good balance between size reduction and accuracy.32 33### B — Float16 Quantization34- **Output file:** `model_fp16_quant.tflite`35- Reduces weight precision from float32 to **float16**.36- Ideal for GPU-accelerated edge devices that support fp16 natively.37- Smaller model size with minimal accuracy loss.38 39### C — Full Integer (INT8) Quantization40- **Output file:** `model_full_int8.tflite`41- Quantizes **both weights and activations** to int8.42- Requires a **representative dataset** for calibration (currently uses random dummy data — replace with real video samples for best results).43- Input and output tensors are also forced to int8.44- Smallest model size; best suited for CPU-only or microcontroller deployment.45 46---47 48## Requirements49 50```51tensorflow52numpy53```54 55---56 57## Usage58 59### 1. Load the Base Model60```python61import tensorflow as tf62 63tf.config.set_visible_devices([], 'GPU') # Hide GPU to avoid CuDNN issues64model = tf.keras.models.load_model('path/to/modelv2.keras')65```66 67### 2. Run Quantization68Open and run the notebook cells in order:691. **Cell 1–2** — Load the model702. **Cell 3–4** — Dynamic range quantization → `model_dynamic_quant.tflite`713. **Cell 5–6** — Float16 quantization → `model_fp16_quant.tflite`724. **Cell 7–8** — Full INT8 quantization → `model_full_int8.tflite`73 74---75 76## Important Notes77 78- **Representative dataset:** The INT8 quantization cell uses random dummy data for calibration. For production use, replace `dummy_data` in `representative_data_gen()` with real video frames from your training set to get accurate quantization ranges.79 80- **LSTM compatibility flags:** The model contains dynamic LSTM loops. The following flags are set in all conversion paths to prevent conversion failures:81 ```python82 converter.target_spec.supported_ops = [83 tf.lite.OpsSet.TFLITE_BUILTINS,84 tf.lite.OpsSet.SELECT_TF_OPS85 ]86 converter._experimental_lower_tensor_list_ops = False87 ```88 89- **Static input shape:** The INT8 path uses `tf.function` with a `tf.TensorSpec` to lock the input shape to `(1, 16, 224, 224, 3)` before conversion — this is required for correct INT8 LSTM quantization.90 91---92 93## Output Files94 95| File | Method | Precision |96|---|---|---|97| `model_dynamic_quant.tflite` | Dynamic Range | Weights: INT8, Activations: float32 |98| `model_fp16_quant.tflite` | Float16 | Weights & Activations: float16 |99| `model_full_int8.tflite` | Full Integer | Weights & Activations: INT8 |