sarathrkrishna/python_coding_assistant
041
1---2license: apache-2.03---4# Model Specifications5 6## Architecture7 8| Property | Value |9| ------------------- | ------------------------ |10| Base Model | IBM Granite 4.1 |11| Architecture Type | Decoder-Only Transformer |12| Transformer Layers | 40 |13| Hidden Size | 4096 |14| Attention Heads | 32 |15| KV Heads (GQA) | 8 |16| Head Dimension | 128 |17| Intermediate Size | 12800 |18| Vocabulary Size | 100,352 |19| Context Length | 131,072 Tokens |20| Activation Function | SiLU |21| RoPE Theta | 10,000,000 |22 23---24 25## Fine-Tuning Statistics26 27| Metric | Value |28| ---------------------- | ------------- |29| Fine-Tuning Method | LoRA |30| Trainable Parameters | 98,959,360 |31| Total Parameters | 4,494,921,728 |32| Trainable Percentage | 2.20% |33| Base Parameters Frozen | 97.80% |34| Training Framework | Unsloth |35| Optimizer | AdamW 8-bit |36 37---38 39## Quantization Details40 41| Property | Value |42| ------------------- | ------------------ |43| Output Format | GGUF |44| Quantization Method | Q4_K_M |45| Quantization Type | K-Quant Medium |46| Deployment Size | ~5 GB |47| Runtime Engine | Ollama / llama.cpp |48 49---50 51## Memory Analysis52 53### KV Cache Formula54 55KV Cache Per Token:56 57KV Cache = 2 × Layers × KV Heads × Head Dimension × 2 Bytes58 59Calculation:60 612 × 40 × 8 × 128 × 262 63= 163,840 Bytes64 65≈ 160 KB per Token66 67---68 69## Estimated Runtime Memory Usage70 71| Context Length | KV Cache | Total Runtime Memory |72| -------------- | -------- | -------------------- |73| 4K Tokens | ~655 MB | ~6.2 GB |74| 8K Tokens | ~1.31 GB | ~7.0 GB |75| 16K Tokens | ~2.62 GB | ~8–9 GB |76| 32K Tokens | ~5.24 GB | ~11 GB |77 78---79 80## Hardware Requirements81 82### Training Environment83 84| Component | Value |85| ------------------ | --------- |86| GPU | NVIDIA T4 |87| VRAM | 16 GB |88| Quantization | 4-bit NF4 |89| Fine-Tuning Method | LoRA |90 91---92 93### Inference Environment94 95| Component | Value |96| ------------------- | ----------- |97| GPU | RTX 2080 |98| VRAM | 8 GB |99| System RAM | 32 GB |100| Recommended Context | 8192 Tokens |101| Quantization | Q4_K_M |102 103---104 105## Deployment Artifacts106 107| Artifact | Purpose |108| -------------------------- | -------------------- |109| granite_python_lora.zip | LoRA Adapter Backup |110| adapter_model.safetensors | Fine-Tuned Weights |111| granite-4.1-8b.Q4_K_M.gguf | Deployable Model |112| Modelfile | Ollama Configuration |113| granite-python | Ollama Model Name |114 115---116 117## Project Workflow118 119Dataset120→ LoRA Fine-Tuning121→ Adapter Export122→ Model Merge123→ GGUF Conversion124→ Q4_K_M Quantization125→ Ollama Deployment126→ VS Code Integration127→ Custom Python Code Generation AI Model128 