Team Ai
Modelpublic

sarathrkrishna/python_coding_assistant

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
0likes41downloads
Model Card

Model Specifications

Architecture

PropertyValue
Base ModelIBM Granite 4.1
Architecture TypeDecoder-Only Transformer
Transformer Layers40
Hidden Size4096
Attention Heads32
KV Heads (GQA)8
Head Dimension128
Intermediate Size12800
Vocabulary Size100,352
Context Length131,072 Tokens
Activation FunctionSiLU
RoPE Theta10,000,000

Fine-Tuning Statistics

MetricValue
Fine-Tuning MethodLoRA
Trainable Parameters98,959,360
Total Parameters4,494,921,728
Trainable Percentage2.20%
Base Parameters Frozen97.80%
Training FrameworkUnsloth
OptimizerAdamW 8-bit

Quantization Details

PropertyValue
Output FormatGGUF
Quantization MethodQ4KM
Quantization TypeK-Quant Medium
Deployment Size~5 GB
Runtime EngineOllama / llama.cpp

Memory Analysis

KV Cache Formula

KV Cache Per Token:

KV Cache = 2 × Layers × KV Heads × Head Dimension × 2 Bytes

Calculation:

2 × 40 × 8 × 128 × 2

= 163,840 Bytes

≈ 160 KB per Token


Estimated Runtime Memory Usage

Context LengthKV CacheTotal Runtime Memory
4K Tokens~655 MB~6.2 GB
8K Tokens~1.31 GB~7.0 GB
16K Tokens~2.62 GB~8–9 GB
32K Tokens~5.24 GB~11 GB

Hardware Requirements

Training Environment

ComponentValue
GPUNVIDIA T4
VRAM16 GB
Quantization4-bit NF4
Fine-Tuning MethodLoRA

Inference Environment

ComponentValue
GPURTX 2080
VRAM8 GB
System RAM32 GB
Recommended Context8192 Tokens
QuantizationQ4KM

Deployment Artifacts

ArtifactPurpose
granitepythonlora.zipLoRA Adapter Backup
adapter_model.safetensorsFine-Tuned Weights
granite-4.1-8b.Q4KM.ggufDeployable Model
ModelfileOllama Configuration
granite-pythonOllama Model Name

Project Workflow

Dataset → LoRA Fine-Tuning → Adapter Export → Model Merge → GGUF Conversion → Q4KM Quantization → Ollama Deployment → VS Code Integration → Custom Python Code Generation AI Model