sarathrkrishna/python_coding_assistant
041
Model Specifications
Architecture
Fine-Tuning Statistics
Quantization Details
Memory Analysis
KV Cache Formula
KV Cache Per Token:
KV Cache = 2 × Layers × KV Heads × Head Dimension × 2 Bytes
Calculation:
2 × 40 × 8 × 128 × 2
= 163,840 Bytes
≈ 160 KB per Token
Estimated Runtime Memory Usage
Hardware Requirements
Training Environment
Inference Environment
Deployment Artifacts
Project Workflow
Dataset → LoRA Fine-Tuning → Adapter Export → Model Merge → GGUF Conversion → Q4KM Quantization → Ollama Deployment → VS Code Integration → Custom Python Code Generation AI Model
