Team Ai
Modelpublic

sarathrkrishna/python_coding_assistant

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
0likes41downloads
Document.md304 linesDownload Raw Back to root
1Dataset = "https://huggingface.co/sarathrkrishna/python_coding_assistant/tree/main"2 3# Fine-Tuning and Deploying a Specialized Python Code Generation AI Model using IBM Granite 4.14 5## Project Overview6 7This project focuses on fine-tuning IBM Granite 4.1 to create a specialized Python code generation AI model capable of generating and assisting with:8 9* FastAPI Development10* Agentic AI Systems11* LangGraph Workflows12* LangChain Applications13* Machine Learning Pipelines14* MLOps Solutions15* General Python Software Engineering Tasks16 17The model was fine-tuned using LoRA, optimized with Unsloth, quantized to GGUF format, deployed using Ollama, and integrated into VS Code as a custom coding agent.18 19---20 21# Base Model22 23Model:24 25IBM Granite 4.126 27Architecture:28 29* 40 Transformer Layers30* 32 Attention Heads31* 8 KV Heads (Grouped Query Attention)32* Head Dimension: 12833* Hidden Size: 409634* Context Length: 131,072 Tokens35 36---37 38# Dataset39 40Training Dataset:41 42Python Code Generation Dataset43 44Examples Included:45 46* Python Functions47* FastAPI APIs48* Object-Oriented Programming49* Data Structures & Algorithms50* Backend Services51* AI Agent Workflows52* Machine Learning Code53* Utility Scripts54 55Dataset Format:56 57```json58{59  "instruction": "Create a FastAPI endpoint for uploading PDF files",60  "response": "Generated Python code..."61}62```63 64---65 66# Fine-Tuning67 68Framework:69 70Unsloth71 72Training Method:73 74LoRA (Low-Rank Adaptation)75 76Benefits:77 78* Lower GPU Memory Usage79* Faster Training80* Parameter Efficient Fine-Tuning81 82Training Statistics:83 84* Trainable Parameters: 98,959,36085* Total Parameters: 4,494,921,72886* Trainable Percentage: 2.2%87 88Training Pipeline:89 90Dataset91→ Formatting92→ Tokenization93→ LoRA Injection94→ SFT Training95→ Adapter Export96 97---98 99# Saving LoRA Adapters100 101```python102model.save_pretrained("granite_python_lora")103tokenizer.save_pretrained("granite_python_lora")104```105 106Generated Files:107 108```text109granite_python_lora/110├── adapter_model.safetensors111├── adapter_config.json112├── tokenizer.json113├── tokenizer_config.json114├── vocab.json115└── merges.txt116```117 118Purpose:119 120* Reproduce Fine-Tuning121* Continue Training122* Upload to Hugging Face123* Rebuild GGUF Models124 125---126 127# GGUF Conversion128 129Convert the fine-tuned model to GGUF format:130 131```python132model.save_pretrained_gguf(133    "granite_finetune_q4",134    tokenizer,135    quantization_method="q4_k_m"136)137```138 139Generated Output:140 141```text142granite_finetune_q4_gguf/143└── granite-4.1-8b.Q4_K_M.gguf144```145 146---147 148# Quantization149 150Selected Quantization:151 152Q4_K_M153 154Advantages:155 156* Small Model Size157* Fast Inference158* Low Memory Usage159* Suitable for Consumer GPUs160 161Deployment Size:162 163~5 GB164 165---166 167# Ollama Deployment168 169## Step 1: Create Modelfile170 171Create a file named:172 173```text174Modelfile175```176 177Contents:178 179```text180FROM ./granite-4.1-8b.Q4_K_M.gguf181 182PARAMETER temperature 0.1183PARAMETER num_ctx 8192184 185SYSTEM """186You are a specialized Python code generation AI model.187 188Expertise:189- Python190- FastAPI191- LangGraph192- LangChain193- AI Agents194- Machine Learning195- MLOps196 197Always generate production-ready code with:198- Type Hints199- Error Handling200- Clean Architecture201- Best Practices202"""203```204 205---206 207## Step 2: Create Ollama Model208 209```bash210ollama create granite-python -f Modelfile211```212 213Expected Output:214 215```text216gathering model components217copying file218parsing GGUF219writing manifest220success221```222 223---224 225## Step 3: Verify Model226 227```bash228ollama list229```230 231Expected:232 233```text234NAME235granite-python236```237 238---239 240## Step 4: Run Model241 242```bash243ollama run granite-python244```245 246Example Prompt:247 248```text249Create a production-ready FastAPI file upload service.250```251 252---253 254# VS Code Integration255 256The Ollama model can be integrated into:257 258* Continue.dev259* Roo Code260* Cline261* Open WebUI262* Custom VS Code AI Workflows263 264Result:265 266A locally running Python code generation AI assistant capable of helping with:267 268* FastAPI269* LangGraph270* LangChain271* Agentic AI272* Machine Learning273* MLOps274 275without relying on cloud-hosted LLM APIs.276 277---278 279# Project Workflow280 281Dataset282↓283LoRA Fine-Tuning284↓285Save LoRA Adapters286↓287Merge Model288↓289GGUF Conversion290↓291Q4_K_M Quantization292↓293Ollama Deployment294↓295VS Code Integration296↓297Custom Python AI Coding Agent298 299---300 301# Final Outcome302 303A specialized fine-tuned Python code generation AI model built using IBM Granite 4.1, optimized for modern Python development workflows and deployed locally using Ollama and VS Code.304