iBotIA/Qwen3.8-4B-Empero-AI-FullStack-GGUF
Qwen3.8 4B Distill by Empero-AI and finetuned by iWebRoot โ GGUF
Model Overview
This repository contains the optimized GGUF quantization of Qwen3.8-4B-Empero-AI-Distill-FullStack, fine-tuned using the Unsloth framework for advanced full-stack web and mobile software development pipelines.
The base architecture features a full-parameter distillation of reasoning traces (Chain-of-Thought via <think>...</think> tags) from the frontier-scale Qwen3.8 2.4T A95B teacher model developed by Empero-AI. This configuration offers advanced local planning, logic, and code compilation compliance within a highly efficient 4-billion parameter footprint.
๐ Repository Links
- Safetensors Version (9.3GB Heavy Build): https://huggingface.co/iBotIA/Qwen3.8-4B-Empero-AI-FullStack
- GGUF Version (3.5GB Optimized Quantization): https://huggingface.co/iBotIA/Qwen3.8-4B-Empero-AI-FullStack-GGUF
๐ Injected Knowledge Stack (Fine-Tuning Data)
The model underwent continuous pre-training on 279,049 curated data segments across 8 strictly isolated developer knowledge directories:
- Mobile / Cross-Platform: Flutter (Modern structural widgets and lifecycle state management).
- Full-Stack Web Architecture: React Router v8 (Framework Mode via Vite, server-loaders, and async server-actions routing).
- Backend & Runtime Engine: NestJS & Node.js (API architecture, scalable middleware, and server streams).
- Database & Persistence Layers: Prisma ORM & Drizzle ORM (Schema modeling, relational builders, and safe SQL migrations).
- Language & System Rigor: TypeScript (Strict typing patterns to enforce self-debugging and runtime stability).
- Design & UI Systems: Tailwind CSS & Shadcn UI / Radix Primitives (Utility class layout embedded in JSX/TSX components).
๐ Training Logs & Learning Curve
The fine-tuning process completed 250 hardware-optimized steps on a T4 GPU. The learning curve showed a definitive late convergence ("Eureka" moment) near step 140, where the weights successfully aligned cross-stack framework logic.
- Step 10 (Start): Loss =
3.643480 - Step 50: Loss =
2.829789 - Step 140 (Logical drop): Loss =
2.788493 - Step 250 (Final score): Loss =
2.508277
๐ฆ Quantization Specifications
- File:
Qwen3.8-4B-Empero-AI-Distill-FullStack-Q6_K.gguf - Format: Q6_K (6-bit quantization)
- Size: ~3.56 GB
- Quality: Near-lossless precision compared to the 16-bit reference build.
๐ป Local Execution Guide (Target: GTX 1050 4GB VRAM)
When deploying this GGUF file inside Unsloth Desktop, LM Studio, Jan, or Ollama, configure these 3 runtime settings to prevent system stuttering:
- GPU Offload: Set your hardware layer slider to 25 layers. This safely loads ~2.5 GB of the model weight into your NVIDIA GTX 1050 VRAM without freezing Windows, while the remaining compute safely overflows into your 16GB system RAM.
- Sampling Settings: Set
temperature=0.6,top_p=0.95, andtop_k=20. Avoid a raw greedy search (temperature=0) to prevent the reasoning tokens from falling into endless structural loops. - Context Window: Set the token length to `16384` or `32768`. This expanded context window allows autonomous agents to evaluate several source files at the same time.
๐ ๏ธ Execution with OpenCode Autonomous Agent
To launch this model as an active developer backend connected to your terminal agent, run the OpenAI-compatible local engine server:
unsloth start opencode --context-length 32000Support / Donate
If this model helped you, consider supporting the project:
- BTC:
18cBC5sFjtctw121ULTkxTbTZPurginJBs - LTC:
ltc1q3jrcwrx66xpz4k92p08u8c5v8zwywk3dqpzdkv - USDT:
TGKVpbbznmvEusKbuZZj4WSK6XxtHcG6FE(TRX chain) - USDT:
0x1059cb5a1F8467e5b56a9bdf082cE86FFB002D15(POL chain) - USDT:
0x18b2AA731daeFD47DFFa278f3F856eAF80376fd6(ETH chain) - USDT:
0x3bEcddC7c49bDba5503eB1677628b4519439884c(BNB chain)
Provenance & Licensing
Quantizations are built upon [empero-ai/Qwen3.8-4B-Distill](https://huggingface.co/empero-ai/Qwen3.8-4B-Distill). Weights inherit the permissive Apache-2.0 license from the base Qwen repository and are shared as-is.
