akshayyy1/small-code-assistant
0
Small Code Assistant
A streaming chat demo for code generation, running Qwen/Qwen2.5-Coder-3B-Instruct (6.2 GB bf16) on ZeroGPU.
Why this model
The demo was originally scoped to CohereLabs/North-Mini-Code-1.0, which is a 30.5B MoE (~61 GB bf16, 49 safetensors shards) — far beyond a free Space. Qwen2.5-Coder-3B is the closest small stand-in that fits the brief (under 10 GB).
To swap models, change MODEL_ID at the top of app.py. North-Mini-Code itself is served remotely by Cohere (see the official demo) and would need a Cohere API key rather than local weights.
API
Open /api or use the MCP endpoint exposed via demo.launch(mcp_server=True).
