iejfiapo/qwen-coder-assistant
๐ป Qwen 2.5 Coder Django Web Assistant
A premium, local-first AI coding companion and chat platform powered by Qwen 2.5 Coder 1.5B (Q8_0 GGUF) running on Ollama, built with a robust Python/Django backend and a beautiful glassmorphism Bootstrap 5 user interface.
This application is configured for both one-click local execution on Windows/macOS/Linux and cloud deployment to Hugging Face Spaces (via Docker).
๐ Project History: From Legacy to Local Powerhouse
This repository has undergone a complete architectural transformation. Here is the timeline of what we did from the beginning:
- The Legacy Stack: Originally, the application was designed with a heavy Next.js frontend and external dependencies.
- The Python/Django Migration: To eliminate complex Node.js build pipelines and maximize local CPU/GPU efficiency, we migrated the entire core server and routing system to Python/Django, using a lightweight SQLite database for account configurations.
- Model Decoupling: We separated the large 1.8GB model binary (
qwen2.5-coder-1.5b-instruct-q8_0.gguf) and the Ollama configuration (Modelfile) into a dedicatedmodel/folder at the root. This keeps development files separate from model weights. - Auto-Creation Launchers: We developed
START_LOCAL.batandSTART_ONLINE.batscripts to check if the Ollama modelqwen-localis registered. If missing, it automatically creates it from theModelfile. - Hugging Face Spaces Integration: We configured a custom
Dockerfileand POSIX-compliantentrypoint.shto allow the entire application (including the model) to run 24/7 on Hugging Face Spaces' free tier. - Compatibility & Shell Fixes:
- Fixed
Dockerfileto download Ollama via the official script (curl -fsSL https://ollama.com/install.sh | sh) instead of broken direct download links. - Fixed
entrypoint.shsyntax errors (e.g. replacing bash-isms like{1..30}with standardwhileloops and fixing statement terminations) for Debiandashshell compatibility. - Added
.gitattributesto enforce UnixLFline endings on shell scripts so they don't crash under Linux.
๐ ๏ธ Developer Self-Questioning (FAQ)
Q: Why is SQLite used instead of a heavier database (like PostgreSQL)?
A: Django is utilized in this architecture for session management, administrative control, and local account setups. Conversational histories are saved in the client's browser using localStorage. This design keeps the backend extremely fast, resource-light, and serverless, making SQLite the perfect zero-configuration database.
Q: How does the application support the massive 32k context size?
A: Qwen 2.5 Coder natively supports a 32,768 token context window. In our views.py file, we configure the parameter num_ctx 32768 inside Ollama queries. To prevent token overflow, we implemented a context truncation strategy: when the total conversation history exceeds 28,672 tokens, the oldest messages are pruned while keeping the main system prompt intact.
Q: How does the webpage scraping work?
A: The Django backend intercepts any URL submitted in the chat input. It performs a request, extracts the raw HTML, strips away script tags, stylesheets, and styling comments, and passes up to 80,000 characters of cleaned markdown text directly into the AI's context.
Q: Why did the Hugging Face Space throw a /usr/bin/ollama: line 1: Not: command not found error?
A: In the original setup, the Dockerfile attempted to download the Ollama binary directly from https://ollama.com/download/ollama-linux-amd64. This URL returned a "Not Found" html string instead of the binary. This string was saved to /usr/bin/ollama and when executed, threw a syntax error. We resolved this by using the official installation script: curl -fsSL https://ollama.com/install.sh | sh.
Q: Should we upload the 1.8GB model binary to Hugging Face or download it dynamically?
A: Pushing a 1.8GB binary file over your local internet connection can be slow and timeout. Because of this, entrypoint.sh includes an automatic fallback: if the model binary is missing from the model/ folder, the container will automatically download it from Hugging Face's official high-speed hub during container startup. However, if you have a stable and fast internet connection, you can copy the GGUF file into the model/ directory, track it with Git LFS, and push it. This guarantees that your Space starts up instantly without downloading anything.
๐ Local Run Instructions
- Requirements: Python 3.10+ and Ollama must be installed.
- Run: Double-click
START_LOCAL.batin the root folder. - This automated launcher checks for Ollama, sets up a virtual environment, installs packages, runs database migrations, registers the custom model with a 32k context, and launches Django.
- Access: Navigate to
http://127.0.0.1:8000(orhttp://localhost:3000depending on port routing).
For manual commands, please see LOCAL_RUN.txt.
๐ Sharing & Hosting Options
- Option A: Public Link via Local Tunnel (ngrok): Double-click
START_ONLINE.batto launch a secure HTTPS tunnel to your local PC. - Option B: 24/7 Cloud Hosting (Hugging Face Spaces): Deploy the Docker image to a free CPU Basic Space. The model will run 24/7 in the cloud.
For step-by-step guides, refer to HOSTING_GUIDE.md.
