Team Ai
Apppublic

iejfiapo/qwen-coder-assistant

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
0likes
App README

๐Ÿ’ป Qwen 2.5 Coder Django Web Assistant

A premium, local-first AI coding companion and chat platform powered by Qwen 2.5 Coder 1.5B (Q8_0 GGUF) running on Ollama, built with a robust Python/Django backend and a beautiful glassmorphism Bootstrap 5 user interface.

This application is configured for both one-click local execution on Windows/macOS/Linux and cloud deployment to Hugging Face Spaces (via Docker).


๐Ÿ“– Project History: From Legacy to Local Powerhouse

This repository has undergone a complete architectural transformation. Here is the timeline of what we did from the beginning:

  1. 1.The Legacy Stack: Originally, the application was designed with a heavy Next.js frontend and external dependencies.
  2. 2.The Python/Django Migration: To eliminate complex Node.js build pipelines and maximize local CPU/GPU efficiency, we migrated the entire core server and routing system to Python/Django, using a lightweight SQLite database for account configurations.
  3. 3.Model Decoupling: We separated the large 1.8GB model binary (qwen2.5-coder-1.5b-instruct-q8_0.gguf) and the Ollama configuration (Modelfile) into a dedicated model/ folder at the root. This keeps development files separate from model weights.
  4. 4.Auto-Creation Launchers: We developed START_LOCAL.bat and START_ONLINE.bat scripts to check if the Ollama model qwen-local is registered. If missing, it automatically creates it from the Modelfile.
  5. 5.Hugging Face Spaces Integration: We configured a custom Dockerfile and POSIX-compliant entrypoint.sh to allow the entire application (including the model) to run 24/7 on Hugging Face Spaces' free tier.
  6. 6.Compatibility & Shell Fixes:
  7. 7.Fixed Dockerfile to download Ollama via the official script (curl -fsSL https://ollama.com/install.sh | sh) instead of broken direct download links.
  8. 8.Fixed entrypoint.sh syntax errors (e.g. replacing bash-isms like {1..30} with standard while loops and fixing statement terminations) for Debian dash shell compatibility.
  9. 9.Added .gitattributes to enforce Unix LF line endings on shell scripts so they don't crash under Linux.

๐Ÿ› ๏ธ Developer Self-Questioning (FAQ)

Q: Why is SQLite used instead of a heavier database (like PostgreSQL)?

A: Django is utilized in this architecture for session management, administrative control, and local account setups. Conversational histories are saved in the client's browser using localStorage. This design keeps the backend extremely fast, resource-light, and serverless, making SQLite the perfect zero-configuration database.

Q: How does the application support the massive 32k context size?

A: Qwen 2.5 Coder natively supports a 32,768 token context window. In our views.py file, we configure the parameter num_ctx 32768 inside Ollama queries. To prevent token overflow, we implemented a context truncation strategy: when the total conversation history exceeds 28,672 tokens, the oldest messages are pruned while keeping the main system prompt intact.

Q: How does the webpage scraping work?

A: The Django backend intercepts any URL submitted in the chat input. It performs a request, extracts the raw HTML, strips away script tags, stylesheets, and styling comments, and passes up to 80,000 characters of cleaned markdown text directly into the AI's context.

Q: Why did the Hugging Face Space throw a /usr/bin/ollama: line 1: Not: command not found error?

A: In the original setup, the Dockerfile attempted to download the Ollama binary directly from https://ollama.com/download/ollama-linux-amd64. This URL returned a "Not Found" html string instead of the binary. This string was saved to /usr/bin/ollama and when executed, threw a syntax error. We resolved this by using the official installation script: curl -fsSL https://ollama.com/install.sh | sh.

Q: Should we upload the 1.8GB model binary to Hugging Face or download it dynamically?

A: Pushing a 1.8GB binary file over your local internet connection can be slow and timeout. Because of this, entrypoint.sh includes an automatic fallback: if the model binary is missing from the model/ folder, the container will automatically download it from Hugging Face's official high-speed hub during container startup. However, if you have a stable and fast internet connection, you can copy the GGUF file into the model/ directory, track it with Git LFS, and push it. This guarantees that your Space starts up instantly without downloading anything.


๐Ÿš€ Local Run Instructions

  1. 1.Requirements: Python 3.10+ and Ollama must be installed.
  2. 2.Run: Double-click START_LOCAL.bat in the root folder.
  3. 3.This automated launcher checks for Ollama, sets up a virtual environment, installs packages, runs database migrations, registers the custom model with a 32k context, and launches Django.
  4. 4.Access: Navigate to http://127.0.0.1:8000 (or http://localhost:3000 depending on port routing).

For manual commands, please see LOCAL_RUN.txt.


๐ŸŒ Sharing & Hosting Options

  • โ€”Option A: Public Link via Local Tunnel (ngrok): Double-click START_ONLINE.bat to launch a secure HTTPS tunnel to your local PC.
  • โ€”Option B: 24/7 Cloud Hosting (Hugging Face Spaces): Deploy the Docker image to a free CPU Basic Space. The model will run 24/7 in the cloud.

For step-by-step guides, refer to HOSTING_GUIDE.md.