badkarma0/fusechat-llama31-8b-instruct
FuseChat Llama 3.1 8B Instruct
A Gradio chat demo for `FuseAI/FuseChat-Llama-3.1-8B-Instruct`, running bfloat16 on ZeroGPU with streaming responses.
FuseChat 3.0 is built by implicit model fusion (IMF) rather than weight merging: Gemma-2-27B, Mistral-Large, Qwen-2.5-72B and Llama-3.1-70B are distilled into a single Llama 3.1 8B via supervised fine-tuning on reward-ranked responses, then preference-optimised — so no source weights are merged and inference cost is unchanged. See the paper.
Each message reserves GPU time from the visitor's own daily ZeroGPU quota, so this Space costs its creator nothing. The first message of a session pays a one-off model load; later messages reuse a warm worker.
Why the parent model, not the 3-bit w3a16g40sym build
The original request pointed at `numen-tech/FuseChat-Llama-3.1-8B-Instruct-w3a16g40sym`, a 3-bit OmniQuant build. It cannot be executed by anything available on the Hub, and its parent model is not reachable through an inference provider either, so the only working path is to host the parent weights on ZeroGPU.
That repo is an artifact of the Private LLM mobile app and has zero downloads. Recompiling it with MLC-LLM against a q3f16_1 target would make it servable, but that is a build pipeline rather than a Space.
