Team Ai
Apppublic

badkarma0/fusechat-llama31-8b-instruct

sourceHugging Facellama3.1updated 13d agoView on Hugging Face
0likes
App README

FuseChat Llama 3.1 8B Instruct

A Gradio chat demo for `FuseAI/FuseChat-Llama-3.1-8B-Instruct`, running bfloat16 on ZeroGPU with streaming responses.

FuseChat 3.0 is built by implicit model fusion (IMF) rather than weight merging: Gemma-2-27B, Mistral-Large, Qwen-2.5-72B and Llama-3.1-70B are distilled into a single Llama 3.1 8B via supervised fine-tuning on reward-ranked responses, then preference-optimised — so no source weights are merged and inference cost is unchanged. See the paper.

BenchmarkScore
IFEval (0-shot, inst + prompt strict acc)72.05
BBH (3-shot, acc_norm)30.85
MATH Lvl 5 (4-shot, exact match)7.02
GPQA (0-shot, acc_norm)7.38

Each message reserves GPU time from the visitor's own daily ZeroGPU quota, so this Space costs its creator nothing. The first message of a session pays a one-off model load; later messages reuse a warm worker.

Why the parent model, not the 3-bit w3a16g40sym build

The original request pointed at `numen-tech/FuseChat-Llama-3.1-8B-Instruct-w3a16g40sym`, a 3-bit OmniQuant build. It cannot be executed by anything available on the Hub, and its parent model is not reachable through an inference provider either, so the only working path is to host the parent weights on ZeroGPU.

RequirementStatus
Routable by an inference providerNo — the repo's inferenceProviderMapping is empty, and the parent's advertises featherless-ai as status: live but the router rejects it (model_not_supported) under both explicit and auto routing. The metadata is stale.
Loadable by transformersNo — its config.json contains only {"quantization_config": {"bits": 3}}, with no architecture. The real config sits in private-llm-config.json, and the weights are 99 MLC-LLM-packed params_shard_*.bin files.
Loadable in-browser by WebLLMNo — the tensor layout is valid MLC-LLM (q_weight/q_scale, 325 tensors, complete for a 32-layer Llama 3.1), but WebLLM needs a compiled WebGPU kernel per architecture and quantization, and no Llama-3_1-8B-Instruct-q3f16_1 WebGPU library exists — only q4f16_1 and q4f32_1.

That repo is an artifact of the Private LLM mobile app and has zero downloads. Recompiling it with MLC-LLM against a q3f16_1 target would make it servable, but that is a build pipeline rather than a Space.