Team Ai
Apppublic

inductiveML/monarch-webgpu

sourceHugging Faceapache-2.0updated 25d agoView on Hugging Face
1likes
README.md107 linesDownload Raw Back to root
1---2title: MONARCH — LFM2.5 WebGPU Kernels3emoji: 🦋4colorFrom: red5colorTo: yellow6sdk: static7app_file: index.html8fullWidth: true9header: mini10license: apache-2.011short_description: Run LFM2.5-230M locally and benchmark our WGSL decoder.12models:13  - inductiveML/LFM2.5-230M-MONARCH14  - LiquidAI/LFM2.5-230M15tags:16  - webgpu17  - wgsl18  - text-generation19  - benchmark20  - javascript21custom_headers:22  cross-origin-embedder-policy: require-corp23  cross-origin-opener-policy: same-origin24  cross-origin-resource-policy: cross-origin25---26# MONARCH — LFM2.5 WebGPU Kernels27 28An open-source browser decoder for Liquid AI's LFM2.5-230M. Load the real model, generate a response, and measure your GPU. Inference happens on your device. No inference API, hosted GPU, or private model credentials are needed.29 30[Experiment and methodology](https://inductive.ml/experiments/monarch) · [Model files](https://huggingface.co/inductiveML/LFM2.5-230M-MONARCH) · [Direct app](https://inductiveml-monarch-webgpu.static.hf.space)31 32## Run locally33 34Requires Node.js 20.19+ (or 22.12+) and npm.35 36```sh37git clone https://huggingface.co/spaces/inductiveML/monarch-webgpu38cd monarch-webgpu39npm ci40npm test41npm run dev42```43 44Open the local URL Vite prints. `npm run build` produces `dist/` and refreshes the committed static files (`index.html`, `assets/`, and `monarch/`). `npm run preview` serves the production build with the required isolation headers. The build downloads no model weights and does not depend on any private research checkout. Model files download on demand when you click **Load model**.45 46For browser support, the runtime checks WebGPU, shader-f16, fixed 32-lane subgroups, 1,024-thread workgroups, and the storage-buffer limit before downloading weights. Current Chromium on a compatible Apple Silicon Mac is the tested path. Other devices may fail the feature or known-answer check. There is no silent server fallback.47 48The Space serves the committed build directly, so duplicating it requires no build credits. After changing source, run `npm ci && npm test && npm run build`, then commit the source and regenerated static files. `source.html` is the Vite input; `index.html` is generated. The dev server serves `source.html` at `/`.49 50## What is included51 52- Nine prompt presets, editable prompts, Markdown output, Stop and Unload controls.53- Real greedy single-stream generation with live decode TPS and time to first token.54- A fixed four-run benchmark at context 192 or 1,024, plus downloadable JSON results.55- The complete source needed to build the worker: **28 pinned TypeScript/WGSL research files**, tokenizer implementation, selected configuration, and correctness fixture.56- Apache-2.0 source license, separate model license, attribution, and recorded research results.57 58The `MIX_M40` decoder uses 72 ordinary dispatches per token. Its kernels fuse convolution work, specialize matrix geometry, use online subgroup attention, and make precision choices explicit. It does not use the experimental cross-workgroup atomic visibility protocol or a persistent token kernel.59 60## Source map61 62| Path | Purpose |63| --- | --- |64| `src/monarch/vendor/monarch/` | Original research kernel and runtime source; unchanged |65| `src/monarch/source-manifest.json` | SHA-256 for each of the 28 research files |66| `src/monarch/model-config.json` | Model shape, 199 buffer sections, allocation limits, and MIX_M40 configuration |67| `src/monarch/worker.ts` | Browser integration, download verification, known-answer check, generation, benchmark |68| `src/monarch/model-source.json` | Model repository, immutable revision, and tokenizer base URL |69| `src/components/MonarchDemo.tsx` | Controls and actual run metrics |70| `src/components/MonarchMarkdown.tsx` | CommonMark/GFM rendering with raw HTML disabled |71| `scripts/build-worker.mjs` | Source-integrity checks and browser worker build |72| `public/monarch/research-results.json` | Recorded lab measurements and research gate status |73 74The vendored code retains its original directory names so its relative imports and source hashes stay intact. An esbuild/Vite build packages it without changing the WGSL kernel bodies. `npm run check:source` verifies those hashes.75 76## Model integrity and provenance77 78The model repository contains nine preconverted transport chunks totaling **168,647,680 bytes**, plus tokenizer data. The runtime verifies every chunk and the reconstructed buffer:79 80```text81ca1fea89fd9f3ca7e5d6d5720705c96f457a92edd0cfc811cf8d9391ac784cb082```83 84The upstream ONNX revision and layout transformations are documented in the [model card](https://huggingface.co/inductiveML/LFM2.5-230M-MONARCH). The Space uses an immutable model repository revision. It runs a known prefill seed and 32-token continuation check before enabling the UI. Errors and device loss stop the run.85 86## Benchmark protocol and limits87 88Generation is greedy, single-turn ChatML with at most 1,024 prompt tokens and 512 output tokens. Model loading and prompt prefill contribute to startup / time to first token. The live decode counter starts after the first token and includes streaming overhead. A short response may end inside the last computed 32-token group.89 90The fixed benchmark:91 921. Uses a pinned prompt and a 64-token warmup.932. Freshly primes four 256-token free-running continuations.943. Reads back every 32 tokens through a 64-token ring.954. Requires all four continuations to match exactly.965. Reports reciprocal median wall time per token and timing spread, excluding prefill.97 98The selected context is the center of the generated window: token positions 64–319 for 192, or 896–1151 for 1,024. Timestamp queries are optional; reported TPS is based on wall time.99 100The **1,574.4 tok/s** figure is a recorded M4 Max / fullscreen Chromium lab observation at context 192. The quick public benchmark uses normal browser settings and does not reproduce the full controlled lab protocol. The research kernel retained `NO_VERDICT_SPLIT`; later B1.11a pricing remained unadjudicated because environmental/thermal conditions failed. These conditions and the distinction between recorded and live measurements are part of the result, not hidden exceptions.101 102## Licenses and credits103 104Code owned by Inductive ML is released under **Apache-2.0** (`LICENSE`). This grant does not relicense model or third-party materials. Model weights and tokenizer data retain **Liquid AI's LFM Open License v1.0**, including its commercial-use conditions. Hugging Face tokenizers.js 0.1.3 is Apache-2.0. React and the Markdown dependencies retain their own licenses. See `NOTICE` and `THIRD_PARTY_NOTICES.md`.105 106The interface was inspired by [webml-community/lfm2-webgpu-kernels](https://huggingface.co/spaces/webml-community/lfm2-webgpu-kernels). The kernel/runtime implementation is MONARCH's; no source from that reference demo is included.107