inductiveML/monarch-webgpu
1
1import {createRoot} from 'react-dom/client';2import MonarchDemo from './components/MonarchDemo';3import './styles/site.css';4import './styles/v9.css';5import './styles/monarch.css';6import './styles/space.css';7 8function App(){return <div className="page monarch-page"><main className="v9 space-page"><div className="pad"><div className="grid12">9 <header className="space-header"><a href="https://inductive.ml/experiments/monarch" target="_blank" rel="noreferrer">← Read the experiment</a><a href="https://huggingface.co/spaces/inductiveML/monarch-webgpu/tree/main" target="_blank" rel="noreferrer">Source code ↗</a></header>10 <div className="secbar"><span className="num">§ 01</span><span className="ttl">// MONARCH</span><span className="meta">WebGPU · open-source kernels</span></div>11 <div className="paper-hero"><div className="head"><h1>Small model.<br/>Your GPU.</h1><p className="sub">Run Liquid AI’s <em>LFM2.5-230M</em> with our WGSL decoder. Generate a response, then measure your own tokens per second.</p></div><div className="meta-rail"><div className="row"><span>// kernel</span><span>MIX_M40</span></div><div className="row"><span>// weights</span><span>int4 / block 32</span></div><div className="row"><span>// dispatches</span><span>72 / token</span></div><div className="row"><span>// inference</span><span>on your device</span></div></div></div>12 <div className="demo-section"><MonarchDemo/></div>13 <div className="secbar" id="compatibility"><span className="num">§ 02</span><span className="ttl">// Before comparing numbers</span></div>14 <section className="paper"><p>The recorded lab result was <strong>1,574.4 tok/s</strong> on an Apple M4 Max in fullscreen Chromium at context 192. The live counter reports only your own run. Loading and prompt prefill are excluded from the fixed decode benchmark; generation shows time to first token separately.</p><p>This decoder needs WebGPU, shader-f16, fixed 32-lane subgroups, 1,024-thread workgroups, and a large model buffer. Current Chromium on a compatible Apple Silicon Mac is the tested path. Unsupported devices are checked before model downloads. If the embedded view restricts browser features, <a href="https://inductiveml-monarch-webgpu.static.hf.space" target="_blank" rel="noreferrer">open the app directly ↗</a>.</p><p>Keep the tab visible, use AC power, and pause other GPU workloads for steadier measurements. Model and tokenizer downloads total about 174 MB and can be cached. Prompts and generated text stay in your browser.</p><p>The research kernel remains experimental: <code>NO_VERDICT_SPLIT</code>. This demo uses ordinary dispatches. No persistent decoder was built. <a href="https://inductive.ml/experiments/monarch" target="_blank" rel="noreferrer">Read the methodology and limitations ↗</a>.</p></section>15 <div className="links space-links"><a className="btn" href="/monarch/research-results.json" download>Recorded results ↓</a><a className="btn" href="https://huggingface.co/inductiveML/LFM2.5-230M-MONARCH" target="_blank" rel="noreferrer">Model & checksums ↗</a><a className="btn" href="https://huggingface.co/spaces/inductiveML/monarch-webgpu/blob/main/README.md" target="_blank" rel="noreferrer">Build it yourself ↗</a></div>16 <footer className="foot"><div className="a">Inductive ML · MONARCH</div><div className="b">Code: Apache-2.0</div><div className="c"><a href="/monarch/LICENSE-LFM.txt">Model: LFM Open License v1.0</a></div></footer>17 </div></div></main></div>}18 19createRoot(document.getElementById('root')!).render(<App/>);20 