Uddom/voxcpm2-khmer-webgpu
VoxCPM2: ONNX graphs for in-browser WebGPU inference
A converted and 4-bit quantized copy of `openbmb/VoxCPM2` (2B text-to-speech model with voice design and voice cloning, 30 languages including Khmer), built to run entirely in the browser with onnxruntime-web on WebGPU. No server needed.
Total download: about 1.9 GB. Start at manifest.json: it lists every file, its size, its SHA-256 and the config the decode loop needs.
Contents
Graph input and output names and the decoding loop are specified by the project that produced these files. The graphs use plain MatMul/Softmax attention (no GroupQueryAttention) and avoid ops that the ORT WebGPU provider lacks. Noise is fed in as an input so results are seedable.
What was changed, and how well it matches the original
- Weights were converted from the upstream checkpoint into ONNX graphs and quantized to int4 with GPTQ calibration (plain round-to-nearest int4 was clearly worse and is not shipped).
- Against the fp32 PyTorch reference, teacher-forced patch cosine similarity was at least 0.974 (mean 0.985) across six test prompts, and the number of generated patches stayed within a few patches of the reference.
- Free-running output is chaotic, so the generated waveform is not identical to the original, even for fp16. Voice and prosody will differ from the PyTorch model for the same seed.
- Khmer intelligibility of this build has not been formally evaluated (no Khmer ASR was used). Please judge by listening.
- Tested on one NVIDIA GPU with Chrome (WebGPU), at roughly real-time speed. Other GPUs, browsers and phones are untested.
Responsible use
VoxCPM2 can clone voices. Upstream forbids use for impersonation, fraud or disinformation, and asks that AI-generated audio be clearly labeled. Only clone voices you own or have permission to use.
License and attribution
Apache License 2.0. See LICENSE and NOTICE. This is a modified derivative of VoxCPM2 by OpenBMB / the VoxCPM Team. It is not endorsed by or affiliated with them. If you use it, please cite the upstream work (see the VoxCPM2 model card).
