FunAudioLLM/Fun-ASR-Nano-2512-vllm
Fun-ASR-Nano-2512 for vLLM
This is the official vLLM-native packaging of `FunAudioLLM/Fun-ASR-Nano-2512`. It preserves the official checkpoint tensors and adds the layout required by vLLM's FunASRForConditionalGeneration implementation.
This repository does not define a new model and does not add LoRA weights. The 1,261 tensors in model.safetensors are bitwise equal to the tensors in the official source model.pt at revision 272c57b82523ada6fd87095e955f8e29100979ab.
Run with vLLM
The validated path uses float32 for the highest transcription fidelity:
python -m pip install "vllm==0.27.1"
vllm serve FunAudioLLM/Fun-ASR-Nano-2512-vllm \
--revision vllm-0.27.1-20260830 \
--served-model-name fun-asr-nano \
--dtype float32 \
--gpu-memory-utilization 0.40 \
--enforce-eagerSend an OpenAI-compatible transcription request:
curl -fL \
https://huggingface.co/FunAudioLLM/Fun-ASR-Nano-2512-vllm/resolve/vllm-0.27.1-20260830/example/zh.mp3 \
-o zh.mp3
curl -sS http://127.0.0.1:8000/v1/audio/transcriptions \
-F file=@zh.mp3 \
-F model=fun-asr-nano \
-F language=zh \
-F temperature=0 \
-F response_format=jsonExpected text for the pinned sample:
开饭时间早上九点至下午五点。Provenance
The complete machine-readable record is in `MODEL_PROVENANCE.json`. The conversion can be reproduced with `convert_from_official.py`.
The vLLM-native layout originated in the community work by `allendou/Fun-ASR-Nano-2512-vllm` and vLLM PR #33247, with subsequent format and initialization fixes in vLLM PRs #36108 and #44215. This official packaging keeps that attribution while anchoring the weights to the official FunAudioLLM checkpoint.
Validation boundary
The published evidence covers vLLM 0.27.1, PyTorch 2.13.0+cu129, Transformers 5.15.0, and one NVIDIA H100 80 GB GPU. The pinned Chinese sample returned the expected text in three consecutive deterministic requests. Other vLLM releases, accelerators, quantizations, and model quality across broader datasets require separate validation.
For the regular FunASR Python runtime, timestamps, speaker diarization, and streaming services, use the canonical `modelscope/FunASR` toolkit and the original checkpoint.
License
The official source model declares Apache License 2.0. See `LICENSE`. Third-party software such as vLLM remains subject to its own license.
