Team Ai
Modelpublic

FunAudioLLM/Fun-ASR-Nano-2512-vllm

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
3likes16kdownloads
Model Card

Fun-ASR-Nano-2512 for vLLM

This is the official vLLM-native packaging of `FunAudioLLM/Fun-ASR-Nano-2512`. It preserves the official checkpoint tensors and adds the layout required by vLLM's FunASRForConditionalGeneration implementation.

This repository does not define a new model and does not add LoRA weights. The 1,261 tensors in model.safetensors are bitwise equal to the tensors in the official source model.pt at revision 272c57b82523ada6fd87095e955f8e29100979ab.

Run with vLLM

The validated path uses float32 for the highest transcription fidelity:

bash
python -m pip install "vllm==0.27.1"

vllm serve FunAudioLLM/Fun-ASR-Nano-2512-vllm \
  --revision vllm-0.27.1-20260830 \
  --served-model-name fun-asr-nano \
  --dtype float32 \
  --gpu-memory-utilization 0.40 \
  --enforce-eager

Send an OpenAI-compatible transcription request:

bash
curl -fL \
  https://huggingface.co/FunAudioLLM/Fun-ASR-Nano-2512-vllm/resolve/vllm-0.27.1-20260830/example/zh.mp3 \
  -o zh.mp3

curl -sS http://127.0.0.1:8000/v1/audio/transcriptions \
  -F file=@zh.mp3 \
  -F model=fun-asr-nano \
  -F language=zh \
  -F temperature=0 \
  -F response_format=json

Expected text for the pinned sample:

text
开饭时间早上九点至下午五点。

Provenance

ItemValue
Official sourceFunAudioLLM/Fun-ASR-Nano-2512
Source revision272c57b82523ada6fd87095e955f8e29100979ab
Source model.pt SHA-25655ae0d2fee369f0f11cce0795f6927934ad17cf11b278a7e56a51272074160bb
Tensor count1,261
LoRA tensors0
Converted model.safetensors SHA-25696dfbec48282dd24d3334369a01e9e909f321ee39a1b0003c528c5379f68c1a6
Sample example/zh.mp3 SHA-2560e64de19e4ff9a02e682955c9112f32d2317cfdbb5bc2f3504664044c993f195

The complete machine-readable record is in `MODEL_PROVENANCE.json`. The conversion can be reproduced with `convert_from_official.py`.

The vLLM-native layout originated in the community work by `allendou/Fun-ASR-Nano-2512-vllm` and vLLM PR #33247, with subsequent format and initialization fixes in vLLM PRs #36108 and #44215. This official packaging keeps that attribution while anchoring the weights to the official FunAudioLLM checkpoint.

Validation boundary

The published evidence covers vLLM 0.27.1, PyTorch 2.13.0+cu129, Transformers 5.15.0, and one NVIDIA H100 80 GB GPU. The pinned Chinese sample returned the expected text in three consecutive deterministic requests. Other vLLM releases, accelerators, quantizations, and model quality across broader datasets require separate validation.

For the regular FunASR Python runtime, timestamps, speaker diarization, and streaming services, use the canonical `modelscope/FunASR` toolkit and the original checkpoint.

License

The official source model declares Apache License 2.0. See `LICENSE`. Third-party software such as vLLM remains subject to its own license.