Team Ai
Modelpublic

AbelWa/Test

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
0likes45downloads
Model Card

<div align="center">

Audar-ASR-V1-Turbo Β· GGUF

Audar's proprietary Arabic speech-recognition model β€” leaderboard-grade, dialect-aware.

From Arabic to the world.

License Task Format Params Open-AR-ASR-brightgreen) CommonVoice Emirati

<p><a href="#-what-it-is"><b>🧭 Overview</b></a> Β· <a href="#-benchmarks"><b>πŸ“Š Benchmarks</b></a> Β· <a href="#-gguf-inference-llamacpp"><b>πŸ’» GGUF Deploy</b></a> Β· <a href="#-real-time-streaming"><b>πŸŽ™οΈ Streaming</b></a> Β· <a href="https://www.audarai.com"><b>☁️ Audar API</b></a> Β· <a href="https://www.audarai.com/license/audarai-community-license-v1.0/"><b>πŸ“œ License</b></a></p>

</div>


🧭 What it is

Audar-ASR-V1-Turbo is Audar's proprietary Arabic speech-recognition model β€” the accuracy tier of the Audar-ASR family. It recasts transcription as audio-conditioned next-token prediction over a unified text vocabulary (a language-model decoder rather than a CTC or transducer objective), and is developed in-house through a proprietary Arabic training program:

  • β€”πŸ§± Large-scale dialectal pretraining β€” 300,000+ hours of Arabic audio spanning MSA, Gulf, Egyptian, Levantine and Maghrebi speech, code-switching, and diverse acoustic channels.
  • β€”πŸŽ― Dialect-targeted fine-tuning β€” hardness sampling and multi-task conditioning focused on proper nouns, code-switching, and dialect-faithful orthography.
  • β€”πŸ§  GRPO reinforcement-learning alignment β€” preference optimization against Arabic-native failure modes (diacritization, code-switching, named-entity preservation, formatting) with trained native annotators.

The result is state-of-the-art dialectal Arabic ASR β€” the lowest average WER of any evaluated system on the Open Universal Arabic ASR Leaderboard. It transcribes MSA and every major Arabic dialect, code-switched Arabic–English, and English, across 30 languages in total. For real-time, edge, or high-throughput deployment, see the smaller **Audar-ASR-V1-Flash**.

Distributed in the widely-supported Qwen3-ASR architecture format for turnkey tooling (llama.cpp / GGUF). The model β€” data, training curriculum, and alignment β€” is Audar's.

Model summary

<table> <tbody> <tr><td width="200"><b>Model</b></td><td>Audar-ASR-V1-Turbo β€” proprietary Arabic ASR (accuracy tier)</td></tr> <tr><td><b>Task</b></td><td>Automatic speech recognition (audio β†’ text)</td></tr> <tr><td><b>Approach</b></td><td>Generative ASR β€” audio encoder + language-model decoder (audio-conditioned next-token prediction)</td></tr> <tr><td><b>Training</b></td><td>300k+ hrs dialectal pretraining β†’ dialect-targeted SFT β†’ GRPO alignment</td></tr> <tr><td><b>Decoder parameters</b></td><td>2,031,739,904 (2.03B)</td></tr> <tr><td><b>Audio encoder parameters</b></td><td>317,477,504 (0.32B)</td></tr> <tr><td><b>Total parameters</b></td><td>2,349,217,408 (2.35B, bf16)</td></tr> <tr><td><b>Audio input</b></td><td>16 kHz mono; 30 s context (longer audio is chunked/streamed)</td></tr> <tr><td><b>Languages</b></td><td>Arabic (MSA + Gulf/Egyptian/Levantine/Maghrebi dialects) + English + 28 more</td></tr> <tr><td><b>Runtime</b></td><td>GGUF / llama.cpp β€” CPU Β· GPU Β· edge</td></tr> <tr><td><b>License</b></td><td>AudarAI Community License v1.0</td></tr> </tbody> </table>

πŸ“Š Benchmarks

Arabic dialectal ASR is hard β€” heavily dialectal, conversational, code-switched speech is the frontier for every system. On the Open Universal Arabic ASR Leaderboard, Audar-ASR-V1-Turbo posts the lowest average WER of any evaluated system on the full test sets β€” 24.7 %, best on four of the six β€” and 3.55 % WER on CommonVoice-18 Arabic. The per-dataset development-protocol results (100 utterances/benchmark) are below.

Open Universal Arabic ASR Leaderboard β€” WER % (lower is better)

Per-dataset WER (%), development protocol (100 utterances/benchmark); baselines are the leaderboard's published full-test scores. Best per column in bold. Authoritative full-test-set average: 24.7 %.

SystemCommonVoice-18MASC-cleanMASC-noisyMGB-2SADACasablanca**Avg**
Audar-ASR-V1-Turbo3.559.1316.8414.0135.2262.8723.60
ElevenLabs Scribe v15.749.8719.7815.1540.8766.9326.39
Qwen3-ASR-1.7B (base)10.8615.0721.1229.2150.5485.2535.34
Whisper-Large-v317.8324.6634.6316.2655.9671.8136.86

Emirati Arabic

SetWER %CER %
Emirati (Mixat, full 1,585-clip test)19.47.3

On Emirati, the real recognition error is β‰ˆ 7.3 % β€” near-parity with spontaneous English β€” while the residual up to 19.4 % WER is largely orthographic convention (near-miss spelling of the same word, e.g. Ψ§Ω†ΨͺΩˆβ†”Ψ§Ω†Ψͺوا, and Latin-vs-Arabic rendering of English loanwords), not misrecognition.

Measured on an internal dialectal validation sample

Same sample and harness as the [Flash card](https://huggingface.co/audarai/Audar-ASR-V1-Flash#-benchmarks) β€” useful for a direct Flash-vs-Turbo comparison (WER/CER %, N clips per set).

Set (dialect)NWER %CER %
SawtArabi (Gulf)2313.72.7
ArzEn (Egyptian ⇄ English code-switch)4019.99.2
MGB-3 (Egyptian broadcast)4027.310.5
Casablanca (Maghrebi / Moroccan Darija)4061.928.6

Casablanca 61.9 WER β‰ˆ the official leaderboard's 62.87 (reproduced in-house) β€” the numbers line up.

πŸ’» GGUF inference (llama.cpp)

Turbo runs on llama.cpp via the multimodal (mtmd) path β€” a quantized decoder GGUF plus a BF16 audio projector (mmproj). Build a recent llama.cpp (with Qwen3-ASR support), then:

bash
./llama-mtmd-cli \
  -m       Audar-ASR-V1-Turbo-Q8_0.gguf \
  --mmproj mmproj-Audar-ASR-V1-Turbo.gguf \
  --audio  clip.wav \
  -sys     "فرّغ Ψ§Ω„ΩƒΩ„Ψ§Ω… Ψ§Ω„ΨΉΨ±Ψ¨ΩŠ Ψ§Ω„ΨͺΨ§Ω„ΩŠ." \
  --temp 0
⚠️ The audio projector (`mmproj`) must stay BF16 (its ClippableLinear is numerically sensitive). The decoder quantizes normally.

Prefer a managed endpoint? The Audar-ASR family is also available via the **Audar API/SDK** β€” streaming, speaker-attributed transcription, and diarization, production-hosted.

GGUF variants

FileApprox. sizeNotes
Audar-ASR-V1-Turbo-Q4_K_M.gguf~1.28 GBSmallest; constrained hardware
Audar-ASR-V1-Turbo-Q8_0.gguf~2.16 GBNear-lossless (recommended)
Audar-ASR-V1-Turbo.gguf (BF16)~4.07 GBFull precision decoder
mmproj-Audar-ASR-V1-Turbo.gguf~0.64 GBBF16 audio encoder β€” required, keep BF16

πŸŽ™οΈ Real-time streaming

Audar-ASR streams via LocalAgreement-2: as audio arrives the trailing window is re-decoded each hop and a word is committed only once two consecutive decodes agree on it β€” giving stable, low-latency incremental output over the GGUF runtime. Audar's production realtime engine serves the same policy over an OpenAI-Realtime-compatible WebSocket with model-based endpointing and β‰₯64 concurrent streams on a single A100-80GB.

🌍 Languages, dialects & tasks

  • β€”Primary: Arabic β€” MSA and dialectal (Gulf/Emirati, Egyptian, Levantine, Maghrebi), plus code-switched Arabic–English; emits dialect-faithful orthography from audio alone.
  • β€”Also: English + 28 additional languages.
  • β€”Task: transcription (audio β†’ UTF-8 text), prompt-steerable for language and formatting.

Intended use & limitations

Intended use. Broadcast/media transcription, meeting & contact-center intelligence, voice agents, captioning, and accessibility β€” cloud or on-prem.

Limitations.

  • β€”Maghrebi / Moroccan Darija (Casablanca) remains the hardest condition (~63 % WER) for all systems.
  • β€”Heavily code-switched telephony and low-SNR audio degrade accuracy relative to clean MSA.
  • β€”Long-form audio can drift on very long recordings.
  • β€”Not evaluated for, and must not be used for, covert speaker identification.

πŸ“œ License

Released under the AudarAI Community License v1.0 β€” research and limited commercial use for qualifying Community Entities; enterprise / large-scale / MaaS use requires an AudarAI Enterprise License. See audarai.com/license/audarai-community-license-v1.0.

Citation

bibtex
@misc{audar-asr-turbo-2026,
  title  = {Audar-ASR: Dialect-Aware Arabic Speech Recognition},
  author = {AudarAI},
  year   = {2026},
  note   = {Audar-ASR-V1-Turbo},
  url    = {https://huggingface.co/audarai/Audar-ASR-V1-Turbo}
}

About AudarAI

<div align="center">

Leading Arabic-First Multilingual Audio Intelligence

AudarAI starts with Arabic β€” and expands to the world.

</div>

We are building advanced multilingual audio intelligence that helps individuals, enterprises, and governments communicate across languages, cultures, and borders. By combining Arabic-first speech technology with global multilingual AI, AudarAI transforms voice into understanding, interaction, and connection.

Our work spans speech recognition, speech understanding, voice-enabled digital assistants, human-computer interaction, and intelligent audio systems designed for real-world impact. From empowering people to access technology in their native language to helping organizations communicate globally, AudarAI is shaping a future where every voice can be heard, understood, and connected.

Arabic-first. Multilingual by design. Human-centered at heart.

<div align="center">

[🌐 www.audarai.com](https://www.audarai.com) Β· πŸ€— Hugging Face Β· GitHub Β· contact@audarai.com

Β© 2026 AUDARAI PTE. LTD. Β· Licensed under the AudarAI Community License v1.0

</div>