AbelWa/Test
<div align="center">
Audar-ASR-V1-Turbo Β· GGUF
Audar's proprietary Arabic speech-recognition model β leaderboard-grade, dialect-aware.
From Arabic to the world.
-brightgreen)
<p><a href="#-what-it-is"><b>π§ Overview</b></a> Β· <a href="#-benchmarks"><b>π Benchmarks</b></a> Β· <a href="#-gguf-inference-llamacpp"><b>π» GGUF Deploy</b></a> Β· <a href="#-real-time-streaming"><b>ποΈ Streaming</b></a> Β· <a href="https://www.audarai.com"><b>βοΈ Audar API</b></a> Β· <a href="https://www.audarai.com/license/audarai-community-license-v1.0/"><b>π License</b></a></p>
</div>
π§ What it is
Audar-ASR-V1-Turbo is Audar's proprietary Arabic speech-recognition model β the accuracy tier of the Audar-ASR family. It recasts transcription as audio-conditioned next-token prediction over a unified text vocabulary (a language-model decoder rather than a CTC or transducer objective), and is developed in-house through a proprietary Arabic training program:
- π§± Large-scale dialectal pretraining β 300,000+ hours of Arabic audio spanning MSA, Gulf, Egyptian, Levantine and Maghrebi speech, code-switching, and diverse acoustic channels.
- π― Dialect-targeted fine-tuning β hardness sampling and multi-task conditioning focused on proper nouns, code-switching, and dialect-faithful orthography.
- π§ GRPO reinforcement-learning alignment β preference optimization against Arabic-native failure modes (diacritization, code-switching, named-entity preservation, formatting) with trained native annotators.
The result is state-of-the-art dialectal Arabic ASR β the lowest average WER of any evaluated system on the Open Universal Arabic ASR Leaderboard. It transcribes MSA and every major Arabic dialect, code-switched ArabicβEnglish, and English, across 30 languages in total. For real-time, edge, or high-throughput deployment, see the smaller **Audar-ASR-V1-Flash**.
Distributed in the widely-supported Qwen3-ASR architecture format for turnkey tooling (llama.cpp / GGUF). The model β data, training curriculum, and alignment β is Audar's.
Model summary
<table> <tbody> <tr><td width="200"><b>Model</b></td><td>Audar-ASR-V1-Turbo β proprietary Arabic ASR (accuracy tier)</td></tr> <tr><td><b>Task</b></td><td>Automatic speech recognition (audio β text)</td></tr> <tr><td><b>Approach</b></td><td>Generative ASR β audio encoder + language-model decoder (audio-conditioned next-token prediction)</td></tr> <tr><td><b>Training</b></td><td>300k+ hrs dialectal pretraining β dialect-targeted SFT β GRPO alignment</td></tr> <tr><td><b>Decoder parameters</b></td><td>2,031,739,904 (2.03B)</td></tr> <tr><td><b>Audio encoder parameters</b></td><td>317,477,504 (0.32B)</td></tr> <tr><td><b>Total parameters</b></td><td>2,349,217,408 (2.35B, bf16)</td></tr> <tr><td><b>Audio input</b></td><td>16 kHz mono; 30 s context (longer audio is chunked/streamed)</td></tr> <tr><td><b>Languages</b></td><td>Arabic (MSA + Gulf/Egyptian/Levantine/Maghrebi dialects) + English + 28 more</td></tr> <tr><td><b>Runtime</b></td><td>GGUF / llama.cpp β CPU Β· GPU Β· edge</td></tr> <tr><td><b>License</b></td><td>AudarAI Community License v1.0</td></tr> </tbody> </table>
π Benchmarks
Arabic dialectal ASR is hard β heavily dialectal, conversational, code-switched speech is the frontier for every system. On the Open Universal Arabic ASR Leaderboard, Audar-ASR-V1-Turbo posts the lowest average WER of any evaluated system on the full test sets β 24.7 %, best on four of the six β and 3.55 % WER on CommonVoice-18 Arabic. The per-dataset development-protocol results (100 utterances/benchmark) are below.
Open Universal Arabic ASR Leaderboard β WER % (lower is better)
Per-dataset WER (%), development protocol (100 utterances/benchmark); baselines are the leaderboard's published full-test scores. Best per column in bold. Authoritative full-test-set average: 24.7 %.
Emirati Arabic
On Emirati, the real recognition error is β 7.3 % β near-parity with spontaneous English β while the residual up to 19.4 % WER is largely orthographic convention (near-miss spelling of the same word, e.g. Ψ§ΩΨͺΩβΨ§ΩΨͺΩΨ§, and Latin-vs-Arabic rendering of English loanwords), not misrecognition.
Measured on an internal dialectal validation sample
Same sample and harness as the [Flash card](https://huggingface.co/audarai/Audar-ASR-V1-Flash#-benchmarks) β useful for a direct Flash-vs-Turbo comparison (WER/CER %, N clips per set).
Casablanca 61.9 WER β the official leaderboard's 62.87 (reproduced in-house) β the numbers line up.
π» GGUF inference (llama.cpp)
Turbo runs on llama.cpp via the multimodal (mtmd) path β a quantized decoder GGUF plus a BF16 audio projector (mmproj). Build a recent llama.cpp (with Qwen3-ASR support), then:
./llama-mtmd-cli \
-m Audar-ASR-V1-Turbo-Q8_0.gguf \
--mmproj mmproj-Audar-ASR-V1-Turbo.gguf \
--audio clip.wav \
-sys "ΩΨ±ΩΨΊ Ψ§ΩΩΩΨ§Ω
Ψ§ΩΨΉΨ±Ψ¨Ω Ψ§ΩΨͺΨ§ΩΩ." \
--temp 0β οΈ The audio projector (`mmproj`) must stay BF16 (its ClippableLinear is numerically sensitive). The decoder quantizes normally.Prefer a managed endpoint? The Audar-ASR family is also available via the **Audar API/SDK** β streaming, speaker-attributed transcription, and diarization, production-hosted.
GGUF variants
ποΈ Real-time streaming
Audar-ASR streams via LocalAgreement-2: as audio arrives the trailing window is re-decoded each hop and a word is committed only once two consecutive decodes agree on it β giving stable, low-latency incremental output over the GGUF runtime. Audar's production realtime engine serves the same policy over an OpenAI-Realtime-compatible WebSocket with model-based endpointing and β₯64 concurrent streams on a single A100-80GB.
π Languages, dialects & tasks
- Primary: Arabic β MSA and dialectal (Gulf/Emirati, Egyptian, Levantine, Maghrebi), plus code-switched ArabicβEnglish; emits dialect-faithful orthography from audio alone.
- Also: English + 28 additional languages.
- Task: transcription (audio β UTF-8 text), prompt-steerable for language and formatting.
Intended use & limitations
Intended use. Broadcast/media transcription, meeting & contact-center intelligence, voice agents, captioning, and accessibility β cloud or on-prem.
Limitations.
- Maghrebi / Moroccan Darija (Casablanca) remains the hardest condition (~63 % WER) for all systems.
- Heavily code-switched telephony and low-SNR audio degrade accuracy relative to clean MSA.
- Long-form audio can drift on very long recordings.
- Not evaluated for, and must not be used for, covert speaker identification.
π License
Released under the AudarAI Community License v1.0 β research and limited commercial use for qualifying Community Entities; enterprise / large-scale / MaaS use requires an AudarAI Enterprise License. See audarai.com/license/audarai-community-license-v1.0.
Citation
@misc{audar-asr-turbo-2026,
title = {Audar-ASR: Dialect-Aware Arabic Speech Recognition},
author = {AudarAI},
year = {2026},
note = {Audar-ASR-V1-Turbo},
url = {https://huggingface.co/audarai/Audar-ASR-V1-Turbo}
}About AudarAI
<div align="center">
Leading Arabic-First Multilingual Audio Intelligence
AudarAI starts with Arabic β and expands to the world.
</div>
We are building advanced multilingual audio intelligence that helps individuals, enterprises, and governments communicate across languages, cultures, and borders. By combining Arabic-first speech technology with global multilingual AI, AudarAI transforms voice into understanding, interaction, and connection.
Our work spans speech recognition, speech understanding, voice-enabled digital assistants, human-computer interaction, and intelligent audio systems designed for real-world impact. From empowering people to access technology in their native language to helping organizations communicate globally, AudarAI is shaping a future where every voice can be heard, understood, and connected.
Arabic-first. Multilingual by design. Human-centered at heart.
<div align="center">
[π www.audarai.com](https://www.audarai.com) Β· π€ Hugging Face Β· GitHub Β· contact@audarai.com
Β© 2026 AUDARAI PTE. LTD. Β· Licensed under the AudarAI Community License v1.0
</div>
