avaeziaiteam/Audio-Encoder-Training-Data
Audio-Encoder-Training-Data Stage-1 data for adapting the audio encoder of Gemma 4 E2B to Persian ASR (audio -> Soniox transcript, no draft in the prompt). Three configs, one schema. Sensitive: the call-center part contains real customer calls (names, phone numbers, order details). Labels are machine-generated (Soniox stt-async-v5), not human transcripts. config rows hours train h val h movies 65,917 201.4 197.2 4.2 youtube 41,149 199.7 195.7 4.0 callcenter 64,895… See the full description on the dataset page: https://huggingface.co/datasets/avaeziaiteam/Audio-Encoder-Training-Data.
Conversations for this repository live on Hugging Face.
Team Ai shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face