NabuxAi/buddi-asr-source-code
Buddi-ASR A privacy-aware training and deployment pipeline for a small Arabic/Gulf-Arabic + English child-speech model. فارسی · GPU training · Data governance · Experiment plan · Raspberry Pi Buddi-ASR turns multilingual Whisper Tiny or Base into a domain model for short child utterances, then exports the merged checkpoint to a quantized whisper.cpp artifact for Raspberry Pi, mobile, or server inference. This repository contains the reproducible pipeline, not fabricated model… See the full description on the dataset page: https://huggingface.co/datasets/NabuxAi/buddi-asr-source-code.
Record the executable bit on the shell scripts
Call the CLI as a module, not through its console script
Stop excluding the Python that has been running this all along
Let the data stage run without a GPU, and split the Colab notebook by intent
Add a Colab path to the same artifact
Add the Whisper sweep this repo was missing
Let evaluation run in batches and on a stratified sample
Resolve tf32 against the GPU, and reject GPUs PyTorch cannot use
Declare a torchao floor so LoRA injection survives the Kaggle image
Cover the warmup-argument fallback with a test
Make GPU training run on the Kaggle image
fix
Initial commit of buddi-asr source
initial commit
