NabuxAi/buddi-asr-source-code
Buddi-ASR A privacy-aware training and deployment pipeline for a small Arabic/Gulf-Arabic + English child-speech model. فارسی · GPU training · Data governance · Experiment plan · Raspberry Pi Buddi-ASR turns multilingual Whisper Tiny or Base into a domain model for short child utterances, then exports the merged checkpoint to a quantized whisper.cpp artifact for Raspberry Pi, mobile, or server inference. This repository contains the reproducible pipeline, not fabricated model… See the full description on the dataset page: https://huggingface.co/datasets/NabuxAi/buddi-asr-source-code.
Buddi-ASR
A privacy-aware training and deployment pipeline for a small Arabic/Gulf-Arabic + English child-speech model.
فارسی · GPU training · Data governance · Experiment plan · Raspberry Pi
Buddi-ASR turns multilingual Whisper Tiny or Base into a domain model for short child utterances, then exports the merged checkpoint to a quantized whisper.cpp artifact for Raspberry Pi, mobile, or server inference.
This repository contains the reproducible pipeline, not fabricated model weights. A real Buddi-ASR-Tiny-ArGulf-v1 checkpoint must be trained on licensed/consented audio and accepted only after speaker-disjoint evaluation.
flowchart TD
A["Licensed + consented audio"] --> B["PII-safe manifests"]
B --> C["Speaker-disjoint split"]
C --> D["Whisper Tiny/Base adaptation"]
D --> E["WER/CER by cohort"]
E --> F["Merge + Q5 export"]
F --> G["Raspberry Pi / mobile"]What is implemented
- Abjad-Kids download, 144-label audit, Romanized-label to Arabic conversion, and WAV duration probing.
- HMAC pseudonymization, anonymized processed filenames, consent-gated local CSV import, and PII-blocking manifest validation.
- Deterministic label-aware splitting with strict speaker isolation.
- Weighted dataset mixing without silently duplicating child utterances.
- Whisper Tiny/Base LoRA and full fine-tuning with Arabic/English per-example prompt tokens, waveform augmentation, early stopping, and reproducible configs.
- One-command CUDA training for RunPod/servers with VRAM-aware batch sizing, strict preflight, multi-GPU launch, automatic checkpoint resume, and runtime provenance.
- Normalized micro-averaged WER/CER overall and by language, dialect, age group, source, and category.
- Safe LoRA merge, official
whisper.cppconversion, Q4/Q5/Q8 quantization, artifact metadata, and SHA-256 checksum. - Dependency-free unit tests for all data-critical paths and a hardened Raspberry Pi service example.
Important dataset finding
The current Abjad-Kids repository reports 40,646 rows and 144 observed labels, while the associated 2026 paper reports 46,397 samples and 141 classes. The upstream CSV files also lack a documented speaker column; this project groups speakers from filenames only as an explicit heuristic. A full CSV audit parsed every row and found 40,568 filename-grouped rows plus 78 session-only rows that require special review.
Do not publish final metrics until:
- the seven ambiguous label mappings in
data/label_maps/abjad_ar.jsonhave been checked by listening; - session-only speaker groups have been resolved with anonymous overrides or acknowledged as a limitation;
- train, validation, and test show zero speaker overlap.
Quick start
Python 3.10–3.12 is supported. Data utilities have no third-party runtime dependencies.
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[data]'Create a stable private salt. Losing or changing it changes speaker groups; committing it defeats pseudonymization.
openssl rand -hex 32 > .buddi-speaker.saltDownload and prepare Abjad-Kids:
buddi-asr download-abjad data/raw/abjad-kids
buddi-asr prepare-abjad \
data/raw/abjad-kids \
data/manifests/abjad.jsonl \
--processed-audio-dir data/processed/abjad \
--salt-file .buddi-speaker.salt \
--audit-output reports/abjad-audit.json
buddi-asr validate data/manifests/abjad.jsonl --report reports/abjad-validation.jsonCreate speaker-disjoint sets:
buddi-asr split data/manifests/abjad.jsonl data/manifests/splits \
--train 0.8 --validation 0.1 --test 0.1 --seed 42Install GPU training dependencies and run the Tiny LoRA baseline:
./scripts/train_gpu.sh tiny-lora --smoke
./scripts/train_gpu.sh tiny-loraThe runner prepares Abjad-Kids when split manifests are absent. See the real GPU training guide before renting a GPU or adding private Omani data.
Evaluate exactly once on the held-out test set after experiment selection:
buddi-asr evaluate \
artifacts/tiny-lora/final \
data/manifests/splits/test.jsonl \
reports/tiny-lora-testMerge the adapter and export Q5 for the Pi:
buddi-asr merge-lora \
openai/whisper-tiny \
artifacts/tiny-lora/final \
artifacts/tiny-merged
git clone https://github.com/openai/whisper.git vendor/openai-whisper
git clone https://github.com/ggml-org/whisper.cpp.git vendor/whisper.cpp
cmake -S vendor/whisper.cpp -B vendor/whisper.cpp/build -DGGML_BLAS=1
cmake --build vendor/whisper.cpp/build -j --config Release
buddi-asr export-whisper-cpp \
artifacts/tiny-merged \
vendor/whisper.cpp \
vendor/openai-whisper \
artifacts/pi \
--name buddi-asr-tiny-argulf-v1 \
--quantization q5_0Add consented Omani/Gulf child speech
Use a private CSV with these required columns:
audio,text,speaker_key,language,consent_basis,dialect,age_group,device,noise_level
audio/001.wav,بودي أبغى قصة,anonymous-local-001,ar,verified_parental_consent,gulf_omani,6-8,iphone,quietspeaker_key is pseudonymized and never written to the output. Unknown columns are rejected to prevent accidental identity export.
buddi-asr import-local private/omani.csv data/manifests/omani.jsonl \
--processed-audio-dir data/processed/omani \
--salt-file .buddi-speaker.salt \
--source-name buddi_omani_childRepository map
configs/ Reproducible Tiny LoRA, Base full, and mixture configs
data/label_maps/ Audited Abjad Romanized-to-Arabic transcript map
deploy/systemd/ Hardened local-only Pi service
docs/ Governance, experiment, model-card, and Pi guides
scripts/train_gpu.sh One-command real CUDA training and recovery
src/buddi_asr/data/ Ingestion, manifest, split, mix, validation
src/buddi_asr/training/ Audio loading, augmentation, collator, trainer
src/buddi_asr/ CLI, metrics, evaluation, merge, export
tests/ Dependency-free unit and privacy regression testsNon-goals
- Training from scratch: Whisper pretraining is deliberately reused.
- Claiming “Gulf child ASR” from Abjad-Kids alone: it is short educational vocabulary, not Omani conversation.
- Speaker recognition, emotion surveillance, or child profiling.
- Sending raw child audio to third-party telemetry by default.
Verification
make checkThe CI suite runs without downloading models or child data. GPU/model integration should additionally be run on the target training image before a release.
License
Code is MIT. Model and dataset releases must separately document every upstream model/dataset license and consent restriction. A code license never grants permission to publish child audio.
