Team Ai
Datasetpublic

NabuxAi/buddi-asr-source-code

Buddi-ASR A privacy-aware training and deployment pipeline for a small Arabic/Gulf-Arabic + English child-speech model. فارسی · GPU training · Data governance · Experiment plan · Raspberry Pi Buddi-ASR turns multilingual Whisper Tiny or Base into a domain model for short child utterances, then exports the merged checkpoint to a quantized whisper.cpp artifact for Raspberry Pi, mobile, or server inference. This repository contains the reproducible pipeline, not fabricated model… See the full description on the dataset page: https://huggingface.co/datasets/NabuxAi/buddi-asr-source-code.

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes31downloads
Dataset Card

Buddi-ASR

A privacy-aware training and deployment pipeline for a small Arabic/Gulf-Arabic + English child-speech model.

فارسی · GPU training · Data governance · Experiment plan · Raspberry Pi

Buddi-ASR turns multilingual Whisper Tiny or Base into a domain model for short child utterances, then exports the merged checkpoint to a quantized whisper.cpp artifact for Raspberry Pi, mobile, or server inference.

This repository contains the reproducible pipeline, not fabricated model weights. A real Buddi-ASR-Tiny-ArGulf-v1 checkpoint must be trained on licensed/consented audio and accepted only after speaker-disjoint evaluation.

mermaid
flowchart TD
    A["Licensed + consented audio"] --> B["PII-safe manifests"]
    B --> C["Speaker-disjoint split"]
    C --> D["Whisper Tiny/Base adaptation"]
    D --> E["WER/CER by cohort"]
    E --> F["Merge + Q5 export"]
    F --> G["Raspberry Pi / mobile"]

What is implemented

  • —Abjad-Kids download, 144-label audit, Romanized-label to Arabic conversion, and WAV duration probing.
  • —HMAC pseudonymization, anonymized processed filenames, consent-gated local CSV import, and PII-blocking manifest validation.
  • —Deterministic label-aware splitting with strict speaker isolation.
  • —Weighted dataset mixing without silently duplicating child utterances.
  • —Whisper Tiny/Base LoRA and full fine-tuning with Arabic/English per-example prompt tokens, waveform augmentation, early stopping, and reproducible configs.
  • —One-command CUDA training for RunPod/servers with VRAM-aware batch sizing, strict preflight, multi-GPU launch, automatic checkpoint resume, and runtime provenance.
  • —Normalized micro-averaged WER/CER overall and by language, dialect, age group, source, and category.
  • —Safe LoRA merge, official whisper.cpp conversion, Q4/Q5/Q8 quantization, artifact metadata, and SHA-256 checksum.
  • —Dependency-free unit tests for all data-critical paths and a hardened Raspberry Pi service example.

Important dataset finding

The current Abjad-Kids repository reports 40,646 rows and 144 observed labels, while the associated 2026 paper reports 46,397 samples and 141 classes. The upstream CSV files also lack a documented speaker column; this project groups speakers from filenames only as an explicit heuristic. A full CSV audit parsed every row and found 40,568 filename-grouped rows plus 78 session-only rows that require special review.

Do not publish final metrics until:

  1. 1.the seven ambiguous label mappings in data/label_maps/abjad_ar.json have been checked by listening;
  2. 2.session-only speaker groups have been resolved with anonymous overrides or acknowledged as a limitation;
  3. 3.train, validation, and test show zero speaker overlap.

Quick start

Python 3.10–3.12 is supported. Data utilities have no third-party runtime dependencies.

bash
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[data]'

Create a stable private salt. Losing or changing it changes speaker groups; committing it defeats pseudonymization.

bash
openssl rand -hex 32 > .buddi-speaker.salt

Download and prepare Abjad-Kids:

bash
buddi-asr download-abjad data/raw/abjad-kids

buddi-asr prepare-abjad \
  data/raw/abjad-kids \
  data/manifests/abjad.jsonl \
  --processed-audio-dir data/processed/abjad \
  --salt-file .buddi-speaker.salt \
  --audit-output reports/abjad-audit.json

buddi-asr validate data/manifests/abjad.jsonl --report reports/abjad-validation.json

Create speaker-disjoint sets:

bash
buddi-asr split data/manifests/abjad.jsonl data/manifests/splits \
  --train 0.8 --validation 0.1 --test 0.1 --seed 42

Install GPU training dependencies and run the Tiny LoRA baseline:

bash
./scripts/train_gpu.sh tiny-lora --smoke
./scripts/train_gpu.sh tiny-lora

The runner prepares Abjad-Kids when split manifests are absent. See the real GPU training guide before renting a GPU or adding private Omani data.

Evaluate exactly once on the held-out test set after experiment selection:

bash
buddi-asr evaluate \
  artifacts/tiny-lora/final \
  data/manifests/splits/test.jsonl \
  reports/tiny-lora-test

Merge the adapter and export Q5 for the Pi:

bash
buddi-asr merge-lora \
  openai/whisper-tiny \
  artifacts/tiny-lora/final \
  artifacts/tiny-merged

git clone https://github.com/openai/whisper.git vendor/openai-whisper
git clone https://github.com/ggml-org/whisper.cpp.git vendor/whisper.cpp
cmake -S vendor/whisper.cpp -B vendor/whisper.cpp/build -DGGML_BLAS=1
cmake --build vendor/whisper.cpp/build -j --config Release

buddi-asr export-whisper-cpp \
  artifacts/tiny-merged \
  vendor/whisper.cpp \
  vendor/openai-whisper \
  artifacts/pi \
  --name buddi-asr-tiny-argulf-v1 \
  --quantization q5_0

Add consented Omani/Gulf child speech

Use a private CSV with these required columns:

csv
audio,text,speaker_key,language,consent_basis,dialect,age_group,device,noise_level
audio/001.wav,بودي أبغى قصة,anonymous-local-001,ar,verified_parental_consent,gulf_omani,6-8,iphone,quiet

speaker_key is pseudonymized and never written to the output. Unknown columns are rejected to prevent accidental identity export.

bash
buddi-asr import-local private/omani.csv data/manifests/omani.jsonl \
  --processed-audio-dir data/processed/omani \
  --salt-file .buddi-speaker.salt \
  --source-name buddi_omani_child

Repository map

text
configs/                 Reproducible Tiny LoRA, Base full, and mixture configs
data/label_maps/         Audited Abjad Romanized-to-Arabic transcript map
deploy/systemd/          Hardened local-only Pi service
docs/                    Governance, experiment, model-card, and Pi guides
scripts/train_gpu.sh     One-command real CUDA training and recovery
src/buddi_asr/data/      Ingestion, manifest, split, mix, validation
src/buddi_asr/training/  Audio loading, augmentation, collator, trainer
src/buddi_asr/           CLI, metrics, evaluation, merge, export
tests/                   Dependency-free unit and privacy regression tests

Non-goals

  • —Training from scratch: Whisper pretraining is deliberately reused.
  • —Claiming “Gulf child ASR” from Abjad-Kids alone: it is short educational vocabulary, not Omani conversation.
  • —Speaker recognition, emotion surveillance, or child profiling.
  • —Sending raw child audio to third-party telemetry by default.

Verification

bash
make check

The CI suite runs without downloading models or child data. GPU/model integration should additionally be run on the target training image before a release.

License

Code is MIT. Model and dataset releases must separately document every upstream model/dataset license and consent restriction. A code license never grants permission to publish child audio.