Team Ai
Modelpublic

eemberda/phcodeswitch-ceb-dvo

sourceHugging Facecc-by-nc-4.0updated 2mo agoView on Hugging Face
0likes152downloads
Model Card

phcodeswitch-ceb-dvo

Whisper small fine-tuned for automatic speech recognition (ASR) of Davao Cebuano, including English–Cebuano code-switching.

  • —Base model: openai/whisper-small
  • —Fine-tuned language: Cebuano (ceb), BCP-47 ceb
  • —Task: Automatic speech recognition (transcribe)
  • —Best WER: 20.86% (test split, step 1200)
  • —License: CC BY-NC 4.0 — non-commercial research / educational use only

How to use

Install dependencies:

bash
pip install --upgrade transformers torch librosa

Option 1 – pipeline (quickstart)

python
from transformers import pipeline
import librosa

pipe = pipeline(
    "automatic-speech-recognition",
    model="eemberda/phcodeswitch-ceb-dvo",
    device=0,  # use -1 for CPU
)

audio, sr = librosa.load("sample.wav", sr=16_000, mono=True)
result = pipe(
    audio,
    generate_kwargs={
        "language": "tl",   # Cebuano is not native to Whisper; Tagalog prompt works best
        "task": "transcribe",
        "num_beams": 5,
    },
)
print(result["text"])

Option 2 – manual inference with processor + model

python
from transformers import WhisperForConditionalGeneration, WhisperProcessor
import librosa
import torch

model_id = "eemberda/phcodeswitch-ceb-dvo"

processor = WhisperProcessor.from_pretrained(model_id)
model = WhisperForConditionalGeneration.from_pretrained(model_id)
device = "cuda" if torch.cuda.is_available() else "cpu"
model.to(device).eval()

audio, sr = librosa.load("sample.wav", sr=16_000, mono=True)
inputs = processor.feature_extractor(
    audio, sampling_rate=16_000, return_tensors="pt"
).input_features.to(device)

forced_decoder_ids = processor.get_decoder_prompt_ids(
    language="tl", task="transcribe"
)

with torch.no_grad():
    predicted_ids = model.generate(
        inputs,
        forced_decoder_ids=forced_decoder_ids,
        num_beams=5,
    )

transcription = processor.batch_decode(
    predicted_ids, skip_special_tokens=True
)[0].strip()
print(transcription)

Language notes

  • —Whisper has no native Cebuano language token. The model is fine-tuned on Cebuano audio but uses the Tagalog (`tl`) decoder prompt, which Whisper treats as the closest supported related language.
  • —The model also handles English and English–Cebuano code-switched speech.
  • —Audio is expected at 16 kHz mono (resampled automatically by librosa in the examples above).

Training details

  • —Base model: openai/whisper-small
  • —Optimizer: AdamW, learning rate 1e-5, warmup 100 steps
  • —Batch size 4 with gradient accumulation 4 (effective batch 16)
  • —Max steps 1800 (best checkpoint at step 1200), early stopping patience 3
  • —Mixed precision (fp16), gradient checkpointing, beam search (5) decoding

Limitations

  • —Trained on a small, community-contributed dataset; coverage of accents and vocabulary is limited.
  • —For non-commercial research and educational use only (CC BY-NC 4.0).
  • —Contributed speaker data must not be used for voice cloning, impersonation, or voice synthesis. See the project repository's compliance documents.