Team Ai
Modelpublic

seniruk/whisper-small-si-cpu

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes50downloads
Model Card

Hi, Iโ€™m Seniru Epasinghe ๐Ÿ‘‹

Iโ€™m a AI undergraduate and an AI enthusiast, working on machine learning projects and open-source contributions. I enjoy exploring AI pipelines, natural language processing, and building tools that make development easier.


๐ŸŒ Connect with me

![Hugging Face](https://huggingface.co/seniruk)    ![Medium](https://medium.com/@senirukepasinghe)    ![LinkedIn](https://www.linkedin.com/in/seniru-epasinghe-b34b86232/)    ![GitHub](https://github.com/seth2k2)

Sinscribe

This model is a fine-tuned version of openai/whisper-small on the Sinhala CSV + FLACs dataset. It achieves the following results on the evaluation set:

  • โ€”Loss: 0.0829
  • โ€”Wer: 29.1387

Intended uses & limitations

Can be used for Sinhala speech to text conversions. Make sure to input noise low audio to the model, to get the best outcome.

Training and evaluation data

Trained on the custom dataset - seniruk/sinscribe-sinhala-stt

Trained on above final dataset with 2 epochs on a device with below spec for 41:00:59 hours

  • โ€”16GB RAM
  • โ€”NVIDIA GeForce RTX 4060 Ti 16GB
  • โ€”12th Gen Intel(R) Core(TM) i5-12400F (2.50 GHz) - 64-bit

Training results

Training LossEpochStepValidation LossWer
0.18710.110210000.183451.9170
0.14290.220420000.151744.7541
0.13450.330730000.133641.0627
0.11830.440940000.123738.6625
0.1140.551150000.115136.9654
0.10560.661360000.108035.2670
0.09680.771570000.103734.4457
0.10110.881780000.098633.2741
0.09710.992090000.096132.7147
0.07131.1022100000.094732.0250
0.07061.2124110000.094032.0766
0.06911.3226120000.090731.2485
0.06841.4328130000.089330.9512
0.07181.5430140000.087530.3592
0.06421.6533150000.085930.0388
0.06671.7635160000.084229.5840
0.06671.8737170000.083529.3193
0.06771.9839180000.082929.1387

Inferencing

python
import torch
import soundfile
import torchaudio
from transformers import WhisperForConditionalGeneration, WhisperProcessor

device = "cpu"
torchaudio.set_audio_backend("soundfile") 

model = WhisperForConditionalGeneration.from_pretrained("seniruk/whisper-small-si-cpu").to(device)
processor = WhisperProcessor.from_pretrained("seniruk/whisper-small-si-cpu")

def transcribe(audio_path):
    try:
        if audio_path is None:
            return "No audio received. Please record something."

        waveform, sample_rate = torchaudio.load(audio_path)

        if waveform.shape[0] > 1:
            waveform = waveform.mean(dim=0, keepdim=True)

        if sample_rate != 16000:
            resampler = torchaudio.transforms.Resample(orig_freq=sample_rate, new_freq=16000)
            waveform = resampler(waveform)
            sample_rate = 16000

        waveform = waveform.squeeze().numpy()

        inputs = processor(waveform, sampling_rate=sample_rate, return_tensors="pt").input_features.to("cpu")
        predicted_ids = model.generate(inputs)
        transcription = processor.batch_decode(predicted_ids, skip_special_tokens=True)[0]

        return transcription

    except Exception as e:
        return f"Error during transcription: {e}"

print(transcribe('audio.wav'))
Gradio UI
python
import torch
import soundfile
import torchaudio
from transformers import WhisperForConditionalGeneration, WhisperProcessor
import gradio as gr

device = "cpu"
torchaudio.set_audio_backend("soundfile")

model = WhisperForConditionalGeneration.from_pretrained("seniruk/whisper-small-si-cpu").to(device)
processor = WhisperProcessor.from_pretrained("seniruk/whisper-small-si-cpu")

MAX_DURATION_SECONDS = 30  # Limit: 30 seconds of audio

def transcribe(audio_path):
    try:
        if audio_path is None:
            return "No audio received. Please record or upload a file."

        # Load audio
        waveform, sample_rate = torchaudio.load(audio_path)

        # Convert to mono
        if waveform.shape[0] > 1:
            waveform = waveform.mean(dim=0, keepdim=True)

        # Duration check
        duration = waveform.shape[1] / sample_rate
        if duration > MAX_DURATION_SECONDS:
            return f"Audio too long ({duration:.1f}s). Please use an audio clip shorter than {MAX_DURATION_SECONDS}s."

        # Resample if necessary
        if sample_rate != 16000:
            resampler = torchaudio.transforms.Resample(orig_freq=sample_rate, new_freq=16000)
            waveform = resampler(waveform)
            sample_rate = 16000

        waveform = waveform.squeeze().numpy()

        # Process through Whisper
        inputs = processor(
            waveform,
            sampling_rate=sample_rate,
            return_tensors="pt"
        ).input_features.to(device)

        with torch.no_grad():
            predicted_ids = model.generate(inputs)
            transcription = processor.batch_decode(predicted_ids, skip_special_tokens=True)[0]

        return transcription

    except Exception as e:
        return f"Error during transcription: {e}"


iface = gr.Interface(
    fn=transcribe,
    inputs=gr.Audio(sources=["microphone", "upload"], type="filepath", label="Record or Upload Audio"),
    outputs=gr.Textbox(label="Transcription"),
    title="Whisper Small Sinhala (CPU)",
    description=(
        "๐ŸŽ™๏ธ Sinhala speech-to-text using the fine-tuned Whisper Small model (Sinscribe).\n"
        "You can record or upload audio up to 30 seconds long."
    ),
)

iface.launch()

GPU runnable version of above model is below

Framework versions

  • โ€”Transformers 4.48.0
  • โ€”Pytorch 2.5.1+cu121
  • โ€”Datasets 3.6.0
  • โ€”Tokenizers 0.21.4