Team Ai
Modelpublic

gitmbio/test

sourceHugging Facemitupdated 3y agoView on Hugging Face
0likes5downloads
README.md92 linesDownload Raw Back to root
1---2license: mit3tags:4- audio5- automatic-speech-recognition6widget:7- example_title: sample 18  src: https://huggingface.co/bangla-speech-processing/BanglaASR/resolve/main/mp3/common_voice_bn_31515636.mp39- example_title: sample 210  src: https://huggingface.co/bangla-speech-processing/BanglaASR/resolve/main/mp3/common_voice_bn_31549899.mp311- example_title: sample 312  src: https://huggingface.co/bangla-speech-processing/BanglaASR/resolve/main/mp3/common_voice_bn_31617644.mp313pipeline_tag: automatic-speech-recognition14---15 16Bangla ASR model which was trained Bangla Mozilla Common Voice Dataset. This is Fine-tuning Whisper model using Bangla mozilla common voice dataset. 17For training this model used 40k training and 7k Validation of around 400 hours of data. We trained 12000 steps and get word 18error rate 4.58%. This model was whisper small[244 M] variant model.19 20 21```py22 23import os24import librosa25import torch26import torchaudio27import numpy as np28 29from transformers import WhisperTokenizer30from transformers import WhisperProcessor31from transformers import WhisperFeatureExtractor32from transformers import WhisperForConditionalGeneration33 34device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')35 36mp3_path = "https://huggingface.co/bangla-speech-processing/BanglaASR/resolve/main/mp3/common_voice_bn_31515636.mp3"37 38model_path = "bangla-speech-processing/BanglaASR"39 40 41feature_extractor = WhisperFeatureExtractor.from_pretrained(model_path)42tokenizer = WhisperTokenizer.from_pretrained(model_path)43processor = WhisperProcessor.from_pretrained(model_path)44model = WhisperForConditionalGeneration.from_pretrained(model_path).to(device)45 46 47speech_array, sampling_rate = torchaudio.load(mp3_path, format="mp3")48speech_array = speech_array[0].numpy()49speech_array = librosa.resample(np.asarray(speech_array), orig_sr=sampling_rate, target_sr=16000)50input_features = feature_extractor(speech_array, sampling_rate=16000, return_tensors="pt").input_features51 52# batch = processor.feature_extractor.pad(input_features, return_tensors="pt")53predicted_ids = model.generate(inputs=input_features.to(device))[0]54 55 56transcription = processor.decode(predicted_ids, skip_special_tokens=True)57 58print(transcription)59 60```61 62 63# Dataset64Used Mozilla common voice dataset around 400 hours data both training[40k] and validation[7k] mp3 samples.65For more information about dataser please [click here](https://commonvoice.mozilla.org/bn/datasets)66 67# Training Model Information68 69 70| Size | Layers | Width | Heads | Parameters | Bangla-only | Training Status |71| ------------- | ------------- | --------    |--------    | ------------- | ------------- | --------    |72tiny   | 4  |384  | 6   | 39 M 	| X |  X73base   | 6 	|512  | 8 	|74 M 	| X	|  X74small  | 12 |768  | 12 	|244 M 	| ✓ |  ✓ 75medium | 24 |1024 | 16 	|769 M 	| X |  X76large  | 32 |1280 | 20 	|1550 M | X |  X77 78# Evaluation79 80Word Error Rate 4.58 %81 82For More please check the [github](https://github.com/saiful9379/BanglaASR/tree/main)83 84```85@misc{BanglaASR ,86  title={Transformer Based Whisper Bangla ASR Model},87  author={Md Saiful Islam},88  howpublished={},89  year={2023}90}91```92