ismaeeelxd/Egyptian-Arabic-Lectures
Egyptian Arabic Lectures Dataset The Egyptian Arabic Lectures dataset is a collection of transcribed audio clips (around 30 hours) extracted from educational lectures delivered in Egyptian Arabic (with mixed English technical terms, such as in Physics, IoT and Operating Systems etc.). It is designed to train, evaluate, and fine-tune Automatic Speech Recognition models for the Egyptian dialect, specifically in educational and academic CS contexts. Alongside the audio and text… See the full description on the dataset page: https://huggingface.co/datasets/ismaeeelxd/Egyptian-Arabic-Lectures.
Egyptian Arabic Lectures Dataset
The Egyptian Arabic Lectures dataset is a collection of transcribed audio clips (around 30 hours) extracted from educational lectures delivered in Egyptian Arabic (with mixed English technical terms, such as in Physics, IoT and Operating Systems etc.). It is designed to train, evaluate, and fine-tune Automatic Speech Recognition models for the Egyptian dialect, specifically in educational and academic CS contexts.
Alongside the audio and text pairs, the dataset provides comprehensive metadata for audio quality and characteristics, including Signal-to-Noise Ratio (SNR), spectral flatness, and speech ratio, making it highly useful for audio processing and filtering.
Dataset Description
This dataset has been developed during our graduation project to fine tune OpenAI's whisper mainly but it can be used to fine tune other ASR models aswell. All samples are resampled at 16khz and segmented into 30 seconds chunks. I'll dive into the dataset curation process since it's important to be noted. The lectures are publicly available on YouTube. All transcriptions are generated by Gemini and reviewed by us to ensure the accuracy of the transcriptions. Transcriptions also have been processed to normalize all slang talk to be unified across the dataset such as "برضو" to be "بردو" etc..
- Curated by: Alhasan Muhammed, Eslam Khaled, Abdelaleam Ehab, Basel Muhammed, Basmala Abdelwahab and me
- Language: Code Switched Egyptian Arabic
- License: MIT
Dataset Sources
- Recorded Lectures on YouTube.
- Recorded Lectures on our university's LMS.
Usage
You can load the dataset directly using ``datasets`` library from Hugging face:
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("ismaeeelxd/Egyptian-Arabic-Lectures")
# Access the first example in the train split
sample = dataset["train"][0]
print("Transcription:", sample["transcription"])
print("Subject:", sample["subject_id"])
print("Audio Array:", sample["audio"]["array"])Audio processing note
The audio files are provided in .mp3 format and are resampled at 16kHz. When feeding this data into ASR models like Whisper or Wav2Vec2, make sure to resample the audio if your specific model requires a different input sampling rate.
