SarahUssama/sada-arabic-test-dataset-sample
๐ฃ๏ธ Arabic Dialect Segmented Speech Dataset (SADA2022 Subset) This dataset contains segmented Arabic speech samples from the SADA2022 corpus, annotated by dialect, gender, age group, speaking rate, environmental condition, and includes ground truth transcriptions. It is intended to support research and applications in Arabic dialect classification, automatic speech recognition (ASR), and spoken language understanding. ๐ Dataset Structure Audio segments areโฆ See the full description on the dataset page: https://huggingface.co/datasets/SarahUssama/sada-arabic-test-dataset-sample.
๐ฃ๏ธ Arabic Dialect Segmented Speech Dataset (SADA2022 Subset)
This dataset contains segmented Arabic speech samples from the SADA2022 corpus, annotated by dialect, gender, age group, speaking rate, environmental condition, and includes ground truth transcriptions.
It is intended to support research and applications in Arabic dialect classification, automatic speech recognition (ASR), and spoken language understanding.
๐ Dataset Structure
- Audio segments are stored in
.wavformat - Accompanied by a CSV file (
dataset.csv) with rich metadata - Organized into dialect folders:
Egyptian/Saudi/
๐Dataset Statistics
๐ Data Insights
๐ข Total Segments: 496
๐ค Speaker Demographics
๐ Recording Conditions
๐ Audio Properties
๐ Ground Truth Transcriptions
Each audio segment is paired with an Arabic text transcription (GroundTruthText column). There are 493 unique phrases, including:
ุงูุณูุงู ุนููููโ 3 occurrencesุฅูุญู ูุง ุนู ุฏุฉโ 2 occurrencesู ุน ุงูุณูุงู ุฉโ 2 occurrences- As well as longer, spontaneous utterances like:
"ุงูู ููุง ููู ู ุงุจุดุฑ ุฅู ุดุงุก ุงูููุ ุงูู ุงููููุฉ ูู ุงูู ุฒุฑุนุฉ ุฅู ุดุงุก ุงูููุ ูุดููู ุนูู ุฎูุฑ ูู ุฃู ุงู ุงููู ู ุน ุงูุณูุงู ุฉ"
๐ ๏ธ How to Use
You can load the dataset using:
from datasets import load_dataset
ds = load_dataset("SarahUssama/sada-arabic-test-dataset-sample")
