Team Ai
Datasetpublic

SarahUssama/sada-arabic-test-dataset-sample

๐Ÿ—ฃ๏ธ Arabic Dialect Segmented Speech Dataset (SADA2022 Subset) This dataset contains segmented Arabic speech samples from the SADA2022 corpus, annotated by dialect, gender, age group, speaking rate, environmental condition, and includes ground truth transcriptions. It is intended to support research and applications in Arabic dialect classification, automatic speech recognition (ASR), and spoken language understanding. ๐Ÿ“ Dataset Structure Audio segments areโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/SarahUssama/sada-arabic-test-dataset-sample.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes41downloads
Dataset Card

๐Ÿ—ฃ๏ธ Arabic Dialect Segmented Speech Dataset (SADA2022 Subset)

This dataset contains segmented Arabic speech samples from the SADA2022 corpus, annotated by dialect, gender, age group, speaking rate, environmental condition, and includes ground truth transcriptions.

It is intended to support research and applications in Arabic dialect classification, automatic speech recognition (ASR), and spoken language understanding.


๐Ÿ“ Dataset Structure

  • โ€”Audio segments are stored in .wav format
  • โ€”Accompanied by a CSV file (dataset.csv) with rich metadata
  • โ€”Organized into dialect folders:
  • โ€”Egyptian/
  • โ€”Saudi/

๐Ÿ“ŠDataset Statistics

MetricValue
Total Segments496
LanguagesArabic (Egyptian, Saudi dialects)
Audio Format.wav
Sampling Rate16kHz

๐Ÿ“Š Data Insights

๐Ÿ”ข Total Segments: 496

๐Ÿ‘ค Speaker Demographics

AttributeCategoryCount
Age GroupAdult (ุจุงู„ุบ)350
Elderly (ูƒุจูŠุฑ ููŠ ุงู„ุณู†)100
Young Adult (ู…ุฑุงู‡ู‚)46
GenderMale397
Female99
DialectSaudi300
Egyptian196

๐ŸŒ Recording Conditions

EnvironmentCount
Clean (ู†ุธูŠู)217
Noisy (ุถูˆุถุงุก)154
Music (ู…ูˆุณูŠู‚ู‰)125

๐Ÿ• Audio Properties

AttributeCategoryCount/Value
Length TypeShort490
Long6
Speaking RateAverage320
Fast166
Slow10
Segment Length (seconds)Min0.5
Max27.51
Mean3.31

๐Ÿ“Œ Ground Truth Transcriptions

Each audio segment is paired with an Arabic text transcription (GroundTruthText column). There are 493 unique phrases, including:

  • โ€”ุงู„ุณู„ุงู… ุนู„ูŠูƒู… โ€” 3 occurrences
  • โ€”ุฅู„ุญู‚ ูŠุง ุนู…ุฏุฉ โ€” 2 occurrences
  • โ€”ู…ุน ุงู„ุณู„ุงู…ุฉ โ€” 2 occurrences
  • โ€”As well as longer, spontaneous utterances like:
"ุงูŠู‡ ูˆู„ุง ูŠู‡ู…ูƒ ุงุจุดุฑ ุฅู† ุดุงุก ุงู„ู„ู‡ุŒ ุงูŠู‡ ุงู„ู„ูŠู„ุฉ ููŠ ุงู„ู…ุฒุฑุนุฉ ุฅู† ุดุงุก ุงู„ู„ู‡ุŒ ู†ุดูˆููƒ ุนู„ู‰ ุฎูŠุฑ ููŠ ุฃู…ุงู† ุงู„ู„ู‡ ู…ุน ุงู„ุณู„ุงู…ุฉ"

๐Ÿ› ๏ธ How to Use

You can load the dataset using:

python
from datasets import load_dataset

ds = load_dataset("SarahUssama/sada-arabic-test-dataset-sample")