arz
Datasets
All datasets matching “arz”african_speech_dataset_arzArzEn-MultiGenre
ArzEn-MultiGenre: A Comprehensive Parallel Dataset
Overview
ArzEn-MultiGenre is a distinctive parallel dataset that encompasses a diverse collection of Egyptian Arabic content. This collection includes song lyrics, novels, and TV show subtitles, all of which have been meticulously translated and aligned with their English counterparts. The dataset serves as an invaluable tool for various linguistic and computational applications.
Published: 28 December 2023Version: 3DOI:… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/ArzEn-MultiGenre.arz-en-parallel-corpus
Egyptian Arabic - English Parallel Corpus 🇪🇬✨🇬🇧
Dataset Description
This dataset is a cleaned and filtered merge of multiple Egyptian Arabic - English parallel corpora, containing ~27,000 aligned sentence pairs. It’s designed for researchers and developers working on machine translation, speech translation, and other NLP tasks involving Egyptian Arabic and English.
Sources 📚
This dataset integrates and refines data from the following publicly available… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimAmin/arz-en-parallel-corpus.news-commentary-eng-arz
Dataset details
In this version of the News Commentary dataset, Standard Arabic text segments are converted into Egyptian Arabic (ARZ) using GPT-4.1-Mini.
We calculated the semantic similarity between the English source and the Egyptian Arabic target and selected the 500 segments with the highest scores for the test split,
while the train split comprises the remaining 83.2K segments.
∙ Dataset columns
"english": original English text
"arabic": original Standard… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/news-commentary-eng-arz.dataset-feature-extractionMBZUAI_ArzEn_wav
