Team Ai
Datasetpublic

irfan-ahmad/AURA-Multi_Label_Classification

AURA-Classification (Multi-Label Version) Dataset Description The AURA (App User Review in Arabic) Classification dataset is a collection of 2,900 Arabic-language app reviews collected from various mobile applications. This dataset is designed for multi-label text classification, where each review can belong to multiple classes simultaneously. Each review in the dataset was independently annotated by five different annotators. To construct the multi-label version… See the full description on the dataset page: https://huggingface.co/datasets/irfan-ahmad/AURA-Multi_Label_Classification.

sourceHugging Facemitupdated 10mo agoView on Hugging Face
1likes51downloads
Dataset Card

AURA-Classification (Multi-Label Version)

Dataset Description

The AURA (App User Review in Arabic) Classification dataset is a collection of 2,900 Arabic-language app reviews collected from various mobile applications. This dataset is designed for multi-label text classification, where each review can belong to multiple classes simultaneously. Each review in the dataset was independently annotated by five different annotators. To construct the multi-label version of the dataset, a review is assigned to a given class if at least one annotator labeled it with that class. As a result, a single review may be associated with up to four labels.

Features

The dataset includes the following fields:

  • —review: The text of the review in Arabic.
  • —appName: The name of the application being reviewed.
  • —platform: The platform (iOS or Android) where the review was posted.
  • —judg_one, judg_two, judg_three, judg_four, judg_five: The labels assigned by each of the five annotators.

The possible classification labels are:

  • —bug_report: The review highlights a bug or issue in the app.
  • —improvement_request: The review suggests improvements or features.
  • —rating: The review expresses a general rating or opinion.
  • —others: Miscellaneous or uncategorized reviews.

Dataset Statistics

  • —Total Reviews: 2,900
  • —Platforms: iOS, Android
  • —Applications: Multiple apps from diverse categories.
  • —Labels Distribution: Four possible classes in a multi-label setting (each review may have multiple labels).

Example Entry

json
{
  "review": "الرجاء تحديث التطبيق لانه اذا سويت البلاغ في النهاية يخرج من الطبيق ولا يتم 
إرسال البلاغ",
  "appname": "تقديم بلاغ مخالفة تجارية",
  "platform": "android",
  "judg_one": "bug_report",
  "judg_two": "improvement_request",
  "judg_three": "bug_report",
  "judg_four": "improvement_request",
  "judg_five": "bug_report"
}

Use Cases

This dataset is suitable for:

  • —Multi-label text classification of app reviews.
  • —Issue identification and requirements elicitation.
  • —Multilingual NLP research focused on Arabic.
  • —Fine-tuning models for app review classification.

Citation

If you use this dataset, please cite it as:

@article{Aljeezani2025arabic,
  title={Arabic App Reviews: Analysis and Classification},
  author={Aljeezani, Othman and  Alomari, Dorieh and Ahmad, Irfan},
  journal={ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP)},
  volume={24},
  number={2},
  pages={1--28},
  year={2025},
  publisher={ACM New York, NY, USA},
}

@article{alansari2025multilabel,
  title   = {Multi-Label Classification of Arabic App Reviews with Data Augmentation and Explainable AI},
  author  = {Alansari, Aisha and Alomari, Dorieh and Mahmood, Sajjad and Ahmad, Irfan},
  journal = {Arabian Journal for Science and Engineering},
  year    = {to-appear},
  publisher = {Springer}
}

License

This dataset is shared under the MIT License. Please ensure appropriate attribution when using this dataset.

Acknowledgments

Special thanks to the contributors and reviewers who made this dataset possible.

Contact

For questions or feedback, please reach out to the corresponding author (Irfan Ahmad).

irfan-ahmad/AURA-Multi_Label_Classification · Team Ai