Team Ai
Datasetpublic

nazimali/quran-question-answer-context

Dataset Card for "quran-question-answer-context" Dataset Summary Translated the original dataset from Arabic to English and added the Surah ayahs to the context column. Usage from datasets import load_dataset dataset = load_dataset("nazimali/quran-question-answer-context") DatasetDict({ train: Dataset({ features: ['q_id', 'question', 'answer', 'q_word', 'q_topic', 'fine_class', 'class', 'ontology_concept', 'ontology_concept2', 'source'… See the full description on the dataset page: https://huggingface.co/datasets/nazimali/quran-question-answer-context.

sourceHugging Facecc-by-4.0updated 2y agoView on Hugging Face
11likes97downloads
Dataset Card

Dataset Card for "quran-question-answer-context"

Dataset Summary

Translated the original dataset from Arabic to English and added the Surah ayahs to the context column.

Usage

python
from datasets import load_dataset

dataset = load_dataset("nazimali/quran-question-answer-context")
python
DatasetDict({
    train: Dataset({
        features: ['q_id', 'question', 'answer', 'q_word', 'q_topic', 'fine_class', 'class', 'ontology_concept', 'ontology_concept2', 'source', 'q_src_id', 'quetion_type', 'chapter_name', 'chapter_no', 'verse', 'answer_en', 'class_en', 'fine_class_en', 'ontology_concept2_en', 'ontology_concept_en', 'q_topic_en', 'q_word_en', 'question_en', 'chapter_name_en', 'verse_list', 'context', 'context_data', 'context_missing_verses'],
        num_rows: 1224
    })
})

Translation Info

  1. 1.Translated the Arabic questions/concept columns to English with Helsinki-NLP/opus-mt-ar-en
  2. 2.Used en-yusufali translations for ayas M-AI-C/quran-en-tafssirs
  3. 3.Renamed Surahs with kheder/quran
  4. 4.Added the ayahs that helped answer the questions
  5. 5.Split the ayah columns string into a list of integers
  6. 6.Concactenated the Surah:Ayah pairs into a sentence to the context column

Columns with the suffix _en contain the translations of the original columns.

TODO

The context column has some null values that needs to be investigated and fixed

Initial Data Collection

The original dataset is from [Annotated Corpus of Arabic Al-Quran Question and Answer](https://archive.researchdata.leeds.ac.uk/464/)

Licensing Information

Original dataset license: Creative Commons Attribution 4.0 International (CC BY 4.0)

Contributions

Original paper authors: Alqahtani, Mohammad and Atwell, Eric (2018) Annotated Corpus of Arabic Al-Quran Question and Answer. University of Leeds. https://doi.org/10.5518/356