AmazonScience/xtr-wiki_qa
Xtr-WikiQA Dataset Summary Xtr-WikiQA is an Answer Sentence Selection (AS2) dataset in 9 non-English languages, proposed in our paper accepted at ACL 2023 (Findings): Cross-Lingual Knowledge Distillation for Answer Sentence Selection in Low-Resource Languages. This dataset is based on an English AS2 dataset, WikiQA (Original, Hugging Face). For translations, we used Amazon Translate. Languages Arabic (ar) Spanish (es) French (fr) German (de) Hindi… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/xtr-wiki_qa.
Xtr-WikiQA
Table of Contents
- Dataset Card Creation Guide
- Table of Contents
- Dataset Description
- Dataset Summary
- Languages
- Dataset Structure
- Data Instances
- Data Fields
- Data Splits
- Dataset Creation
- Source Data
- Additional Information
- Licensing Information
- Citation Information
- Contributions
Dataset Description
- Homepage: Amazon Science
- Paper: Cross-Lingual Knowledge Distillation for Answer Sentence Selection in Low-Resource Languages
- Point of Contact: Yoshitomo Matsubara
Dataset Summary
*Xtr-WikiQA is an Answer Sentence Selection (AS2) dataset in 9 non-English languages, proposed in our paper accepted at ACL 2023 (Findings): [Cross-Lingual Knowledge Distillation for Answer Sentence Selection in Low-Resource Languages*](https://aclanthology.org/2023.findings-acl.885/). This dataset is based on an English AS2 dataset, WikiQA (Original, Hugging Face). For translations, we used Amazon Translate.
Languages
- Arabic (ar)
- Spanish (es)
- French (fr)
- German (de)
- Hindi (hi)
- Italian (it)
- Japanese (ja)
- Dutch (nl)
- Portuguese (pt)
File location: `tsv/`
Dataset Structure
Data Instances
This is an example instance from the Arabic training split of Xtr-WikiQA dataset.
{
"QuestionID": "Q1",
"Question": "كيف تتشكل الكهوف الجليدية؟",
"DocumentID": "D1",
"DocumentTitle": "كهف جليدي",
"SentenceID": "D1-0",
"Sentence": "كهف جليدي مغمور جزئيًا على نهر بيريتو مورينو الجليدي.",
"Label": 0
}All the translated instances in tsv files are listed in the same order of the original (native) instances in the WikiQA dataset.
For example, the 2nd instance in `tsv/ar-train.tsv` (Arabic-translated from English) corresponds to the 2nd instance in `WikiQA-train.tsv` (English).
Data Fields
Each instance (a QA pair) consists of the following fields:
QuestionID: Question ID (str)Question: Question to be answered (str)DocumentID: Document ID (str)DocumentTitle: Document title (str)SentenceID: Answer sentence in the document (str)Sentence: Answer sentence in the document (str)Label: Label that indicates the answer sentence correctly answers the question (int, 1: correct, 0: incorrect)
Data Splits
See our paper for more details about the statistics of the datasets.
Dataset Creation
Source Data
The source of Xtr-WikiQA dataset is WikiQA.
Additional Information
Licensing Information
CDLA-Permissive-2.0
Citation Information
@inproceedings{gupta2023cross-lingual,
title={{Cross-Lingual Knowledge Distillation for Answer Sentence Selection in Low-Resource Languages}},
author={Gupta, Shivanshu and Matsubara, Yoshitomo and Chadha, Ankit and Moschitti, Alessandro},
booktitle={Findings of the Association for Computational Linguistics: ACL 2023},
pages={14078--14092},
year={2023}
}Contributions
- Shivanshu Gupta
- Yoshitomo Matsubara
- Ankit Chadha
- Alessandro Moschitti
