Team Ai
Datasetpublic

AmazonScience/tydi-as2

TyDi-AS2 Dataset Summary TyDi-AS2 and Xtr-TyDi-AS2 are multilingual Answer Sentence Selection (AS2) datasets comprising 8 diverse languages, proposed in our paper accepted at ACL 2023 (Findings): Cross-Lingual Knowledge Distillation for Answer Sentence Selection in Low-Resource Languages. Both the datasets were created from TyDi-QA, a multilingual question-answering dataset. TyDi-AS2 was created by converting the QA instances in TyDi-QA to AS2 instances (see… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/tydi-as2.

sourceHugging Facecdla-permissive-2.0updated 3y agoView on Hugging Face
1likes305downloads
Dataset Card

TyDi-AS2

Table of Contents

Dataset Description

Dataset Summary

*TyDi-AS2 and Xtr-TyDi-AS2 are multilingual Answer Sentence Selection (AS2) datasets comprising 8 diverse languages, proposed in our paper accepted at ACL 2023 (Findings): [Cross-Lingual Knowledge Distillation for Answer Sentence Selection in Low-Resource Languages*](https://aclanthology.org/2023.findings-acl.885/). Both the datasets were created from TyDi-QA, a multilingual question-answering dataset. TyDi-AS2 was created by converting the QA instances in TyDi-QA to AS2 instances (see Dataset Creation for details). Xtr-TyDi-AS2 was created by translating the non-English TyDi-AS2 instances to English and vise versa. For translations, we used Amazon Translate.

Languages

TyDi-AS2 (original)
  • —bn: Bengali
  • —en: English
  • —fi: Finnish
  • —id: Indonesian
  • —ja: Japanese
  • —ko: Korean
  • —ru: Russian
  • —sw: Swahili

File location: `jsonl/original/`

For non-English sets, we also have English-translated samples used for the cross-lingual knowledge distillation (CLKD) experiments in our paper.

File location: `jsonl/x-to-en/`

Xtr-TyDi-AS2 (translationese)

Xtr-TyDi-AS2 (X-translated TyDi-AS2) dataset consists of non-English AS2 instances translated from the English set of TyDi-AS2.

  • —bn: Bengali
  • —fi: Finnish
  • —id: Indonesian
  • —ja: Japanese
  • —ko: Korean
  • —ru: Russian
  • —sw: Swahili

File location: `jsonl/en-to-x/`

Dataset Structure

Data Instances

This is an example instance from the English training split of TyDi-AS2 dataset.

{
  "Question": "When was the Argentine Basketball Federation formed?",
  "Title": "History of the Argentina national basketball team",
  "Sentence": "The Argentina national basketball team represents Argentina in basketball international competitions, and is controlled by the Argentine Basketball Federation.",
  "Label": 0
}

For English-translated TyDi-AS2 dataset and Xtr-TyDi-AS2 dataset, the translated instances in JSONL files are listed in the same order of the original (native) instances in the original TyDi-AS2 dataset.

For example, the 2nd instance in `jsonl/x-to-en/en_from_bn-train.jsonl` (English-translated from Bengali) corresponds to the 2nd instance in `jsonl/original/bn-train.jsonl` (Bengali).

Similarly, the 2nd instance in `jsonl/en-to-x/bn_from_en-train.jsonl` (Bengali-translated from English) corresponds to the 2nd instance in `jsonl/original/en-train.jsonl` (English).

Data Fields

Each instance (a QA pair) consists of the following fields:

  • —Question: Question to be answered (str)
  • —Title: Document title (str)
  • —Sentence: Answer sentence in the document (str)
  • —Label: Label that indicates the answer sentence correctly answers the question (int, 1: correct, 0: incorrect)

Data Splits

**#Questions****#Sentences**
traindevtesttraindevtest
Bengali (bn)7,9782,0563161,376,432351,18637,465
English (en)6,7301,6869181,643,702420,899249,513
Finnish (fi)10,8592,7311,8701,567,695408,205298,093
Indonesian (id)9,3102,3391,355960,270236,07697,057
Japanese (ja)11,8482,9811,5043,183,037822,654444,106
Korean (ko)7,3541,9431,3891,558,191392,361199,043
Russian (ru)9,1872,2941,3953,190,650820,668367,595
Swahili (sw)8,3502,8501,8961,048,303269,89474,775

See our paper for more details about the statistics of the datasets.

Dataset Creation

Source Data

The source of TyDi-AS2 dataset is TyDi QA, which is a question answering dataset.

Annotations

Annotation process

TyDi QA is a QA dataset spanning questions from 11 typologically diverse languages. Each instance comprises a human-generated question, a single Wikipedia document as context, and one or more spans from the document containing the answer. To convert each instance into AS2 instances, we split the context document into sentences and heuristically identify the correct asnwer sentences using the annotated answer spans. To split documents, we use multiple different sentence tokenizers for the diverse languages and omit languages for which we could not find a suitable sentence tokenizer:

  1. 1.bltk for Bengali
  2. 2.blingfire for Swahili, Indonesian, and Korean
  3. 3.pysdb for English and Russian
  4. 4.nltk for Finnish
  5. 5.Konoha for Japanese
Who are the annotators?

Shivanshu Gupta converted TyDi QA to TyDi-AS2. Yoshitomo Matsubara translated non-English samples to English and vice versa for Xtr-TyDi-AS2 dataset Since sentence tokenization and identifying answer sentences can introduce errors, we conducted a manual validation of the AS2 datasets. For each language, we randomly selected 50 instances and verified the accuracy of the answer sentences through manual inspection. Our findings revealed that the answer sentences were accurate in 98% of the cases.

Additional Information

Dataset Curators

Shivanshu Gupta (@shivanshu)

Licensing Information

CDLA-Permissive-2.0

Citation Information

bibtex
@inproceedings{gupta2023cross-lingual,
  title={{Cross-Lingual Knowledge Distillation for Answer Sentence Selection in Low-Resource Languages}},
  author={Gupta, Shivanshu and Matsubara, Yoshitomo and Chadha, Ankit and Moschitti, Alessandro},
  booktitle={Findings of the Association for Computational Linguistics: ACL 2023},
  pages={14078--14092},
  year={2023}
}

Contributions