Team Ai
Datasetpublic

clips/VaccinChatNL

Dataset Card for VaccinChatNL Dataset Description Point of Contact: Jeska Buhmann Dataset Summary VaccinChatNL is a Flemish Dutch FAQ dataset on the topic of COVID-19 vaccinations in Flanders. It consists of 12,833 user questions divided over 181 answer labels, thus providing large groups of semantically equivalent paraphrases (a many-to-one mapping of user questions to answer labels). VaccinChatNL is the first Dutch many-to-one FAQ dataset of… See the full description on the dataset page: https://huggingface.co/datasets/clips/VaccinChatNL.

sourceHugging Facecc-by-4.0updated 4y agoView on Hugging Face
0likes364downloads
Dataset Card

Dataset Card for VaccinChatNL

Table of Contents

Dataset Description

<!-- - Homepage:

  • —Repository:
  • —Paper: [To be added]
  • —Leaderboard: -->
  • —Point of Contact: Jeska Buhmann

Dataset Summary

VaccinChatNL is a Flemish Dutch FAQ dataset on the topic of COVID-19 vaccinations in Flanders. It consists of 12,833 user questions divided over 181 answer labels, thus providing large groups of semantically equivalent paraphrases (a many-to-one mapping of user questions to answer labels). VaccinChatNL is the first Dutch many-to-one FAQ dataset of this size.

Supported Tasks and Leaderboards

  • —'text-classification': the dataset can be used to train a classification model for Dutch frequently asked questions on the topic of COVID-19 vaccination in Flanders.

Languages

Dutch (Flemish): the BCP-47 code for Dutch as generally spoken in Flanders (Belgium) is nl-BE.

Dataset Structure

Data Instances

For each instance, there is a string for the user question and a string for the label of the annotated answer. See the CLiPS / VaccinChatNL dataset viewer.

{"sentence1": "Waar kan ik de bijsluiters van de vaccins vinden?", "label": "faq_ask_bijsluiter"}

Data Fields

  • —sentence1: a string containing the user question
  • —label: a string containing the name of the intent (the answer class)

Data Splits

The VaccinChatNL dataset has 3 splits: train, valid, and test. Below are the statistics for the dataset.

Dataset SplitNumber of Labeled User Questions in Split
Train10,542
Validation1,171
Test1,170

Dataset Creation

<!-- ### Curation Rationale

[More Information Needed] -->

<!-- ### Source Data

[Perhaps a link to vaccinchat.be and some of the website that were used for information] -->

<!-- #### Initial Data Collection and Normalization

[More Information Needed]

Who are the source language producers?

[More Information Needed] -->

Annotations

Annotation process

Annotation was an iterative semi-automatic process. Starting from a very limited dataset with approximately 50 question-answer pairs (sentence1-label pairs) a text classification model was trained and implemented in a publicly available chatbot. When the chatbot was used, the predicted labels for the new questions were checked and corrected if necessary. In addition, new answers were added to the dataset. After each round of corrections, the model was retrained on the updated dataset. This iterative approach led to the final dataset containing 12,883 user questions divided over 181 answer labels.

Who are the annotators?

The VaccinChatNL data were annotated by members and students of CLiPS. All annotators have a background in Computational Linguistics.

Personal and Sensitive Information

The data are anonymized in the sense that a user question can never be traced back to a specific individual.

Considerations for Using the Data

<!-- ### Social Impact of Dataset

[More Information Needed] -->

Discussion of Biases

This dataset contains real user questions, including a rather large section (7%) of out-of-domain questions or remarks (label: nlufallback_). This class of user questions consists of ununderstandable questions, but also jokes and insulting remarks.

<!-- ### Other Known Limitations

[Perhaps some information of % of exact overlap between train and test set] -->

Additional Information

<!-- ### Dataset Curators

[More Information Needed] -->

<!-- ### Licensing Information

[More Information Needed] -->

Citation Information

@inproceedings{buhmann-etal-2022-domain,
    title = "Domain- and Task-Adaptation for {V}accin{C}hat{NL}, a {D}utch {COVID}-19 {FAQ} Answering Corpus and Classification Model",
    author = "Buhmann, Jeska and De Bruyn, Maxime and Lotfi, Ehsan and Daelemans, Walter",
    booktitle = "Proceedings of the 29th International Conference on Computational Linguistics",
    month = oct,
    year = "2022",
    address = "Gyeongju, Republic of Korea",
    publisher = "International Committee on Computational Linguistics",
    url = "https://aclanthology.org/2022.coling-1.312",
    pages = "3539--3549"
}

<!-- ### Contributions

Thanks to @github-username for adding this dataset. -->