google-research-datasets/discofuse
Dataset Card for "discofuse" Dataset Summary DiscoFuse is a large scale dataset for discourse-based sentence fusion. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances discofuse-sport Size of downloaded dataset files: 4.33 GB Size of the generated dataset: 15.04 GB Total amount of disk used: 19.36 GB An example of… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/discofuse.
Dataset Card for "discofuse"
Table of Contents
- Dataset Description
- Dataset Summary
- Supported Tasks and Leaderboards
- Languages
- Dataset Structure
- Data Instances
- Data Fields
- Data Splits
- Dataset Creation
- Curation Rationale
- Source Data
- Annotations
- Personal and Sensitive Information
- Considerations for Using the Data
- Social Impact of Dataset
- Discussion of Biases
- Other Known Limitations
- Additional Information
- Dataset Curators
- Licensing Information
- Citation Information
- Contributions
Dataset Description
- Repository: https://github.com/google-research-datasets/discofuse
- Paper: DiscoFuse: A Large-Scale Dataset for Discourse-Based Sentence Fusion
- Point of Contact: More Information Needed
- Size of downloaded dataset files: 6.04 GB
- Size of the generated dataset: 21.55 GB
- Total amount of disk used: 27.59 GB
Dataset Summary
DiscoFuse is a large scale dataset for discourse-based sentence fusion.
Supported Tasks and Leaderboards
Languages
Dataset Structure
Data Instances
discofuse-sport
- Size of downloaded dataset files: 4.33 GB
- Size of the generated dataset: 15.04 GB
- Total amount of disk used: 19.36 GB
An example of 'train' looks as follows.
{
"coherent_first_sentence": "Four LPr and three LC2000r HP Netservers handle customer management and web server functions .",
"coherent_second_sentence": "Finally , an HP Netserver LT6000r hosts i2 Demand Planner and i2 Collaboration Planner .",
"connective_string": "finally ,",
"discourse_type": "PAIR_CONN",
"has_coref_type_nominal": 0.0,
"has_coref_type_pronoun": 0.0,
"incoherent_first_sentence": "Four LPr and three LC2000r HP Netservers handle customer management and web server functions .",
"incoherent_second_sentence": "An HP Netserver LT6000r hosts i2 Demand Planner and i2 Collaboration Planner ."
}discofuse-wikipedia
- Size of downloaded dataset files: 1.72 GB
- Size of the generated dataset: 6.51 GB
- Total amount of disk used: 8.23 GB
An example of 'validation' looks as follows.
{
"coherent_first_sentence": "Four LPr and three LC2000r HP Netservers handle customer management and web server functions .",
"coherent_second_sentence": "Finally , an HP Netserver LT6000r hosts i2 Demand Planner and i2 Collaboration Planner .",
"connective_string": "finally ,",
"discourse_type": "PAIR_CONN",
"has_coref_type_nominal": 0.0,
"has_coref_type_pronoun": 0.0,
"incoherent_first_sentence": "Four LPr and three LC2000r HP Netservers handle customer management and web server functions .",
"incoherent_second_sentence": "An HP Netserver LT6000r hosts i2 Demand Planner and i2 Collaboration Planner ."
}Data Fields
The data fields are the same among all splits.
discofuse-sport
connective_string: astringfeature.discourse_type: astringfeature.coherent_second_sentence: astringfeature.has_coref_type_pronoun: afloat32feature.incoherent_first_sentence: astringfeature.incoherent_second_sentence: astringfeature.has_coref_type_nominal: afloat32feature.coherent_first_sentence: astringfeature.
discofuse-wikipedia
connective_string: astringfeature.discourse_type: astringfeature.coherent_second_sentence: astringfeature.has_coref_type_pronoun: afloat32feature.incoherent_first_sentence: astringfeature.incoherent_second_sentence: astringfeature.has_coref_type_nominal: afloat32feature.coherent_first_sentence: astringfeature.
Data Splits
Dataset Creation
Curation Rationale
Source Data
Initial Data Collection and Normalization
Who are the source language producers?
Annotations
Annotation process
Who are the annotators?
Personal and Sensitive Information
Considerations for Using the Data
Social Impact of Dataset
Discussion of Biases
Other Known Limitations
Additional Information
Dataset Curators
Licensing Information
The data is licensed under Creative Commons Attribution-ShareAlike 3.0 license.
Citation Information
@InProceedings{GevaEtAl2019,
title = {DiscoFuse: A Large-Scale Dataset for Discourse-Based Sentence Fusion},
author = {Geva, Mor and Malmi, Eric and Szpektor, Idan and Berant, Jonathan},
booktitle = {Proceedings of the 2019 Annual Conference of the North American Chapter of the Association for Computational Linguistics},
note = {arXiv preprint arXiv:1902.10526},
year = {2019}
}Contributions
Thanks to @thomwolf, @patrickvonplaten, @mariamabarham, @lewtun for adding this dataset.
