Dasool/DC_inside_comments
DC_inside_comments This dataset contains 110,000 raw comments collected from DC Inside. It is intended for unsupervised learning or pretraining purposes. Dataset Summary Data Type: Unlabeled raw comments Number of Examples: 110,000 Source: DC Inside Related Dataset For labeled data and multi-task annotated examples, please refer to the KoMultiText dataset. How to Load the Dataset from datasets import load_dataset # Load the… See the full description on the dataset page: https://huggingface.co/datasets/Dasool/DC_inside_comments.
DCinsidecomments
This dataset contains 110,000 raw comments collected from DC Inside. It is intended for unsupervised learning or pretraining purposes.
Dataset Summary
- Data Type: Unlabeled raw comments
- Number of Examples: 110,000
- Source: DC Inside
Related Dataset
For labeled data and multi-task annotated examples, please refer to the KoMultiText dataset.
How to Load the Dataset
from datasets import load_dataset
# Load the unlabeled dataset
dataset = load_dataset("Dasool/DC_inside_comments")
print(dataset)Citation
@misc{choi2023largescale,
title={Large-Scale Korean Text Dataset for Classifying Biased Speech in Real-World Online Services},
author={Dasol Choi and Jooyoung Song and Eunsun Lee and Jinwoo Seo and Heejune Park and Dongbin Na},
year={2023},
eprint={2310.04313},
archivePrefix={arXiv},
primaryClass={cs.CL}
}Contact
- dasolchoi@yonsei.ac.kr
