Team Ai
Datasetpublic

Dasool/DC_inside_comments

DC_inside_comments This dataset contains 110,000 raw comments collected from DC Inside. It is intended for unsupervised learning or pretraining purposes. Dataset Summary Data Type: Unlabeled raw comments Number of Examples: 110,000 Source: DC Inside Related Dataset For labeled data and multi-task annotated examples, please refer to the KoMultiText dataset. How to Load the Dataset from datasets import load_dataset # Load the… See the full description on the dataset page: https://huggingface.co/datasets/Dasool/DC_inside_comments.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes52downloads
Dataset Card

DCinsidecomments

This dataset contains 110,000 raw comments collected from DC Inside. It is intended for unsupervised learning or pretraining purposes.

Dataset Summary

  • —Data Type: Unlabeled raw comments
  • —Number of Examples: 110,000
  • —Source: DC Inside

Related Dataset

For labeled data and multi-task annotated examples, please refer to the KoMultiText dataset.

How to Load the Dataset

python
from datasets import load_dataset

# Load the unlabeled dataset
dataset = load_dataset("Dasool/DC_inside_comments")
print(dataset)

Citation

bibtex
@misc{choi2023largescale,
      title={Large-Scale Korean Text Dataset for Classifying Biased Speech in Real-World Online Services}, 
      author={Dasol Choi and Jooyoung Song and Eunsun Lee and Jinwoo Seo and Heejune Park and Dongbin Na},
      year={2023},
      eprint={2310.04313},
      archivePrefix={arXiv},
      primaryClass={cs.CL}
}

Contact

  • —dasolchoi@yonsei.ac.kr