datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
civil_comments
Dataset Card for "civil_comments"
Dataset Summary
The comments in this dataset come from an archive of the Civil Comments
platform, a commenting plugin for independent news sites. These public comments
were created from 2015 - 2017 and appeared on approximately 50 English-language
news sites across the world. When Civil Comments shut down in 2017, they chose
to make the public comments available in a lasting open archive to enable future
research. The original data… See the full description on the dataset page: https://huggingface.co/datasets/google/civil_comments.jigsaw-toxic-comment-classification-challenge
Dataset Description
You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior. The types of toxicity are:
toxic
severe_toxic
obscene
threat
insult
identity_hate
You must create a model which predicts a probability of each type of toxicity for each comment.
File descriptions
train.csv - the training set, contains comments with their binary labels
test.csv - the test set, you must predict the toxicity… See the full description on the dataset page: https://huggingface.co/datasets/thesofakillers/jigsaw-toxic-comment-classification-challenge.REDDIT_comments
Dataset Card for "REDDIT_comments"
Dataset Summary
Comments of 50 high-quality subreddits, extracted from the REDDIT PushShift data dumps (from 2006 to Jan 2023).
Supported Tasks
These comments can be used for text generation and language modeling, as well as dialogue modeling.
Dataset Structure
Data Splits
Each split corresponds to a specific subreddit in the following list: "tifu", "explainlikeimfive", "WritingPrompts", "changemyview"… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceGECLM/REDDIT_comments.pushshift-reddit-comments
Dataset Card for "pushshift-reddit"
More Information needed
news_commentary
Dataset Card for OPUS News-Commentary
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/news_commentary.pali-commentary-thai-script-siamrath-version
Multi-File CSV Dataset
คำอธิบาย
อรรถกถาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๘ เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
คำอธิบายของแต่ละเล่ม
เล่ม ๑: วินยฏฺกถา (สมนฺตปาสาทิกา ๑)
เล่ม ๒: วินยฏฺกถา (สมนฺตปาสาทิกา ๒)
เล่ม ๓: วินยฏฺกถา (สมนฺตปาสาทิกา ๓)
เล่ม ๔: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๑)
เล่ม ๕: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๒)
เล่ม ๖: ทีฆนิกายฏฺกถา… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-commentary-thai-script-siamrath-version.leading-comments
Leading Comments
We release Leading Comments, a collection of opening comment blocks from source code in The Stack v1, GitHub-Code, CodeParrot, The Pile, and RedPajama. The dataset brings together file-level license notices, copyright statements, redistribution conditions, and other introductory comments to support research on code licensing and dataset curation.
Paper · Preprint · Code and replication package
Getting started
Install the Hugging Face Datasets… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/leading-comments.mm2_level_comments
Mario Maker 2 level comments
Part of the Mario Maker 2 Dataset Collection
Dataset Description
The Mario Maker 2 level comment dataset consists of 31.9 million level comments from Nintendo's online service totaling around 20GB of data. The dataset was created using the self-hosted Mario Maker 2 api over the course of 1 month in February 2022.
How to use it
The Mario Maker 2 level comment dataset is a very large dataset so for most use cases it is recommended to… See the full description on the dataset page: https://huggingface.co/datasets/TheGreatRambler/mm2_level_comments.ids-paragraph-commentary-corpusThe data was collected from GitHub using https://github.com/cheop-byeon/RFCRationaleBuilder.
This is our first version of data for RFC Rationale.
Dataset Citation
If you find this dataset useful and include it in your studies, please cite our paper:
@inproceedings{bian2024tell,
title={Tell Me Why: Language Models Help Explain the Rationale Behind Internet Protocol Design},
author={Bian, Jie and Welzl, Michael and Kutuzov, Andrey and Arefyev, Nikolay},
booktitle={2024 IEEE… See the full description on the dataset page: https://huggingface.co/datasets/jiebi/ids-paragraph-commentary-corpus.code-comments-small
Comment Dataset
Opening comments extracted from code datasets with CommentMiner and ML4SE-toolkit.
Files are grouped as <dataset>/<language>/part-*.parquet.
The Hugging Face dataset card declares one config per source dataset and one split-safe language name per language.
Each row contains dataset, record_id, opening_comment, language, path, repo, extracted_at, and metadata.
For Parquet exports, metadata is stored as a JSON string so every source dataset shares one stable… See the full description on the dataset page: https://huggingface.co/datasets/Jkatzy/code-comments-small.danish_political_comments
Dataset Card for DanishPoliticalComments
Dataset Summary
The dataset consists of 9008 sentences that are labeled with fine-grained polarity in the range from -2 to 2 (negative to positive). The quality of the fine-grained is not cross-validated and is therefore subject to uncertainties; however, the simple polarity has been cross-validated and therefore is considered to be more correct.
Supported Tasks and Leaderboards
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/danish_political_comments.parallel-sentences-news-commentary
Dataset Card for Parallel Sentences - News Commentary
This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website.
In particular, this dataset contains the News-Commentary dataset.
Related Datasets
The following datasets are also a part of the Parallel Sentences collection:
parallel-sentences-europarl
parallel-sentences-global-voices… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-news-commentary.hackernews-comments
Hackernews Comments Dataset
A dataset of all HN API items from id=0 till id=41422887 (so from 2006 till 02 Sep 2024). The dataset is build by scraping the HN API according to its official schema and docs. Scraper code is also available on github: nixiesearch/hnscrape
Dataset contents
No cleaning, validation or filtering was performed. The resulting data files are raw JSON API response dumps in zstd-compressed JSONL files. An example payload:
{
"by": "goldfish"… See the full description on the dataset page: https://huggingface.co/datasets/nixiesearch/hackernews-comments.treaty-bodies-general-comments
Treaty Bodies General Comments
A paragraph-level dataset of General Comments and General Recommendations
adopted by the nine UN human-rights Treaty Bodies, with concerned-group
labels and document metadata. Companion to the
UNHRD search interface.
Licence
The curated dataset (paragraph segmentation, label annotation, document
metadata enrichment, footnote and section work) is released under
Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International… See the full description on the dataset page: https://huggingface.co/datasets/lszoszk/treaty-bodies-general-comments.OPUS_News-CommentaryComments_200K
Comments-200k 📚
1. Introduction
Raw reviews often intermix critical points with extraneous content, such as salutations and summaries. Directly feeding this unprocessed text into a model introduces significant noise and redundancy, which can compromise the precision of the generated rebuttal. Furthermore, due to diverse reviewer writing styles and varying conference formats, comments are typically presented in an unstructured manner. Therefore, to address these… See the full description on the dataset page: https://huggingface.co/datasets/RebuttalAgent/Comments_200K.Algerian-Youtube-Comments
Algerian YouTube Comments Dataset & Extraction Pipeline
A large-scale, curated, and high-entropy NLP corpus of 55,365 authentic Algerian dialect comments collected from popular Algerian YouTube channels, cleaned through an 18-step NLP preprocessing engine, and classified using Google Gemini AI models to isolate genuine Algerian Darija (الدارجة الجزائرية) and Algerian Arabizi (العرابيزي).
Hugging Face Hub: touati-kamel/Algerian-Youtube-Comments
Primary Extraction Script:… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/Algerian-Youtube-Comments.Youtube_news_comments_analysis_ver2
YouTube 뉴스 댓글 감정 분석 — 최종 통합본
기존 분석과 2026-10-07 추가 분석을 댓글 생성일 기준으로 통합했습니다.
총 74,136,841개 댓글, 2014-04-21~2026-09-29, 관측 날짜 3,467일.
구조와 날짜
final/raw/json/YYYY/MM/01-10/news_comments.parquet
final/raw/json/YYYY/MM/11-20/news_comments.parquet
final/raw/json/YYYY/MM/21-말일/news_comments.parquet
경로와 파일 내 날짜 정렬 기준은 **dataset_date(댓글 작성일)**입니다. 윤년과 실제 말일을 반영합니다.
news_date와 기존 호환용 date는 영상 게시일입니다. 댓글 시계열 집계에 사용하지 마세요.
신규 원본 comment_created_at, video_uploaded_at은 UTC 시각입니다.… See the full description on the dataset page: https://huggingface.co/datasets/MindCastSogang/Youtube_news_comments_analysis_ver2.civil-comments-wildsIn this dataset, given a textual dialogue i.e. an utterance along with two previous turns of context, the goal was to infer the underlying emotion of the utterance by choosing from four emotion classes - Happy, Sad, Angry and Others.reddit-comments-uwaterloo--- Generated Part of README Below ---
Dataset Overview
The goal is to have an open dataset of r/uwaterloo submissions, leveraging PRAW and the Reddit API to get downloads.
Posts are here
Comments are here
Creation Details
This dataset was created by alvanlii/dataset-creator-reddit-uwaterloo
Update Frequency
The dataset is updated custom with the most recent update being 2024-12-12 23:00:00 UTC+0000 where we added 72 new rows.
Licensing… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/reddit-comments-uwaterloo.Toxic_Russian_Commentshttps://www.kaggle.com/datasets/alexandersemiletov/toxic-russian-comments
0 - neutral user comments
1 - toxic user comments
Toxic Russian Comments Dataset
This dataset contains labelled comments from the popular Russian social network ok.ru.
The data was used in a competition where participants had to automatically label each comment with at least one of the four predefined classes. The classes represent different levels of toxicity. The competition was held on the All Cups platform.
Each… See the full description on the dataset page: https://huggingface.co/datasets/AlexSham/Toxic_Russian_Comments.fufufafa-comment-datasetsmart_contract_code_commentsAlgerian-Youtube-Comments
Algerian Youtube Comments
55,365 raw YouTube comments on Algeria-related videos for Darija social-text modeling, from the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/Algerian-Youtube-Comments: 55,365 train rows) and re-counted row-by-row with datasets streaming (load_dataset("algerian-nlp/Algerian-Youtube-Comments", split="train", streaming=True): 55,365 rows).
The default config answers: how do Algerians actually write in… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/Algerian-Youtube-Comments.RedditDataFor353_2026_COMMENTSToxic_Comment_Classification_Challengeregulatory_commentsUnited States governmental agencies often make proposed regulations open to the public for comment.
Proposed regulations are organized into "dockets". This project will use Regulation.gov public API
to aggregate and clean public comments for dockets that mention opioid use.
Each example will consist of one docket, and include metadata such as docket id, docket title, etc.
Each docket entry will also include information about the top 10 comments, including comment metadata
and comment text.youtube-comment-sentiment
YouTube Comments Sentiment Analysis Dataset (1M+ Labeled Comments)
Overview
This dataset comprises over one million YouTube comments, each annotated with sentiment labels—Positive, Neutral, or Negative. The comments span a diverse range of topics including programming, news, sports, politics and more, and are enriched with comprehensive metadata to facilitate various NLP and sentiment analysis tasks.
How to use:
import pandas as pd
df =… See the full description on the dataset page: https://huggingface.co/datasets/AmaanP314/youtube-comment-sentiment.multilingual-code-comments-fixed-8
Multilingual code comments
This dataset contains 500 source-code examples for each of Chinese, Dutch, English, Greek and Polish (2,500 examples total), with comments generated by five models and human correctness ratings and error annotations. Each language has a train split.
Human ratings use Correct, Partial and Incorrect. Error fields contain comma-separated taxonomy codes. A generated comment can carry multiple error codes.
Annotation interpretation
Stored… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/multilingual-code-comments-fixed-8.jigsaw-toxic-comment
Dataset Card for "jigsaw-toxic-comment"
More Information needed
