datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
civil_comments
Dataset Card for "civil_comments"
Dataset Summary
The comments in this dataset come from an archive of the Civil Comments
platform, a commenting plugin for independent news sites. These public comments
were created from 2015 - 2017 and appeared on approximately 50 English-language
news sites across the world. When Civil Comments shut down in 2017, they chose
to make the public comments available in a lasting open archive to enable future
research. The original data… See the full description on the dataset page: https://huggingface.co/datasets/google/civil_comments.REDDIT_comments
Dataset Card for "REDDIT_comments"
Dataset Summary
Comments of 50 high-quality subreddits, extracted from the REDDIT PushShift data dumps (from 2006 to Jan 2023).
Supported Tasks
These comments can be used for text generation and language modeling, as well as dialogue modeling.
Dataset Structure
Data Splits
Each split corresponds to a specific subreddit in the following list: "tifu", "explainlikeimfive", "WritingPrompts", "changemyview"… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceGECLM/REDDIT_comments.pushshift-reddit-comments
Dataset Card for "pushshift-reddit"
More Information needed
leading-comments
Leading Comments
We release Leading Comments, a collection of opening comment blocks from source code in The Stack v1, GitHub-Code, CodeParrot, The Pile, and RedPajama. The dataset brings together file-level license notices, copyright statements, redistribution conditions, and other introductory comments to support research on code licensing and dataset curation.
Paper · Preprint · Code and replication package
Getting started
Install the Hugging Face Datasets… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/leading-comments.mm2_level_comments
Mario Maker 2 level comments
Part of the Mario Maker 2 Dataset Collection
Dataset Description
The Mario Maker 2 level comment dataset consists of 31.9 million level comments from Nintendo's online service totaling around 20GB of data. The dataset was created using the self-hosted Mario Maker 2 api over the course of 1 month in February 2022.
How to use it
The Mario Maker 2 level comment dataset is a very large dataset so for most use cases it is recommended to… See the full description on the dataset page: https://huggingface.co/datasets/TheGreatRambler/mm2_level_comments.code-comments-small
Comment Dataset
Opening comments extracted from code datasets with CommentMiner and ML4SE-toolkit.
Files are grouped as <dataset>/<language>/part-*.parquet.
The Hugging Face dataset card declares one config per source dataset and one split-safe language name per language.
Each row contains dataset, record_id, opening_comment, language, path, repo, extracted_at, and metadata.
For Parquet exports, metadata is stored as a JSON string so every source dataset shares one stable… See the full description on the dataset page: https://huggingface.co/datasets/Jkatzy/code-comments-small.danish_political_comments
Dataset Card for DanishPoliticalComments
Dataset Summary
The dataset consists of 9008 sentences that are labeled with fine-grained polarity in the range from -2 to 2 (negative to positive). The quality of the fine-grained is not cross-validated and is therefore subject to uncertainties; however, the simple polarity has been cross-validated and therefore is considered to be more correct.
Supported Tasks and Leaderboards
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/danish_political_comments.hackernews-comments
Hackernews Comments Dataset
A dataset of all HN API items from id=0 till id=41422887 (so from 2006 till 02 Sep 2024). The dataset is build by scraping the HN API according to its official schema and docs. Scraper code is also available on github: nixiesearch/hnscrape
Dataset contents
No cleaning, validation or filtering was performed. The resulting data files are raw JSON API response dumps in zstd-compressed JSONL files. An example payload:
{
"by": "goldfish"… See the full description on the dataset page: https://huggingface.co/datasets/nixiesearch/hackernews-comments.treaty-bodies-general-comments
Treaty Bodies General Comments
A paragraph-level dataset of General Comments and General Recommendations
adopted by the nine UN human-rights Treaty Bodies, with concerned-group
labels and document metadata. Companion to the
UNHRD search interface.
Licence
The curated dataset (paragraph segmentation, label annotation, document
metadata enrichment, footnote and section work) is released under
Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International… See the full description on the dataset page: https://huggingface.co/datasets/lszoszk/treaty-bodies-general-comments.Comments_200K
Comments-200k 📚
1. Introduction
Raw reviews often intermix critical points with extraneous content, such as salutations and summaries. Directly feeding this unprocessed text into a model introduces significant noise and redundancy, which can compromise the precision of the generated rebuttal. Furthermore, due to diverse reviewer writing styles and varying conference formats, comments are typically presented in an unstructured manner. Therefore, to address these… See the full description on the dataset page: https://huggingface.co/datasets/RebuttalAgent/Comments_200K.Algerian-Youtube-Comments
Algerian YouTube Comments Dataset & Extraction Pipeline
A large-scale, curated, and high-entropy NLP corpus of 55,365 authentic Algerian dialect comments collected from popular Algerian YouTube channels, cleaned through an 18-step NLP preprocessing engine, and classified using Google Gemini AI models to isolate genuine Algerian Darija (الدارجة الجزائرية) and Algerian Arabizi (العرابيزي).
Hugging Face Hub: touati-kamel/Algerian-Youtube-Comments
Primary Extraction Script:… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/Algerian-Youtube-Comments.Youtube_news_comments_analysis_ver2
YouTube 뉴스 댓글 감정 분석 — 최종 통합본
기존 분석과 2026-10-07 추가 분석을 댓글 생성일 기준으로 통합했습니다.
총 74,136,841개 댓글, 2014-04-21~2026-09-29, 관측 날짜 3,467일.
구조와 날짜
final/raw/json/YYYY/MM/01-10/news_comments.parquet
final/raw/json/YYYY/MM/11-20/news_comments.parquet
final/raw/json/YYYY/MM/21-말일/news_comments.parquet
경로와 파일 내 날짜 정렬 기준은 **dataset_date(댓글 작성일)**입니다. 윤년과 실제 말일을 반영합니다.
news_date와 기존 호환용 date는 영상 게시일입니다. 댓글 시계열 집계에 사용하지 마세요.
신규 원본 comment_created_at, video_uploaded_at은 UTC 시각입니다.… See the full description on the dataset page: https://huggingface.co/datasets/MindCastSogang/Youtube_news_comments_analysis_ver2.civil-comments-wildsIn this dataset, given a textual dialogue i.e. an utterance along with two previous turns of context, the goal was to infer the underlying emotion of the utterance by choosing from four emotion classes - Happy, Sad, Angry and Others.reddit-comments-uwaterloo--- Generated Part of README Below ---
Dataset Overview
The goal is to have an open dataset of r/uwaterloo submissions, leveraging PRAW and the Reddit API to get downloads.
Posts are here
Comments are here
Creation Details
This dataset was created by alvanlii/dataset-creator-reddit-uwaterloo
Update Frequency
The dataset is updated custom with the most recent update being 2024-12-12 23:00:00 UTC+0000 where we added 72 new rows.
Licensing… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/reddit-comments-uwaterloo.Toxic_Russian_Commentshttps://www.kaggle.com/datasets/alexandersemiletov/toxic-russian-comments
0 - neutral user comments
1 - toxic user comments
Toxic Russian Comments Dataset
This dataset contains labelled comments from the popular Russian social network ok.ru.
The data was used in a competition where participants had to automatically label each comment with at least one of the four predefined classes. The classes represent different levels of toxicity. The competition was held on the All Cups platform.
Each… See the full description on the dataset page: https://huggingface.co/datasets/AlexSham/Toxic_Russian_Comments.smart_contract_code_commentsAlgerian-Youtube-Comments
Algerian Youtube Comments
55,365 raw YouTube comments on Algeria-related videos for Darija social-text modeling, from the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/Algerian-Youtube-Comments: 55,365 train rows) and re-counted row-by-row with datasets streaming (load_dataset("algerian-nlp/Algerian-Youtube-Comments", split="train", streaming=True): 55,365 rows).
The default config answers: how do Algerians actually write in… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/Algerian-Youtube-Comments.RedditDataFor353_2026_COMMENTSregulatory_commentsUnited States governmental agencies often make proposed regulations open to the public for comment.
Proposed regulations are organized into "dockets". This project will use Regulation.gov public API
to aggregate and clean public comments for dockets that mention opioid use.
Each example will consist of one docket, and include metadata such as docket id, docket title, etc.
Each docket entry will also include information about the top 10 comments, including comment metadata
and comment text.multilingual-code-comments-fixed-8
Multilingual code comments
This dataset contains 500 source-code examples for each of Chinese, Dutch, English, Greek and Polish (2,500 examples total), with comments generated by five models and human correctness ratings and error annotations. Each language has a train split.
Human ratings use Correct, Partial and Incorrect. Error fields contain comma-separated taxonomy codes. A generated comment can carry multiple error codes.
Annotation interpretation
Stored… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/multilingual-code-comments-fixed-8.tiktok-product-video-comments
TikTok Product Video Comments 2026
999,930 TikTok comments and inline replies from 17,730 product videos across 12 shopping categories, collected September 10-13, 2026. This is the corpus behind Crawlora's study 17,774 TikTok 'Where Did You Get That?' Comments: Who Gets an Answer.
Files
File
Rows
Bytes
SHA-256
data/videos.parquet
17,730
4,229,777
5aef621072ada8ff852007e810532c13069b3244ffae6c4fe47ff32549e83020
data/comments.parquet
999,930
81,410,979… See the full description on the dataset page: https://huggingface.co/datasets/crawlora-net/tiktok-product-video-comments.TrojanLoc-gpt-4.1-embeddings-wo-commentsJigsaw-Toxic-Comments
Jigsaw Toxic Comments — canonical 2017–2018 Kaggle release
A verbatim mirror of the Jigsaw Toxic Comment Classification Challenge training and test data, repackaged as three gzip-compressed CSVs. No rows added, removed, or reordered relative to the upstream Kaggle release — only the hosting moved and each file was gzipped.
Re-hosted under Heliosoph for ingestion-pipeline stability — the upstream files live behind Kaggle's competition-rules click-through and require an… See the full description on the dataset page: https://huggingface.co/datasets/Heliosoph/Jigsaw-Toxic-Comments.russian-toxic-comments-multilabel
Russian Toxic Comments Multi-label Dataset
Dataset Description
Этот датасет содержит размеченные комментарии на русском языке для задачи многозадачной (multi-task) и мультилейбл (multi-label) бинарной классификации токсичности.
Цель
Обучение модели для автоматического обнаружения трех типов токсичного контента:
Profanity (ненормативная лексика) — мат, оскорбления, нецензурная брань
Threat (угрозы) — явные или скрытые угрозы в адрес других людей… See the full description on the dataset page: https://huggingface.co/datasets/IvanFed/russian-toxic-comments-multilabel.task1720_civil_comments_toxicity_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1720_civil_comments_toxicity_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1720_civil_comments_toxicity_classification.850_commentscivil_comments
Dataset Card for "civil_comments"
Dataset Summary
The comments in this dataset come from an archive of the Civil Comments
platform, a commenting plugin for independent news sites. These public comments
were created from 2015 - 2017 and appeared on approximately 50 English-language
news sites across the world. When Civil Comments shut down in 2017, they chose
to make the public comments available in a lasting open archive to enable future
research. The original data… See the full description on the dataset page: https://huggingface.co/datasets/abdulahad-dev/civil_comments.regulatory_comments_apiUnited States governmental agencies often make proposed regulations open to the public for comment.
Proposed regulations are organized into "dockets". This dataset will use Regulation.gov public API
to aggregate and clean public comments for dockets that mention opioid use.
Each example will consist of one docket, and include metadata such as docket id, docket title, etc.
Each docket entry will also include information about the top 10 comments, including comment metadata
and comment text.civil_comments_helmswe-bench-commentsIncludes the SWE-bench dataset along with the corresponding comments for each instance_id
