datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hackernews-comments
Hackernews Comments Dataset
A dataset of all HN API items from id=0 till id=41422887 (so from 2006 till 02 Sep 2024). The dataset is build by scraping the HN API according to its official schema and docs. Scraper code is also available on github: nixiesearch/hnscrape
Dataset contents
No cleaning, validation or filtering was performed. The resulting data files are raw JSON API response dumps in zstd-compressed JSONL files. An example payload:
{
"by": "goldfish"… See the full description on the dataset page: https://huggingface.co/datasets/nixiesearch/hackernews-comments.Toxic_Russian_Commentshttps://www.kaggle.com/datasets/alexandersemiletov/toxic-russian-comments
0 - neutral user comments
1 - toxic user comments
Toxic Russian Comments Dataset
This dataset contains labelled comments from the popular Russian social network ok.ru.
The data was used in a competition where participants had to automatically label each comment with at least one of the four predefined classes. The classes represent different levels of toxicity. The competition was held on the All Cups platform.
Each… See the full description on the dataset page: https://huggingface.co/datasets/AlexSham/Toxic_Russian_Comments.game_commentary_sftswiss-law-commentary
Swiss Law Commentary Dataset
Open-access, AI-generated legal commentary on 8 Swiss federal laws.
Source: openlegalcommentary.ch
Laws: BV, ZGB, OR, ZPO, StGB, StPO, SchKG, VwVG
Layers: Summary (B1 level), Doctrine (academic), Case Law (BGE digest)
Languages: DE, FR, IT, EN
Articles exported: 215
License: CC BY-SA 4.0
Built with ♥ by Jonas Hertner
Files
One JSONL file per law (e.g., or.jsonl, zgb.jsonl).
Schema
Field
Type
Description
law… See the full description on the dataset page: https://huggingface.co/datasets/voilaj/swiss-law-commentary.Korean-YouTube-Comment-Sentiment-Dataset
Korean YouTube Comment Sentiment Dataset
Data Overview
Summary
본 데이터셋은 유튜브에서 수집된 한국어 댓글 5,482개와 이에 대응하는 감정 레이블(긍정, 부정, 중립, 불명확)로 구성된 감정 분류용 데이터셋입니다.
주요 레이블: 긍정, 부정, 중립, 불명확
Features
수집 대상: 요리, 뷰티, 게임, 여행, 쇼핑 등 분야의 10만 명 이상 구독자를 보유한 유튜브 채널
형식: JSON (id, text, label)
검수: 한국인 검수자에 의한 수작업 라벨링 및 교차 검토
본 데이터셋은 구어체, 이모지, 줄임말 등 실제 사용자 표현이 반영되어 있습니다.
Dataset Structure
Dataset Fields
Field
Type
Description
id
string
각 댓글의… See the full description on the dataset page: https://huggingface.co/datasets/LLM-SocialMedia/Korean-YouTube-Comment-Sentiment-Dataset.douban_movie_commentscricket-commentary-ball-datacommentinject
CommentInject: Testing Whether Code Comments Mislead AI Security Reviewers
A benchmark of adversarial code-comment injection against LLM-based code vulnerability detectors. The vulnerable code stays exactly the same; only a comment claiming the code is safe is added.
30 vulnerable Python functions across 14 CWE categories (samples).
12 adversarial comment strategies in four families (strategies).
3,360 recorded trials from four studies, run locally through Ollama at temperature… See the full description on the dataset page: https://huggingface.co/datasets/sunny-chokshi/commentinject.swiss-law-commentary
Swiss Law Commentary Dataset
Open-access, AI-generated legal commentary on 8 Swiss federal laws.
Source: openlegalcommentary.ch
Laws: BV, ZGB, OR, ZPO, StGB, StPO, SchKG, VwVG
Layers: Summary (B1 level), Doctrine (academic), Case Law (BGE digest)
Languages: DE, FR, IT, EN
Articles exported: 215
License: CC BY-SA 4.0
Built with ♥ by Jonas Hertner
Files
One JSONL file per law (e.g., or.jsonl, zgb.jsonl).
Schema
Field
Type
Description
law… See the full description on the dataset page: https://huggingface.co/datasets/AccountVerify/swiss-law-commentary.comment-translation-01This dataset is based on Kaggle.
This dataset includes translations of 69,000 Reddit comments into 17 languages (English to 16 languages):
Belarusian, Czech, German,
English, Spanish, Finnish,
French, Italian, Japanese,
Kazakh, Korean, Latvian,
Polish, Russian, Swedish,
Ukrainian, and Chinese.
It contains 50% regular comments and 50% highly negative ones.
Enjoy using it!
ru-comments-classificationtruexa-comments
Truha audience comments
45,918 real audience comments from the public satirical news channel
«Труха⚡️Україна» (June 30 – July 22, 2026), collected from posts and
cleaned (ads, links, mentions, ultra-short fragments, duplicates and
phone numbers dropped — 50,000 raw → 45,918 clean).
The audience writes in both Ukrainian and Russian (a mix, not a
clean split — a share of comments code-switch between the two), so the
corpus carries language: [ru, uk].
These are the "taste anchor"… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/truexa-comments.youtube-comment-exposure
YouTube Comment Exposure Corpus v1
1,345,353 top-level YouTube comments from 633 fully-enumerated videos across 34
channels, with the field the existing public corpora leave out: when each comment was
posted.
Why this exists
The two large public YouTube comment datasets, YT-30M and YTCommentVerse (~32M comments
between them), carry like counts with no timestamps and no reply counts. A like count
without a post time is close to unusable for anything causal, because… See the full description on the dataset page: https://huggingface.co/datasets/brianhliou/youtube-comment-exposure.xhs-comment-responsenews_commentary_tw本資料集是來自QingySi所搜集的中英對照新聞評論,一共有 252,776 對中英語翻譯的句子,是使用Alpaca的指令資料集格式製成。本資料集利用了OpenCC 進行簡轉繁。
bilibili_comments_sharegpt
林亦LYi B站留言 sharegpt 格式
把 train-test-validation 全合并了,因为使用上是混合其他对话资料训练,没有 overfitting 问题。如果你只是单训练这一份资料,小心overfitting
资料清理也把 B站表情符号去掉了,本来想保留但是无法都放到 system prompt 里,所以还是下次吧
youtube-comment-insights-chatml
YouTube Comment Insights - ChatML
Overview
This dataset contains instruction-tuning samples for structured YouTube comment analysis.
The dataset is formatted in ChatML conversational format and is intended for supervised fine-tuning (SFT), QLoRA, and instruction tuning of large language models.
Each sample contains:
sentiment
tone
pros
cons
Dataset Statistics
~20k training samples
~2k validation samples
Multilingual YouTube comments
Structured JSON… See the full description on the dataset page: https://huggingface.co/datasets/AnandforU/youtube-comment-insights-chatml.cs_facebook-comments
Dataset Card for Czech Facebook comments
Dataset Description
The dataset contains user comments from Facebook. Each comment contains text, sentiment (positive/negative/neutral).
The dataset has in total (train+validation+test) 6,600 reviews. The data is balanced.
Dataset Features
Each sample contains:
comment_id: unique string identifier of the comment.
sentiment_str: string representation of the rating - "pozitivní" / "neutrální" / "negativní"
sentiment_int:… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_facebook-comments.test-sync-commentsreddit_news_articles_commentsreddit-comment-generation-v1
🤖 Official MentionBroker Research: Legacy V1 Comment Generation Dataset
📌 Dataset Summary
This is an official Legacy V1 Release from the MentionBroker Research Lab. This corpus represents our foundational work in mapping Community Linguistic Variance and Conversational Response Efficacy.
This dataset is specifically curated for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs) to generate high-authority, authentic community responses. By releasing this V1… See the full description on the dataset page: https://huggingface.co/datasets/MentionBroker/reddit-comment-generation-v1.Youtube_shorts_comments
Fine-tuned distilgpt2 on this dataset
The average amount of emojis in a YouTube short comment is 4.59 (Based on this dataset)
12 millions view 😂😂😂 good god 🤦♂️
lemmy-world-commentsThis is a data set of lemmy.world comments.
commenttrendyol-electronics-products-features-and-commentsgenerated-news-commentsThis dataset has been generated by an LLM.
Convert to a DuckDB
# 1) Download the JSONL file:
wget "https://huggingface.co/datasets/speech-uk/generated-news-comments/resolve/main/generated_comments.jsonl"
# 2) Open a DuckDB session in the terminal, then import the JSONL file into a table:
duckdb
$ CREATE TABLE generated_comments AS SELECT * FROM read_json_auto('generated_comments.jsonl');
# 3) Export the data from memory to a file:
$ ATTACH 'my_database.db';
$ COPY FROM… See the full description on the dataset page: https://huggingface.co/datasets/speech-uk/generated-news-comments.synthetic-football-commentary-qwen
Synthetic Passionate Football Commentary
Dataset Summary
This dataset contains synthetic conversational data designed to fine-tune large language models for creative writing and persona adoption. Specifically, it trains models to act as a passionate football commentator. The data pairs factual football match events with highly dramatic, emotional, and tactical commentary.
Data Generation
Base Data: The raw input features (Minute, Match, Team, Player, Action)… See the full description on the dataset page: https://huggingface.co/datasets/Alpaczyk/synthetic-football-commentary-qwen.rstp_comments_2UX-comments
Cross-Domain Polarity Models to Evaluate User eXperience in E-learning
Abstract
Virtual learning environments are growing in importance as fast as e-learning, which is becoming highly demanded by universities and students worldwide.
This paper investigates how to automatically evaluate User eXperience in this domain using sentiment analysis techniques.
For this purpose, a corpus has been built with the opinions of 583 users (107 English speakers and 476 Spanish… See the full description on the dataset page: https://huggingface.co/datasets/ELiRF/UX-comments.eval-corm_black_comments
Ranking Evaluation Dataset
