datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
REDDIT_comments
Dataset Card for "REDDIT_comments"
Dataset Summary
Comments of 50 high-quality subreddits, extracted from the REDDIT PushShift data dumps (from 2006 to Jan 2023).
Supported Tasks
These comments can be used for text generation and language modeling, as well as dialogue modeling.
Dataset Structure
Data Splits
Each split corresponds to a specific subreddit in the following list: "tifu", "explainlikeimfive", "WritingPrompts", "changemyview"… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceGECLM/REDDIT_comments.mm2_level_comments
Mario Maker 2 level comments
Part of the Mario Maker 2 Dataset Collection
Dataset Description
The Mario Maker 2 level comment dataset consists of 31.9 million level comments from Nintendo's online service totaling around 20GB of data. The dataset was created using the self-hosted Mario Maker 2 api over the course of 1 month in February 2022.
How to use it
The Mario Maker 2 level comment dataset is a very large dataset so for most use cases it is recommended to… See the full description on the dataset page: https://huggingface.co/datasets/TheGreatRambler/mm2_level_comments.Algerian-Youtube-Comments
Algerian YouTube Comments Dataset & Extraction Pipeline
A large-scale, curated, and high-entropy NLP corpus of 55,365 authentic Algerian dialect comments collected from popular Algerian YouTube channels, cleaned through an 18-step NLP preprocessing engine, and classified using Google Gemini AI models to isolate genuine Algerian Darija (الدارجة الجزائرية) and Algerian Arabizi (العرابيزي).
Hugging Face Hub: touati-kamel/Algerian-Youtube-Comments
Primary Extraction Script:… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/Algerian-Youtube-Comments.Algerian-Youtube-Comments
Algerian Youtube Comments
55,365 raw YouTube comments on Algeria-related videos for Darija social-text modeling, from the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/Algerian-Youtube-Comments: 55,365 train rows) and re-counted row-by-row with datasets streaming (load_dataset("algerian-nlp/Algerian-Youtube-Comments", split="train", streaming=True): 55,365 rows).
The default config answers: how do Algerians actually write in… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/Algerian-Youtube-Comments.task1720_civil_comments_toxicity_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1720_civil_comments_toxicity_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1720_civil_comments_toxicity_classification.reddit-comments-sample
Reddit Comment Trees Sample — Initial snapshot
Initial sample: 43,913 posts and 59,874 comments across three communities.
This is a selected, structurally checked snapshot, with incomplete subreddit coverage. It is not a complete three-month archive.
Overview
Posts and associated comments from r/LocalLLaMA, r/wallstreetbets, and r/SkincareAddiction.
The requested post window is June 1–August 31, 2026 in Asia/Shanghai, with UTC bounds 2026-05-31 16:00:00… See the full description on the dataset page: https://huggingface.co/datasets/mashu-data/reddit-comments-sample.task1724_civil_comments_insult_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1724_civil_comments_insult_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1724_civil_comments_insult_classification.jira-comments-nsp
Dataset Card for Dataset Name
Dataset Summary
Dataset contains pairs of sentences with next_sentence_label for NSP. Sentences was given from public jira projects dataset. Next sentence is always next sentence in one comment or sentence from reply to the comment.
Supported Tasks and Leaderboards
NSP, MLM
Languages
English
Dataset Structure
sentence_a, sentence_b, next_sentence_label
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/pheepa/jira-comments-nsp.task1722_civil_comments_threat_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1722_civil_comments_threat_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1722_civil_comments_threat_classification.task1723_civil_comments_sexuallyexplicit_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1723_civil_comments_sexuallyexplicit_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1723_civil_comments_sexuallyexplicit_classification.truexa-comments
Truha audience comments
45,918 real audience comments from the public satirical news channel
«Труха⚡️Україна» (June 30 – July 22, 2026), collected from posts and
cleaned (ads, links, mentions, ultra-short fragments, duplicates and
phone numbers dropped — 50,000 raw → 45,918 clean).
The audience writes in both Ukrainian and Russian (a mix, not a
clean split — a share of comments code-switch between the two), so the
corpus carries language: [ru, uk].
These are the "taste anchor"… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/truexa-comments.task1721_civil_comments_obscenity_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1721_civil_comments_obscenity_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1721_civil_comments_obscenity_classification.hacker_news_with_comments
Dataset Card for [Dataset Name]
Dataset Summary
Hacker news until 2015 with comments. Collect from Google BigQuery open dataset. We didn't do any pre-processing except remove HTML tags.
Supported Tasks and Leaderboards
Comment Generation; News analysis with comments; Other comment-based NLP tasks.
Languages
English
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Linkseed/hacker_news_with_comments.DC_inside_comments
DC_inside_comments
This dataset contains 110,000 raw comments collected from DC Inside. It is intended for unsupervised learning or pretraining purposes.
Dataset Summary
Data Type: Unlabeled raw comments
Number of Examples: 110,000
Source: DC Inside
Related Dataset
For labeled data and multi-task annotated examples, please refer to the KoMultiText dataset.
How to Load the Dataset
from datasets import load_dataset
# Load the unlabeled dataset… See the full description on the dataset page: https://huggingface.co/datasets/Dasool/DC_inside_comments.LocalLLaMA-comments
LocalLLaMA-comments
A companion dataset to pszemraj/LocalLLaMA-posts. Time frame is in sync (up through Tue Mar 3 9PM EST 2026)
football-commentary-dataset
⚽ YallaShoot Football Commentary Dataset
A curated dataset of football (soccer) match commentary, player mentions, and match events — built to power NLP models for the Arabic and global football community.
📌 Dataset Description
This dataset contains structured football match commentary collected from live match feeds, covering top leagues including:
🏴 English Premier League (EPL)
🇪🇸 La Liga
🏆 UEFA Champions League
🌍 Arab World Leagues (Saudi Pro… See the full description on the dataset page: https://huggingface.co/datasets/yallashoot/football-commentary-dataset.news_commentary_tw本資料集是來自QingySi所搜集的中英對照新聞評論,一共有 252,776 對中英語翻譯的句子,是使用Alpaca的指令資料集格式製成。本資料集利用了OpenCC 進行簡轉繁。
youtube-comment-insights-chatml
YouTube Comment Insights - ChatML
Overview
This dataset contains instruction-tuning samples for structured YouTube comment analysis.
The dataset is formatted in ChatML conversational format and is intended for supervised fine-tuning (SFT), QLoRA, and instruction tuning of large language models.
Each sample contains:
sentiment
tone
pros
cons
Dataset Statistics
~20k training samples
~2k validation samples
Multilingual YouTube comments
Structured JSON… See the full description on the dataset page: https://huggingface.co/datasets/AnandforU/youtube-comment-insights-chatml.stack-dedup-alt-comments
Dataset Description
This is the Python, Java and JavaScript subsets of The Stack (v1.1) after cleaning* and agressive deduplication from stack-dedup-alt-decontaminate with
filtering on comment to code ratio with minimum of 0.01 and maximum of 0.8.
The additional comments filtering removes 26.5% of the dataset's volume which goes from 215GB of text to 170GB.
(*) cleaning: near deduplication + PII redaction + line length & percentage of alphanumeric characters filtering + data… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/stack-dedup-alt-comments.toxic_commentsafcon2025-commentary
AFCON 2025 Match Commentary Dataset
High-quality French match commentary training data for the Africa Cup of Nations 2025.
Dataset Details
Size: 2,000 training examples
Language: French
Format: JSONL (chat template)
Use Case: Fine-tuning LLMs for realistic African football match commentary
Event Distribution
82% General commentary
10% Goals
5% Substitutions
2% Penalties
1% Cards (yellow/red)
Teams Covered
Morocco, Senegal, Egypt, Nigeria, Côte… See the full description on the dataset page: https://huggingface.co/datasets/oxmo88/afcon2025-commentary.telegram-post-comments_data
Telegram Post-Comment Pairs (Russian)
Описание
Датасет пар «пост — комментарий» из публичных русскоязычных Telegram-каналов. Предназначен для задачи стилевой адаптации генеративных языковых моделей: модель получает текст поста и должна сгенерировать комментарий, стилистически согласованный с реальными пользовательскими откликами.
Структура
Split
Примеров
train
8 246
validation
1 034
test
1 020
Каждый пример содержит два поля:
post — текст… See the full description on the dataset page: https://huggingface.co/datasets/r9zhenka/telegram-post-comments_data.youtube-comment-insights-clean
YouTube Comment Insights - Clean Dataset
Overview
This dataset contains structured YouTube comment analytics data designed for visualization, analytics, and machine learning workflows.
Each sample contains:
comment
sentiment
tone
pros
cons
The dataset is intended for easy readability and downstream analytics tasks.
Dataset Statistics
~20k training samples
~2k validation samples
Multilingual YouTube comments
Structured JSON format
Files… See the full description on the dataset page: https://huggingface.co/datasets/AnandforU/youtube-comment-insights-clean.synthetic-football-commentary-qwen
Synthetic Passionate Football Commentary
Dataset Summary
This dataset contains synthetic conversational data designed to fine-tune large language models for creative writing and persona adoption. Specifically, it trains models to act as a passionate football commentator. The data pairs factual football match events with highly dramatic, emotional, and tactical commentary.
Data Generation
Base Data: The raw input features (Minute, Match, Team, Player, Action)… See the full description on the dataset page: https://huggingface.co/datasets/Alpaczyk/synthetic-football-commentary-qwen.digikala-commentsjira-commentaries-mlmDataset of jira comments from different projects of Apache and more.YouTube-Comment-Master-2024-v1
🎮 Roblox MM2 YouTube Comment Dataset (2024 Master)
A curated dataset of 27,089 clean, deduplicated, and length-filtered YouTube Short comments scraped from top Roblox Murder Mystery 2 (MM2) videos across 2024.
This dataset captures real-world internet gaming culture, short-form video engagement patterns, emoji distributions, trader slang, and brainrot banter—making it ideal for fine-tuning compact LLMs (such as Qwen2.5 or Llama 3) for casual gaming roleplay, comment generation… See the full description on the dataset page: https://huggingface.co/datasets/DinoResearch/YouTube-Comment-Master-2024-v1.task1725_civil_comments_severtoxicity_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1725_civil_comments_severtoxicity_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1725_civil_comments_severtoxicity_classification.cricket-commentary-dataset
Cricket Commentary Dataset
Description
A curated dataset of cricket commentary examples for fine-tuning language models to generate exciting sports commentary.
Dataset Structure
Each example contains:
instruction: Task description
input: Match situation (batsman, bowler, action, result)
output: Professional commentary text
Example
{
"instruction": "Generate exciting cricket commentary for this moment",
"input": "Batsman: Kohli, Bowler: Starc… See the full description on the dataset page: https://huggingface.co/datasets/siva-gunasehkaran/cricket-commentary-dataset.illegal-job-titles-comments
My Job Sounds Illegal Dataset
A dataset of ~42,000 humorous social media comments collected from a viral reel prompt:
"Comment your job but make it sound illegal."
The dataset contains creative, exaggerated, and comedic descriptions of professions and daily work activities written to sound suspicious, criminal, or absurd while remaining harmless and humorous.
Examples include:
"I manipulate vulnerable people into buying things they don't need."
"I convince tiny humans to obey me… See the full description on the dataset page: https://huggingface.co/datasets/BibbyResearch/illegal-job-titles-comments.
