datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mm2_level_comments
Mario Maker 2 level comments
Part of the Mario Maker 2 Dataset Collection
Dataset Description
The Mario Maker 2 level comment dataset consists of 31.9 million level comments from Nintendo's online service totaling around 20GB of data. The dataset was created using the self-hosted Mario Maker 2 api over the course of 1 month in February 2022.
How to use it
The Mario Maker 2 level comment dataset is a very large dataset so for most use cases it is recommended to… See the full description on the dataset page: https://huggingface.co/datasets/TheGreatRambler/mm2_level_comments.reddit-comments-sample
Reddit Comment Trees Sample — Initial snapshot
Initial sample: 43,913 posts and 59,874 comments across three communities.
This is a selected, structurally checked snapshot, with incomplete subreddit coverage. It is not a complete three-month archive.
Overview
Posts and associated comments from r/LocalLLaMA, r/wallstreetbets, and r/SkincareAddiction.
The requested post window is June 1–August 31, 2026 in Asia/Shanghai, with UTC bounds 2026-05-31 16:00:00… See the full description on the dataset page: https://huggingface.co/datasets/mashu-data/reddit-comments-sample.truexa-comments
Truha audience comments
45,918 real audience comments from the public satirical news channel
«Труха⚡️Україна» (June 30 – July 22, 2026), collected from posts and
cleaned (ads, links, mentions, ultra-short fragments, duplicates and
phone numbers dropped — 50,000 raw → 45,918 clean).
The audience writes in both Ukrainian and Russian (a mix, not a
clean split — a share of comments code-switch between the two), so the
corpus carries language: [ru, uk].
These are the "taste anchor"… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/truexa-comments.hacker_news_with_comments
Dataset Card for [Dataset Name]
Dataset Summary
Hacker news until 2015 with comments. Collect from Google BigQuery open dataset. We didn't do any pre-processing except remove HTML tags.
Supported Tasks and Leaderboards
Comment Generation; News analysis with comments; Other comment-based NLP tasks.
Languages
English
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Linkseed/hacker_news_with_comments.LocalLLaMA-comments
LocalLLaMA-comments
A companion dataset to pszemraj/LocalLLaMA-posts. Time frame is in sync (up through Tue Mar 3 9PM EST 2026)
football-commentary-dataset
⚽ YallaShoot Football Commentary Dataset
A curated dataset of football (soccer) match commentary, player mentions, and match events — built to power NLP models for the Arabic and global football community.
📌 Dataset Description
This dataset contains structured football match commentary collected from live match feeds, covering top leagues including:
🏴 English Premier League (EPL)
🇪🇸 La Liga
🏆 UEFA Champions League
🌍 Arab World Leagues (Saudi Pro… See the full description on the dataset page: https://huggingface.co/datasets/yallashoot/football-commentary-dataset.stack-dedup-alt-comments
Dataset Description
This is the Python, Java and JavaScript subsets of The Stack (v1.1) after cleaning* and agressive deduplication from stack-dedup-alt-decontaminate with
filtering on comment to code ratio with minimum of 0.01 and maximum of 0.8.
The additional comments filtering removes 26.5% of the dataset's volume which goes from 215GB of text to 170GB.
(*) cleaning: near deduplication + PII redaction + line length & percentage of alphanumeric characters filtering + data… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/stack-dedup-alt-comments.digikala-comments
