Team Ai
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TheGreatRambler /mm2_level_comments Mario Maker 2 level comments Part of the Mario Maker 2 Dataset Collection Dataset Description The Mario Maker 2 level comment dataset consists of 31.9 million level comments from Nintendo's online service totaling around 20GB of data. The dataset was created using the self-hosted Mario Maker 2 api over the course of 1 month in February 2022. How to use it The Mario Maker 2 level comment dataset is a very large dataset so for most use cases it is recommended to… See the full description on the dataset page: https://huggingface.co/datasets/TheGreatRambler/mm2_level_comments.tabularother10M<n<100M3 likes1.1k downloads4y agoHugging Face02mashu-data /reddit-comments-sample Reddit Comment Trees Sample — Initial snapshot Initial sample: 43,913 posts and 59,874 comments across three communities. This is a selected, structurally checked snapshot, with incomplete subreddit coverage. It is not a complete three-month archive. Overview Posts and associated comments from r/LocalLLaMA, r/wallstreetbets, and r/SkincareAddiction. The requested post window is June 1–August 31, 2026 in Asia/Shanghai, with UTC bounds 2026-05-31 16:00:00… See the full description on the dataset page: https://huggingface.co/datasets/mashu-data/reddit-comments-sample.tabulartext-generation100K<n<1M0 likes101 downloads28d agoHugging Face03hausmer /truexa-comments Truha audience comments 45,918 real audience comments from the public satirical news channel «Труха⚡️Україна» (June 30 – July 22, 2026), collected from posts and cleaned (ads, links, mentions, ultra-short fragments, duplicates and phone numbers dropped — 50,000 raw → 45,918 clean). The audience writes in both Ukrainian and Russian (a mix, not a clean split — a share of comments code-switch between the two), so the corpus carries language: [ru, uk]. These are the "taste anchor"… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/truexa-comments.tabulartext-classification10K<n<100K0 likes65 downloads25d agoHugging Face04Linkseed /hacker_news_with_comments Dataset Card for [Dataset Name] Dataset Summary Hacker news until 2015 with comments. Collect from Google BigQuery open dataset. We didn't do any pre-processing except remove HTML tags. Supported Tasks and Leaderboards Comment Generation; News analysis with comments; Other comment-based NLP tasks. Languages English Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Linkseed/hacker_news_with_comments.tabulartext-generation1M<n<10M6 likes62 downloads4y agoHugging Face05pszemraj /LocalLLaMA-comments LocalLLaMA-comments A companion dataset to pszemraj/LocalLLaMA-posts. Time frame is in sync (up through Tue Mar 3 9PM EST 2026) tabulartext-generation1M<n<10M1 likes51 downloads7mo agoHugging Face06yallashoot /football-commentary-dataset ⚽ YallaShoot Football Commentary Dataset A curated dataset of football (soccer) match commentary, player mentions, and match events — built to power NLP models for the Arabic and global football community. 📌 Dataset Description This dataset contains structured football match commentary collected from live match feeds, covering top leagues including: 🏴󠁧󠁢󠁥󠁮󠁧󠁿 English Premier League (EPL) 🇪🇸 La Liga 🏆 UEFA Champions League 🌍 Arab World Leagues (Saudi Pro… See the full description on the dataset page: https://huggingface.co/datasets/yallashoot/football-commentary-dataset.tabulartext-classificationn<1K1 likes47 downloads6mo agoHugging Face07bigcode /stack-dedup-alt-commentsgated Dataset Description This is the Python, Java and JavaScript subsets of The Stack (v1.1) after cleaning* and agressive deduplication from stack-dedup-alt-decontaminate with filtering on comment to code ratio with minimum of 0.01 and maximum of 0.8. The additional comments filtering removes 26.5% of the dataset's volume which goes from 215GB of text to 170GB. (*) cleaning: near deduplication + PII redaction + line length & percentage of alphanumeric characters filtering + data… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/stack-dedup-alt-comments.tabulartext-generation10M<n<100M0 likes25 downloads3y agoHugging Face08EhsanShahbazi /digikala-commentsgatedtabulartext-classification10M<n<100M1 likes16 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.