Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceGECLM /REDDIT_comments Dataset Card for "REDDIT_comments" Dataset Summary Comments of 50 high-quality subreddits, extracted from the REDDIT PushShift data dumps (from 2006 to Jan 2023). Supported Tasks These comments can be used for text generation and language modeling, as well as dialogue modeling. Dataset Structure Data Splits Each split corresponds to a specific subreddit in the following list: "tifu", "explainlikeimfive", "WritingPrompts", "changemyview"… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceGECLM/REDDIT_comments.texttext-generation100M<n<1B27 likes8.1k downloads4y agoHugging Face02TheGreatRambler /mm2_level_comments Mario Maker 2 level comments Part of the Mario Maker 2 Dataset Collection Dataset Description The Mario Maker 2 level comment dataset consists of 31.9 million level comments from Nintendo's online service totaling around 20GB of data. The dataset was created using the self-hosted Mario Maker 2 api over the course of 1 month in February 2022. How to use it The Mario Maker 2 level comment dataset is a very large dataset so for most use cases it is recommended to… See the full description on the dataset page: https://huggingface.co/datasets/TheGreatRambler/mm2_level_comments.tabularother10M<n<100M3 likes1.1k downloads4y agoHugging Face03touati-kamel /Algerian-Youtube-Comments Algerian YouTube Comments Dataset & Extraction Pipeline A large-scale, curated, and high-entropy NLP corpus of 55,365 authentic Algerian dialect comments collected from popular Algerian YouTube channels, cleaned through an 18-step NLP preprocessing engine, and classified using Google Gemini AI models to isolate genuine Algerian Darija (الدارجة الجزائرية) and Algerian Arabizi (العرابيزي). Hugging Face Hub: touati-kamel/Algerian-Youtube-Comments Primary Extraction Script:… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/Algerian-Youtube-Comments.texttext-classification10K<n<100K0 likes280 downloads12d agoHugging Face04algerian-nlp /Algerian-Youtube-Comments Algerian Youtube Comments 55,365 raw YouTube comments on Algeria-related videos for Darija social-text modeling, from the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/Algerian-Youtube-Comments: 55,365 train rows) and re-counted row-by-row with datasets streaming (load_dataset("algerian-nlp/Algerian-Youtube-Comments", split="train", streaming=True): 55,365 rows). The default config answers: how do Algerians actually write in… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/Algerian-Youtube-Comments.texttext-generation10K<n<100K0 likes197 downloads23d agoHugging Face05Lots-of-LoRAs /task1720_civil_comments_toxicity_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1720_civil_comments_toxicity_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1720_civil_comments_toxicity_classification.texttext-generationn<1K0 likes131 downloads2y agoHugging Face06mashu-data /reddit-comments-sample Reddit Comment Trees Sample — Initial snapshot Initial sample: 43,913 posts and 59,874 comments across three communities. This is a selected, structurally checked snapshot, with incomplete subreddit coverage. It is not a complete three-month archive. Overview Posts and associated comments from r/LocalLLaMA, r/wallstreetbets, and r/SkincareAddiction. The requested post window is June 1–August 31, 2026 in Asia/Shanghai, with UTC bounds 2026-05-31 16:00:00… See the full description on the dataset page: https://huggingface.co/datasets/mashu-data/reddit-comments-sample.tabulartext-generation100K<n<1M0 likes101 downloads28d agoHugging Face07Lots-of-LoRAs /task1724_civil_comments_insult_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1724_civil_comments_insult_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1724_civil_comments_insult_classification.texttext-generation1K<n<10K0 likes92 downloads2y agoHugging Face08pheepa /jira-comments-nsp Dataset Card for Dataset Name Dataset Summary Dataset contains pairs of sentences with next_sentence_label for NSP. Sentences was given from public jira projects dataset. Next sentence is always next sentence in one comment or sentence from reply to the comment. Supported Tasks and Leaderboards NSP, MLM Languages English Dataset Structure sentence_a, sentence_b, next_sentence_label Source Data… See the full description on the dataset page: https://huggingface.co/datasets/pheepa/jira-comments-nsp.texttext-generationn<1K0 likes90 downloads4y agoHugging Face09Lots-of-LoRAs /task1722_civil_comments_threat_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1722_civil_comments_threat_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1722_civil_comments_threat_classification.texttext-generationn<1K0 likes77 downloads2y agoHugging Face10Lots-of-LoRAs /task1723_civil_comments_sexuallyexplicit_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1723_civil_comments_sexuallyexplicit_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1723_civil_comments_sexuallyexplicit_classification.texttext-generationn<1K0 likes67 downloads2y agoHugging Face11hausmer /truexa-comments Truha audience comments 45,918 real audience comments from the public satirical news channel «Труха⚡️Україна» (June 30 – July 22, 2026), collected from posts and cleaned (ads, links, mentions, ultra-short fragments, duplicates and phone numbers dropped — 50,000 raw → 45,918 clean). The audience writes in both Ukrainian and Russian (a mix, not a clean split — a share of comments code-switch between the two), so the corpus carries language: [ru, uk]. These are the "taste anchor"… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/truexa-comments.tabulartext-classification10K<n<100K0 likes65 downloads25d agoHugging Face12Lots-of-LoRAs /task1721_civil_comments_obscenity_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1721_civil_comments_obscenity_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1721_civil_comments_obscenity_classification.texttext-generationn<1K1 likes64 downloads2y agoHugging Face13Linkseed /hacker_news_with_comments Dataset Card for [Dataset Name] Dataset Summary Hacker news until 2015 with comments. Collect from Google BigQuery open dataset. We didn't do any pre-processing except remove HTML tags. Supported Tasks and Leaderboards Comment Generation; News analysis with comments; Other comment-based NLP tasks. Languages English Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Linkseed/hacker_news_with_comments.tabulartext-generation1M<n<10M6 likes62 downloads4y agoHugging Face14Dasool /DC_inside_comments DC_inside_comments This dataset contains 110,000 raw comments collected from DC Inside. It is intended for unsupervised learning or pretraining purposes. Dataset Summary Data Type: Unlabeled raw comments Number of Examples: 110,000 Source: DC Inside Related Dataset For labeled data and multi-task annotated examples, please refer to the KoMultiText dataset. How to Load the Dataset from datasets import load_dataset # Load the unlabeled dataset… See the full description on the dataset page: https://huggingface.co/datasets/Dasool/DC_inside_comments.texttext-generation100K<n<1M0 likes52 downloads2y agoHugging Face15pszemraj /LocalLLaMA-comments LocalLLaMA-comments A companion dataset to pszemraj/LocalLLaMA-posts. Time frame is in sync (up through Tue Mar 3 9PM EST 2026) tabulartext-generation1M<n<10M1 likes51 downloads7mo agoHugging Face16yallashoot /football-commentary-dataset ⚽ YallaShoot Football Commentary Dataset A curated dataset of football (soccer) match commentary, player mentions, and match events — built to power NLP models for the Arabic and global football community. 📌 Dataset Description This dataset contains structured football match commentary collected from live match feeds, covering top leagues including: 🏴󠁧󠁢󠁥󠁮󠁧󠁿 English Premier League (EPL) 🇪🇸 La Liga 🏆 UEFA Champions League 🌍 Arab World Leagues (Saudi Pro… See the full description on the dataset page: https://huggingface.co/datasets/yallashoot/football-commentary-dataset.tabulartext-classificationn<1K1 likes47 downloads6mo agoHugging Face17jslin09 /news_commentary_tw本資料集是來自QingySi所搜集的中英對照新聞評論,一共有 252,776 對中英語翻譯的句子,是使用Alpaca的指令資料集格式製成。本資料集利用了OpenCC 進行簡轉繁。 texttranslation100K<n<1M3 likes46 downloads3y agoHugging Face18AnandforU /youtube-comment-insights-chatml YouTube Comment Insights - ChatML Overview This dataset contains instruction-tuning samples for structured YouTube comment analysis. The dataset is formatted in ChatML conversational format and is intended for supervised fine-tuning (SFT), QLoRA, and instruction tuning of large language models. Each sample contains: sentiment tone pros cons Dataset Statistics ~20k training samples ~2k validation samples Multilingual YouTube comments Structured JSON… See the full description on the dataset page: https://huggingface.co/datasets/AnandforU/youtube-comment-insights-chatml.texttext-classification10K<n<100K0 likes34 downloads5mo agoHugging Face19bigcode /stack-dedup-alt-commentsgated Dataset Description This is the Python, Java and JavaScript subsets of The Stack (v1.1) after cleaning* and agressive deduplication from stack-dedup-alt-decontaminate with filtering on comment to code ratio with minimum of 0.01 and maximum of 0.8. The additional comments filtering removes 26.5% of the dataset's volume which goes from 215GB of text to 170GB. (*) cleaning: near deduplication + PII redaction + line length & percentage of alphanumeric characters filtering + data… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/stack-dedup-alt-comments.tabulartext-generation10M<n<100M0 likes25 downloads3y agoHugging Face20hramphul /toxic_commentstext-classification1M<n<10M0 likes20 downloads2y agoHugging Face21oxmo88 /afcon2025-commentary AFCON 2025 Match Commentary Dataset High-quality French match commentary training data for the Africa Cup of Nations 2025. Dataset Details Size: 2,000 training examples Language: French Format: JSONL (chat template) Use Case: Fine-tuning LLMs for realistic African football match commentary Event Distribution 82% General commentary 10% Goals 5% Substitutions 2% Penalties 1% Cards (yellow/red) Teams Covered Morocco, Senegal, Egypt, Nigeria, Côte… See the full description on the dataset page: https://huggingface.co/datasets/oxmo88/afcon2025-commentary.texttext-generation1K<n<10K0 likes20 downloads10mo agoHugging Face22r9zhenka /telegram-post-comments_data Telegram Post-Comment Pairs (Russian) Описание Датасет пар «пост — комментарий» из публичных русскоязычных Telegram-каналов. Предназначен для задачи стилевой адаптации генеративных языковых моделей: модель получает текст поста и должна сгенерировать комментарий, стилистически согласованный с реальными пользовательскими откликами. Структура Split Примеров train 8 246 validation 1 034 test 1 020 Каждый пример содержит два поля: post — текст… See the full description on the dataset page: https://huggingface.co/datasets/r9zhenka/telegram-post-comments_data.texttext-generation10K<n<100K0 likes18 downloads7mo agoHugging Face23AnandforU /youtube-comment-insights-clean YouTube Comment Insights - Clean Dataset Overview This dataset contains structured YouTube comment analytics data designed for visualization, analytics, and machine learning workflows. Each sample contains: comment sentiment tone pros cons The dataset is intended for easy readability and downstream analytics tasks. Dataset Statistics ~20k training samples ~2k validation samples Multilingual YouTube comments Structured JSON format Files… See the full description on the dataset page: https://huggingface.co/datasets/AnandforU/youtube-comment-insights-clean.texttext-classification10K<n<100K0 likes18 downloads5mo agoHugging Face24Alpaczyk /synthetic-football-commentary-qwen Synthetic Passionate Football Commentary Dataset Summary This dataset contains synthetic conversational data designed to fine-tune large language models for creative writing and persona adoption. Specifically, it trains models to act as a passionate football commentator. The data pairs factual football match events with highly dramatic, emotional, and tactical commentary. Data Generation Base Data: The raw input features (Minute, Match, Team, Player, Action)… See the full description on the dataset page: https://huggingface.co/datasets/Alpaczyk/synthetic-football-commentary-qwen.texttext-generation1K<n<10K0 likes17 downloads5mo agoHugging Face25EhsanShahbazi /digikala-commentsgatedtabulartext-classification10M<n<100M1 likes16 downloads10mo agoHugging Face26pheepa /jira-commentaries-mlmDataset of jira comments from different projects of Apache and more.texttext-generation100K<n<1M2 likes15 downloads3y agoHugging Face27DinoResearch /YouTube-Comment-Master-2024-v1 🎮 Roblox MM2 YouTube Comment Dataset (2024 Master) A curated dataset of 27,089 clean, deduplicated, and length-filtered YouTube Short comments scraped from top Roblox Murder Mystery 2 (MM2) videos across 2024. This dataset captures real-world internet gaming culture, short-form video engagement patterns, emoji distributions, trader slang, and brainrot banter—making it ideal for fine-tuning compact LLMs (such as Qwen2.5 or Llama 3) for casual gaming roleplay, comment generation… See the full description on the dataset page: https://huggingface.co/datasets/DinoResearch/YouTube-Comment-Master-2024-v1.texttext-generation10K<n<100K0 likes13 downloads2mo agoHugging Face28Lots-of-LoRAs /task1725_civil_comments_severtoxicity_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1725_civil_comments_severtoxicity_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1725_civil_comments_severtoxicity_classification.texttext-generationn<1K0 likes12 downloads2y agoHugging Face29siva-gunasehkaran /cricket-commentary-dataset Cricket Commentary Dataset Description A curated dataset of cricket commentary examples for fine-tuning language models to generate exciting sports commentary. Dataset Structure Each example contains: instruction: Task description input: Match situation (batsman, bowler, action, result) output: Professional commentary text Example { "instruction": "Generate exciting cricket commentary for this moment", "input": "Batsman: Kohli, Bowler: Starc… See the full description on the dataset page: https://huggingface.co/datasets/siva-gunasehkaran/cricket-commentary-dataset.texttext-generationn<1K0 likes11 downloads9mo agoHugging Face30BibbyResearch /illegal-job-titles-commentsgated My Job Sounds Illegal Dataset A dataset of ~42,000 humorous social media comments collected from a viral reel prompt: "Comment your job but make it sound illegal." The dataset contains creative, exaggerated, and comedic descriptions of professions and daily work activities written to sound suspicious, criminal, or absurd while remaining harmless and humorous. Examples include: "I manipulate vulnerable people into buying things they don't need." "I convince tiny humans to obey me… See the full description on the dataset page: https://huggingface.co/datasets/BibbyResearch/illegal-job-titles-comments.text-generation10K<n<100K3 likes9 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.