Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google /civil_comments Dataset Card for "civil_comments" Dataset Summary The comments in this dataset come from an archive of the Civil Comments platform, a commenting plugin for independent news sites. These public comments were created from 2015 - 2017 and appeared on approximately 50 English-language news sites across the world. When Civil Comments shut down in 2017, they chose to make the public comments available in a lasting open archive to enable future research. The original data… See the full description on the dataset page: https://huggingface.co/datasets/google/civil_comments.tabulartext-classification1M<n<10M41 likes10k downloads3y agoHugging Face02thesofakillers /jigsaw-toxic-comment-classification-challenge Dataset Description You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior. The types of toxicity are: toxic severe_toxic obscene threat insult identity_hate You must create a model which predicts a probability of each type of toxicity for each comment. File descriptions train.csv - the training set, contains comments with their binary labels test.csv - the test set, you must predict the toxicity… See the full description on the dataset page: https://huggingface.co/datasets/thesofakillers/jigsaw-toxic-comment-classification-challenge.tabular100K<n<1M13 likes9.7k downloads2y agoHugging Face03fddemarco /pushshift-reddit-comments Dataset Card for "pushshift-reddit" More Information needed tabular1B<n<10B27 likes4.6k downloads3y agoHugging Face04uisp /pali-commentary-thai-script-siamrath-version Multi-File CSV Dataset คำอธิบาย อรรถกถาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๘ เล่ม ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์ 01/010001.csv: เล่ม 1 หน้า 1 01/010002.csv: เล่ม 1 หน้า 2 ... 02/020001.csv: เล่ม 2 หน้า 1 คำอธิบายของแต่ละเล่ม เล่ม ๑: วินยฏฺกถา (สมนฺตปาสาทิกา ๑) เล่ม ๒: วินยฏฺกถา (สมนฺตปาสาทิกา ๒) เล่ม ๓: วินยฏฺกถา (สมนฺตปาสาทิกา ๓) เล่ม ๔: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๑) เล่ม ๕: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๒) เล่ม ๖: ทีฆนิกายฏฺกถา… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-commentary-thai-script-siamrath-version.tabular100K<n<1M3 likes1.2k downloads2y agoHugging Face05TheGreatRambler /mm2_level_comments Mario Maker 2 level comments Part of the Mario Maker 2 Dataset Collection Dataset Description The Mario Maker 2 level comment dataset consists of 31.9 million level comments from Nintendo's online service totaling around 20GB of data. The dataset was created using the self-hosted Mario Maker 2 api over the course of 1 month in February 2022. How to use it The Mario Maker 2 level comment dataset is a very large dataset so for most use cases it is recommended to… See the full description on the dataset page: https://huggingface.co/datasets/TheGreatRambler/mm2_level_comments.tabularother10M<n<100M3 likes1.1k downloads4y agoHugging Face06nixiesearch /hackernews-comments Hackernews Comments Dataset A dataset of all HN API items from id=0 till id=41422887 (so from 2006 till 02 Sep 2024). The dataset is build by scraping the HN API according to its official schema and docs. Scraper code is also available on github: nixiesearch/hnscrape Dataset contents No cleaning, validation or filtering was performed. The resulting data files are raw JSON API response dumps in zstd-compressed JSONL files. An example payload: { "by": "goldfish"… See the full description on the dataset page: https://huggingface.co/datasets/nixiesearch/hackernews-comments.tabular10M<n<100M3 likes403 downloads2y agoHugging Face07lszoszk /treaty-bodies-general-comments Treaty Bodies General Comments A paragraph-level dataset of General Comments and General Recommendations adopted by the nine UN human-rights Treaty Bodies, with concerned-group labels and document metadata. Companion to the UNHRD search interface. Licence The curated dataset (paragraph segmentation, label annotation, document metadata enrichment, footnote and section work) is released under Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International… See the full description on the dataset page: https://huggingface.co/datasets/lszoszk/treaty-bodies-general-comments.tabulartext-classification1K<n<10K0 likes392 downloads1mo agoHugging Face08MindCastSogang /Youtube_news_comments_analysis_ver2 YouTube 뉴스 댓글 감정 분석 — 최종 통합본 기존 분석과 2026-10-07 추가 분석을 댓글 생성일 기준으로 통합했습니다. 총 74,136,841개 댓글, 2014-04-21~2026-09-29, 관측 날짜 3,467일. 구조와 날짜 final/raw/json/YYYY/MM/01-10/news_comments.parquet final/raw/json/YYYY/MM/11-20/news_comments.parquet final/raw/json/YYYY/MM/21-말일/news_comments.parquet 경로와 파일 내 날짜 정렬 기준은 **dataset_date(댓글 작성일)**입니다. 윤년과 실제 말일을 반영합니다. news_date와 기존 호환용 date는 영상 게시일입니다. 댓글 시계열 집계에 사용하지 마세요. 신규 원본 comment_created_at, video_uploaded_at은 UTC 시각입니다.… See the full description on the dataset page: https://huggingface.co/datasets/MindCastSogang/Youtube_news_comments_analysis_ver2.tabulartext-classification10M<n<100M0 likes274 downloads16h agoHugging Face09alvanlii /reddit-comments-uwaterloo--- Generated Part of README Below --- Dataset Overview The goal is to have an open dataset of r/uwaterloo submissions, leveraging PRAW and the Reddit API to get downloads. Posts are here Comments are here Creation Details This dataset was created by alvanlii/dataset-creator-reddit-uwaterloo Update Frequency The dataset is updated custom with the most recent update being 2024-12-12 23:00:00 UTC+0000 where we added 72 new rows. Licensing… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/reddit-comments-uwaterloo.tabular1M<n<10M2 likes236 downloads2y agoHugging Face10Jennyyahoo /RedditDataFor353_2026_COMMENTStabular1M<n<10M0 likes191 downloads2mo agoHugging Face11AmaanP314 /youtube-comment-sentiment YouTube Comments Sentiment Analysis Dataset (1M+ Labeled Comments) Overview This dataset comprises over one million YouTube comments, each annotated with sentiment labels—Positive, Neutral, or Negative. The comments span a diverse range of topics including programming, news, sports, politics and more, and are enriched with comprehensive metadata to facilitate various NLP and sentiment analysis tasks. How to use: import pandas as pd df =… See the full description on the dataset page: https://huggingface.co/datasets/AmaanP314/youtube-comment-sentiment.tabulartext-classification1M<n<10M5 likes175 downloads7mo agoHugging Face12crawlora-net /tiktok-product-video-comments TikTok Product Video Comments 2026 999,930 TikTok comments and inline replies from 17,730 product videos across 12 shopping categories, collected September 10-13, 2026. This is the corpus behind Crawlora's study 17,774 TikTok 'Where Did You Get That?' Comments: Who Gets an Answer. Files File Rows Bytes SHA-256 data/videos.parquet 17,730 4,229,777 5aef621072ada8ff852007e810532c13069b3244ffae6c4fe47ff32549e83020 data/comments.parquet 999,930 81,410,979… See the full description on the dataset page: https://huggingface.co/datasets/crawlora-net/tiktok-product-video-comments.tabular1M<n<10M1 likes147 downloads27d agoHugging Face13Heliosoph /Jigsaw-Toxic-Comments Jigsaw Toxic Comments — canonical 2017–2018 Kaggle release A verbatim mirror of the Jigsaw Toxic Comment Classification Challenge training and test data, repackaged as three gzip-compressed CSVs. No rows added, removed, or reordered relative to the upstream Kaggle release — only the hosting moved and each file was gzipped. Re-hosted under Heliosoph for ingestion-pipeline stability — the upstream files live behind Kaggle's competition-rules click-through and require an… See the full description on the dataset page: https://huggingface.co/datasets/Heliosoph/Jigsaw-Toxic-Comments.tabulartext-classification100K<n<1M0 likes133 downloads4mo agoHugging Face14IvanFed /russian-toxic-comments-multilabel Russian Toxic Comments Multi-label Dataset Dataset Description Этот датасет содержит размеченные комментарии на русском языке для задачи многозадачной (multi-task) и мультилейбл (multi-label) бинарной классификации токсичности. Цель Обучение модели для автоматического обнаружения трех типов токсичного контента: Profanity (ненормативная лексика) — мат, оскорбления, нецензурная брань Threat (угрозы) — явные или скрытые угрозы в адрес других людей… See the full description on the dataset page: https://huggingface.co/datasets/IvanFed/russian-toxic-comments-multilabel.tabulartext-classification100K<n<1M1 likes133 downloads3mo agoHugging Face15voilaj /swiss-law-commentary Swiss Law Commentary Dataset Open-access, AI-generated legal commentary on 8 Swiss federal laws. Source: openlegalcommentary.ch Laws: BV, ZGB, OR, ZPO, StGB, StPO, SchKG, VwVG Layers: Summary (B1 level), Doctrine (academic), Case Law (BGE digest) Languages: DE, FR, IT, EN Articles exported: 215 License: CC BY-SA 4.0 Built with ♥ by Jonas Hertner Files One JSONL file per law (e.g., or.jsonl, zgb.jsonl). Schema Field Type Description law… See the full description on the dataset page: https://huggingface.co/datasets/voilaj/swiss-law-commentary.tabular1K<n<10K1 likes126 downloads7mo agoHugging Face16abdulahad-dev /civil_comments Dataset Card for "civil_comments" Dataset Summary The comments in this dataset come from an archive of the Civil Comments platform, a commenting plugin for independent news sites. These public comments were created from 2015 - 2017 and appeared on approximately 50 English-language news sites across the world. When Civil Comments shut down in 2017, they chose to make the public comments available in a lasting open archive to enable future research. The original data… See the full description on the dataset page: https://huggingface.co/datasets/abdulahad-dev/civil_comments.tabulartext-classification1M<n<10M0 likes125 downloads5mo agoHugging Face17LLM-SocialMedia /Korean-YouTube-Comment-Sentiment-Dataset Korean YouTube Comment Sentiment Dataset Data Overview Summary 본 데이터셋은 유튜브에서 수집된 한국어 댓글 5,482개와 이에 대응하는 감정 레이블(긍정, 부정, 중립, 불명확)로 구성된 감정 분류용 데이터셋입니다. 주요 레이블: 긍정, 부정, 중립, 불명확 Features 수집 대상: 요리, 뷰티, 게임, 여행, 쇼핑 등 분야의 10만 명 이상 구독자를 보유한 유튜브 채널 형식: JSON (id, text, label) 검수: 한국인 검수자에 의한 수작업 라벨링 및 교차 검토 본 데이터셋은 구어체, 이모지, 줄임말 등 실제 사용자 표현이 반영되어 있습니다. Dataset Structure Dataset Fields Field Type Description id string 각 댓글의… See the full description on the dataset page: https://huggingface.co/datasets/LLM-SocialMedia/Korean-YouTube-Comment-Sentiment-Dataset.tabulartext-classification10K<n<100K3 likes118 downloads1y agoHugging Face18SteveTran /naruto-reddit-commentsThis dataset is extracted from Reddit datasets for research purpose image100K<n<1M0 likes106 downloads2y agoHugging Face19kattymandy /cricket-commentary-ball-datatabular100K<n<1M0 likes104 downloads4mo agoHugging Face20mashu-data /reddit-comments-sample Reddit Comment Trees Sample — Initial snapshot Initial sample: 43,913 posts and 59,874 comments across three communities. This is a selected, structurally checked snapshot, with incomplete subreddit coverage. It is not a complete three-month archive. Overview Posts and associated comments from r/LocalLLaMA, r/wallstreetbets, and r/SkincareAddiction. The requested post window is June 1–August 31, 2026 in Asia/Shanghai, with UTC bounds 2026-05-31 16:00:00… See the full description on the dataset page: https://huggingface.co/datasets/mashu-data/reddit-comments-sample.tabulartext-generation100K<n<1M0 likes101 downloads28d agoHugging Face21Proyecto-charlie-kirk-reddit /Charlie-Kirk-comments Dataset Summary This dataset is used in the following Github repository: "Charlie-Kirk-comments-sentiment-analysis". This dataset contains Reddit comments related to Charlie Kirk, founder of Turning Point USA, from 11-2024 to 10-2025, which includes the event of his death in 10th September 2025. Charlie Kirk was a highly polarizing political figure in the United States, often drawing both strong support and harsh criticism. Following his death, many media platforms, included… See the full description on the dataset page: https://huggingface.co/datasets/Proyecto-charlie-kirk-reddit/Charlie-Kirk-comments.tabular1K<n<10K3 likes91 downloads1y agoHugging Face22asdfceeegdag /civil_comments Dataset Card for "civil_comments" Dataset Summary The comments in this dataset come from an archive of the Civil Comments platform, a commenting plugin for independent news sites. These public comments were created from 2015 - 2017 and appeared on approximately 50 English-language news sites across the world. When Civil Comments shut down in 2017, they chose to make the public comments available in a lasting open archive to enable future research. The original data… See the full description on the dataset page: https://huggingface.co/datasets/asdfceeegdag/civil_comments.tabulartext-classification1M<n<10M0 likes87 downloads7mo agoHugging Face23patrikgerard /vk-post-comment-replies-translatedtabular1M<n<10M0 likes82 downloads9mo agoHugging Face24irfanhossainsust /bangla-youtube-comments-sentiment Bangla YouTube Comments Sentiment 271,902 public YouTube comments in Bangla: Bengali script and romanized Bangla ("Banglish"). Each comment is labeled negative, neutral or positive, for building and testing Bangla sentiment classifiers. These are silver (machine-generated) labels. The train/validation/test labels were predicted by an open model, not by human annotators. Expect about 78% of them to be correct overall, and about 84% of those with confidence >= 0.8. See Label… See the full description on the dataset page: https://huggingface.co/datasets/irfanhossainsust/bangla-youtube-comments-sentiment.tabulartext-classification100K<n<1M0 likes79 downloads11d agoHugging Face25gamusa /VOZ-HSD-Hate-Comments Dataset Card for Dataset Name A subset of VOZ-HSD dataset, consisting of only hate comments (labels: '1').For more information on the original dataset: https://huggingface.co/datasets/tarudesu/VOZ-HSD. NOTE: The original dataset is labeled automatically using fine-tuned ViSoBERT-HSD. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/gamusa/VOZ-HSD-Hate-Comments.tabulartext-classification100K<n<1M1 likes76 downloads5mo agoHugging Face26AccountVerify /swiss-law-commentary Swiss Law Commentary Dataset Open-access, AI-generated legal commentary on 8 Swiss federal laws. Source: openlegalcommentary.ch Laws: BV, ZGB, OR, ZPO, StGB, StPO, SchKG, VwVG Layers: Summary (B1 level), Doctrine (academic), Case Law (BGE digest) Languages: DE, FR, IT, EN Articles exported: 215 License: CC BY-SA 4.0 Built with ♥ by Jonas Hertner Files One JSONL file per law (e.g., or.jsonl, zgb.jsonl). Schema Field Type Description law… See the full description on the dataset page: https://huggingface.co/datasets/AccountVerify/swiss-law-commentary.tabular1K<n<10K0 likes76 downloads24d agoHugging Face27SubMaroon /DTF_Comments_Responses_CountsThis dataset contains data from mid-2016 to the end of 2024 from the website DTF.ru Structure:– post_title - body of the post;– parent_comment - parent comment :); – parent_author - author of parent comment;– child_comment - response (child) comment to parent comment;– child_author - author of child comment;– subsite_name - subsite name (like a theme);– comment_id_parent - id of parent comment on dtf.ru– comment_id_child - id of child comment on dtf.ru– replyTo - id of parent comment what… See the full description on the dataset page: https://huggingface.co/datasets/SubMaroon/DTF_Comments_Responses_Counts.tabular100K<n<1M0 likes70 downloads2y agoHugging Face28anitamaxvim /jigsaw-toxic-comments Dataset Card for Jigsaw Toxic Comments Dataset Dataset Description The Jigsaw Toxic Comments dataset is a benchmark dataset created for the Toxic Comment Classification Challenge on Kaggle. It is designed to help develop machine learning models that can identify and classify toxic online comments across multiple categories of toxicity. Curated by: Jigsaw (a technology incubator within Alphabet Inc.) Shared by: Kaggle Language(s) (NLP): English License: CC0 1.0… See the full description on the dataset page: https://huggingface.co/datasets/anitamaxvim/jigsaw-toxic-comments.tabulartext-classification100K<n<1M0 likes67 downloads1y agoHugging Face29tcapelle /jigsaw-toxic-comment-classification-challengetabular100K<n<1M1 likes66 downloads2y agoHugging Face30patrikgerard /edreddit-comments-with-roottabular1M<n<10M0 likes66 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.