Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google /civil_comments Dataset Card for "civil_comments" Dataset Summary The comments in this dataset come from an archive of the Civil Comments platform, a commenting plugin for independent news sites. These public comments were created from 2015 - 2017 and appeared on approximately 50 English-language news sites across the world. When Civil Comments shut down in 2017, they chose to make the public comments available in a lasting open archive to enable future research. The original data… See the full description on the dataset page: https://huggingface.co/datasets/google/civil_comments.tabulartext-classification1M<n<10M41 likes10k downloads3y agoHugging Face02thesofakillers /jigsaw-toxic-comment-classification-challenge Dataset Description You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior. The types of toxicity are: toxic severe_toxic obscene threat insult identity_hate You must create a model which predicts a probability of each type of toxicity for each comment. File descriptions train.csv - the training set, contains comments with their binary labels test.csv - the test set, you must predict the toxicity… See the full description on the dataset page: https://huggingface.co/datasets/thesofakillers/jigsaw-toxic-comment-classification-challenge.tabular100K<n<1M13 likes9.7k downloads2y agoHugging Face03HuggingFaceGECLM /REDDIT_comments Dataset Card for "REDDIT_comments" Dataset Summary Comments of 50 high-quality subreddits, extracted from the REDDIT PushShift data dumps (from 2006 to Jan 2023). Supported Tasks These comments can be used for text generation and language modeling, as well as dialogue modeling. Dataset Structure Data Splits Each split corresponds to a specific subreddit in the following list: "tifu", "explainlikeimfive", "WritingPrompts", "changemyview"… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceGECLM/REDDIT_comments.texttext-generation100M<n<1B27 likes8.1k downloads4y agoHugging Face04fddemarco /pushshift-reddit-comments Dataset Card for "pushshift-reddit" More Information needed tabular1B<n<10B27 likes4.6k downloads3y agoHugging Face05Helsinki-NLP /news_commentary Dataset Card for OPUS News-Commentary Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/news_commentary.texttranslation1M<n<10M39 likes3.4k downloads3y agoHugging Face06uisp /pali-commentary-thai-script-siamrath-version Multi-File CSV Dataset คำอธิบาย อรรถกถาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๘ เล่ม ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์ 01/010001.csv: เล่ม 1 หน้า 1 01/010002.csv: เล่ม 1 หน้า 2 ... 02/020001.csv: เล่ม 2 หน้า 1 คำอธิบายของแต่ละเล่ม เล่ม ๑: วินยฏฺกถา (สมนฺตปาสาทิกา ๑) เล่ม ๒: วินยฏฺกถา (สมนฺตปาสาทิกา ๒) เล่ม ๓: วินยฏฺกถา (สมนฺตปาสาทิกา ๓) เล่ม ๔: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๑) เล่ม ๕: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๒) เล่ม ๖: ทีฆนิกายฏฺกถา… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-commentary-thai-script-siamrath-version.tabular100K<n<1M3 likes1.2k downloads2y agoHugging Face07AISE-TUDelft /leading-comments Leading Comments We release Leading Comments, a collection of opening comment blocks from source code in The Stack v1, GitHub-Code, CodeParrot, The Pile, and RedPajama. The dataset brings together file-level license notices, copyright statements, redistribution conditions, and other introductory comments to support research on code licensing and dataset curation. Paper · Preprint · Code and replication package Getting started Install the Hugging Face Datasets… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/leading-comments.text100M<n<1B0 likes1.1k downloads9d agoHugging Face08TheGreatRambler /mm2_level_comments Mario Maker 2 level comments Part of the Mario Maker 2 Dataset Collection Dataset Description The Mario Maker 2 level comment dataset consists of 31.9 million level comments from Nintendo's online service totaling around 20GB of data. The dataset was created using the self-hosted Mario Maker 2 api over the course of 1 month in February 2022. How to use it The Mario Maker 2 level comment dataset is a very large dataset so for most use cases it is recommended to… See the full description on the dataset page: https://huggingface.co/datasets/TheGreatRambler/mm2_level_comments.tabularother10M<n<100M3 likes1.1k downloads4y agoHugging Face09community-datasets /danish_political_comments Dataset Card for DanishPoliticalComments Dataset Summary The dataset consists of 9008 sentences that are labeled with fine-grained polarity in the range from -2 to 2 (negative to positive). The quality of the fine-grained is not cross-validated and is therefore subject to uncertainties; however, the simple polarity has been cross-validated and therefore is considered to be more correct. Supported Tasks and Leaderboards [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/danish_political_comments.texttext-classification1K<n<10K3 likes851 downloads2y agoHugging Face10sentence-transformers /parallel-sentences-news-commentary Dataset Card for Parallel Sentences - News Commentary This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website. In particular, this dataset contains the News-Commentary dataset. Related Datasets The following datasets are also a part of the Parallel Sentences collection: parallel-sentences-europarl parallel-sentences-global-voices… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-news-commentary.textfeature-extraction1M<n<10M2 likes771 downloads2y agoHugging Face11nixiesearch /hackernews-comments Hackernews Comments Dataset A dataset of all HN API items from id=0 till id=41422887 (so from 2006 till 02 Sep 2024). The dataset is build by scraping the HN API according to its official schema and docs. Scraper code is also available on github: nixiesearch/hnscrape Dataset contents No cleaning, validation or filtering was performed. The resulting data files are raw JSON API response dumps in zstd-compressed JSONL files. An example payload: { "by": "goldfish"… See the full description on the dataset page: https://huggingface.co/datasets/nixiesearch/hackernews-comments.tabular10M<n<100M3 likes403 downloads2y agoHugging Face12lszoszk /treaty-bodies-general-comments Treaty Bodies General Comments A paragraph-level dataset of General Comments and General Recommendations adopted by the nine UN human-rights Treaty Bodies, with concerned-group labels and document metadata. Companion to the UNHRD search interface. Licence The curated dataset (paragraph segmentation, label annotation, document metadata enrichment, footnote and section work) is released under Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International… See the full description on the dataset page: https://huggingface.co/datasets/lszoszk/treaty-bodies-general-comments.tabulartext-classification1K<n<10K0 likes392 downloads1mo agoHugging Face13wecover /OPUS_News-Commentarytext1M<n<10M0 likes357 downloads3y agoHugging Face14touati-kamel /Algerian-Youtube-Comments Algerian YouTube Comments Dataset & Extraction Pipeline A large-scale, curated, and high-entropy NLP corpus of 55,365 authentic Algerian dialect comments collected from popular Algerian YouTube channels, cleaned through an 18-step NLP preprocessing engine, and classified using Google Gemini AI models to isolate genuine Algerian Darija (الدارجة الجزائرية) and Algerian Arabizi (العرابيزي). Hugging Face Hub: touati-kamel/Algerian-Youtube-Comments Primary Extraction Script:… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/Algerian-Youtube-Comments.texttext-classification10K<n<100K0 likes280 downloads12d agoHugging Face15MindCastSogang /Youtube_news_comments_analysis_ver2 YouTube 뉴스 댓글 감정 분석 — 최종 통합본 기존 분석과 2026-10-07 추가 분석을 댓글 생성일 기준으로 통합했습니다. 총 74,136,841개 댓글, 2014-04-21~2026-09-29, 관측 날짜 3,467일. 구조와 날짜 final/raw/json/YYYY/MM/01-10/news_comments.parquet final/raw/json/YYYY/MM/11-20/news_comments.parquet final/raw/json/YYYY/MM/21-말일/news_comments.parquet 경로와 파일 내 날짜 정렬 기준은 **dataset_date(댓글 작성일)**입니다. 윤년과 실제 말일을 반영합니다. news_date와 기존 호환용 date는 영상 게시일입니다. 댓글 시계열 집계에 사용하지 마세요. 신규 원본 comment_created_at, video_uploaded_at은 UTC 시각입니다.… See the full description on the dataset page: https://huggingface.co/datasets/MindCastSogang/Youtube_news_comments_analysis_ver2.tabulartext-classification10M<n<100M0 likes274 downloads13h agoHugging Face16alvanlii /reddit-comments-uwaterloo--- Generated Part of README Below --- Dataset Overview The goal is to have an open dataset of r/uwaterloo submissions, leveraging PRAW and the Reddit API to get downloads. Posts are here Comments are here Creation Details This dataset was created by alvanlii/dataset-creator-reddit-uwaterloo Update Frequency The dataset is updated custom with the most recent update being 2024-12-12 23:00:00 UTC+0000 where we added 72 new rows. Licensing… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/reddit-comments-uwaterloo.tabular1M<n<10M2 likes236 downloads2y agoHugging Face17AlexSham /Toxic_Russian_Commentshttps://www.kaggle.com/datasets/alexandersemiletov/toxic-russian-comments 0 - neutral user comments 1 - toxic user comments Toxic Russian Comments Dataset This dataset contains labelled comments from the popular Russian social network ok.ru. The data was used in a competition where participants had to automatically label each comment with at least one of the four predefined classes. The classes represent different levels of toxicity. The competition was held on the All Cups platform. Each… See the full description on the dataset page: https://huggingface.co/datasets/AlexSham/Toxic_Russian_Comments.texttext-classification100K<n<1M10 likes233 downloads3y agoHugging Face18andstor /smart_contract_code_commentstext100K<n<1M4 likes201 downloads3y agoHugging Face19algerian-nlp /Algerian-Youtube-Comments Algerian Youtube Comments 55,365 raw YouTube comments on Algeria-related videos for Darija social-text modeling, from the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/Algerian-Youtube-Comments: 55,365 train rows) and re-counted row-by-row with datasets streaming (load_dataset("algerian-nlp/Algerian-Youtube-Comments", split="train", streaming=True): 55,365 rows). The default config answers: how do Algerians actually write in… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/Algerian-Youtube-Comments.texttext-generation10K<n<100K0 likes197 downloads23d agoHugging Face20Jennyyahoo /RedditDataFor353_2026_COMMENTStabular1M<n<10M0 likes191 downloads2mo agoHugging Face21AmaanP314 /youtube-comment-sentiment YouTube Comments Sentiment Analysis Dataset (1M+ Labeled Comments) Overview This dataset comprises over one million YouTube comments, each annotated with sentiment labels—Positive, Neutral, or Negative. The comments span a diverse range of topics including programming, news, sports, politics and more, and are enriched with comprehensive metadata to facilitate various NLP and sentiment analysis tasks. How to use: import pandas as pd df =… See the full description on the dataset page: https://huggingface.co/datasets/AmaanP314/youtube-comment-sentiment.tabulartext-classification1M<n<10M5 likes175 downloads7mo agoHugging Face22AISE-TUDelft /multilingual-code-comments-fixed-8 Multilingual code comments This dataset contains 500 source-code examples for each of Chinese, Dutch, English, Greek and Polish (2,500 examples total), with comments generated by five models and human correctness ratings and error annotations. Each language has a train split. Human ratings use Correct, Partial and Incorrect. Error fields contain comma-separated taxonomy codes. A generated comment can carry multiple error codes. Annotation interpretation Stored… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/multilingual-code-comments-fixed-8.text1K<n<10K0 likes163 downloads10h agoHugging Face23affahrizain /jigsaw-toxic-comment Dataset Card for "jigsaw-toxic-comment" More Information needed text100K<n<1M4 likes152 downloads4y agoHugging Face24crawlora-net /tiktok-product-video-comments TikTok Product Video Comments 2026 999,930 TikTok comments and inline replies from 17,730 product videos across 12 shopping categories, collected September 10-13, 2026. This is the corpus behind Crawlora's study 17,774 TikTok 'Where Did You Get That?' Comments: Who Gets an Answer. Files File Rows Bytes SHA-256 data/videos.parquet 17,730 4,229,777 5aef621072ada8ff852007e810532c13069b3244ffae6c4fe47ff32549e83020 data/comments.parquet 999,930 81,410,979… See the full description on the dataset page: https://huggingface.co/datasets/crawlora-net/tiktok-product-video-comments.tabular1M<n<10M1 likes147 downloads27d agoHugging Face25oaaoaa /game_commentary_sftgatedtextn<1K1 likes139 downloads7mo agoHugging Face26Heliosoph /Jigsaw-Toxic-Comments Jigsaw Toxic Comments — canonical 2017–2018 Kaggle release A verbatim mirror of the Jigsaw Toxic Comment Classification Challenge training and test data, repackaged as three gzip-compressed CSVs. No rows added, removed, or reordered relative to the upstream Kaggle release — only the hosting moved and each file was gzipped. Re-hosted under Heliosoph for ingestion-pipeline stability — the upstream files live behind Kaggle's competition-rules click-through and require an… See the full description on the dataset page: https://huggingface.co/datasets/Heliosoph/Jigsaw-Toxic-Comments.tabulartext-classification100K<n<1M0 likes133 downloads4mo agoHugging Face27IvanFed /russian-toxic-comments-multilabel Russian Toxic Comments Multi-label Dataset Dataset Description Этот датасет содержит размеченные комментарии на русском языке для задачи многозадачной (multi-task) и мультилейбл (multi-label) бинарной классификации токсичности. Цель Обучение модели для автоматического обнаружения трех типов токсичного контента: Profanity (ненормативная лексика) — мат, оскорбления, нецензурная брань Threat (угрозы) — явные или скрытые угрозы в адрес других людей… See the full description on the dataset page: https://huggingface.co/datasets/IvanFed/russian-toxic-comments-multilabel.tabulartext-classification100K<n<1M1 likes133 downloads3mo agoHugging Face28Lots-of-LoRAs /task1720_civil_comments_toxicity_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1720_civil_comments_toxicity_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1720_civil_comments_toxicity_classification.texttext-generationn<1K0 likes131 downloads2y agoHugging Face29voilaj /swiss-law-commentary Swiss Law Commentary Dataset Open-access, AI-generated legal commentary on 8 Swiss federal laws. Source: openlegalcommentary.ch Laws: BV, ZGB, OR, ZPO, StGB, StPO, SchKG, VwVG Layers: Summary (B1 level), Doctrine (academic), Case Law (BGE digest) Languages: DE, FR, IT, EN Articles exported: 215 License: CC BY-SA 4.0 Built with ♥ by Jonas Hertner Files One JSONL file per law (e.g., or.jsonl, zgb.jsonl). Schema Field Type Description law… See the full description on the dataset page: https://huggingface.co/datasets/voilaj/swiss-law-commentary.tabular1K<n<10K1 likes126 downloads7mo agoHugging Face30abdulahad-dev /civil_comments Dataset Card for "civil_comments" Dataset Summary The comments in this dataset come from an archive of the Civil Comments platform, a commenting plugin for independent news sites. These public comments were created from 2015 - 2017 and appeared on approximately 50 English-language news sites across the world. When Civil Comments shut down in 2017, they chose to make the public comments available in a lasting open archive to enable future research. The original data… See the full description on the dataset page: https://huggingface.co/datasets/abdulahad-dev/civil_comments.tabulartext-classification1M<n<10M0 likes125 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.