Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google /civil_comments Dataset Card for "civil_comments" Dataset Summary The comments in this dataset come from an archive of the Civil Comments platform, a commenting plugin for independent news sites. These public comments were created from 2015 - 2017 and appeared on approximately 50 English-language news sites across the world. When Civil Comments shut down in 2017, they chose to make the public comments available in a lasting open archive to enable future research. The original data… See the full description on the dataset page: https://huggingface.co/datasets/google/civil_comments.tabulartext-classification1M<n<10M41 likes10k downloads3y agoHugging Face02HuggingFaceGECLM /REDDIT_comments Dataset Card for "REDDIT_comments" Dataset Summary Comments of 50 high-quality subreddits, extracted from the REDDIT PushShift data dumps (from 2006 to Jan 2023). Supported Tasks These comments can be used for text generation and language modeling, as well as dialogue modeling. Dataset Structure Data Splits Each split corresponds to a specific subreddit in the following list: "tifu", "explainlikeimfive", "WritingPrompts", "changemyview"… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceGECLM/REDDIT_comments.texttext-generation100M<n<1B27 likes8.1k downloads4y agoHugging Face03fddemarco /pushshift-reddit-comments Dataset Card for "pushshift-reddit" More Information needed tabular1B<n<10B27 likes4.6k downloads3y agoHugging Face04AISE-TUDelft /leading-comments Leading Comments We release Leading Comments, a collection of opening comment blocks from source code in The Stack v1, GitHub-Code, CodeParrot, The Pile, and RedPajama. The dataset brings together file-level license notices, copyright statements, redistribution conditions, and other introductory comments to support research on code licensing and dataset curation. Paper · Preprint · Code and replication package Getting started Install the Hugging Face Datasets… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/leading-comments.text100M<n<1B0 likes1.1k downloads9d agoHugging Face05TheGreatRambler /mm2_level_comments Mario Maker 2 level comments Part of the Mario Maker 2 Dataset Collection Dataset Description The Mario Maker 2 level comment dataset consists of 31.9 million level comments from Nintendo's online service totaling around 20GB of data. The dataset was created using the self-hosted Mario Maker 2 api over the course of 1 month in February 2022. How to use it The Mario Maker 2 level comment dataset is a very large dataset so for most use cases it is recommended to… See the full description on the dataset page: https://huggingface.co/datasets/TheGreatRambler/mm2_level_comments.tabularother10M<n<100M3 likes1.1k downloads4y agoHugging Face06Jkatzy /code-comments-small Comment Dataset Opening comments extracted from code datasets with CommentMiner and ML4SE-toolkit. Files are grouped as <dataset>/<language>/part-*.parquet. The Hugging Face dataset card declares one config per source dataset and one split-safe language name per language. Each row contains dataset, record_id, opening_comment, language, path, repo, extracted_at, and metadata. For Parquet exports, metadata is stored as a JSON string so every source dataset shares one stable… See the full description on the dataset page: https://huggingface.co/datasets/Jkatzy/code-comments-small.0 likes895 downloads3mo agoHugging Face07community-datasets /danish_political_comments Dataset Card for DanishPoliticalComments Dataset Summary The dataset consists of 9008 sentences that are labeled with fine-grained polarity in the range from -2 to 2 (negative to positive). The quality of the fine-grained is not cross-validated and is therefore subject to uncertainties; however, the simple polarity has been cross-validated and therefore is considered to be more correct. Supported Tasks and Leaderboards [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/danish_political_comments.texttext-classification1K<n<10K3 likes851 downloads2y agoHugging Face08nixiesearch /hackernews-comments Hackernews Comments Dataset A dataset of all HN API items from id=0 till id=41422887 (so from 2006 till 02 Sep 2024). The dataset is build by scraping the HN API according to its official schema and docs. Scraper code is also available on github: nixiesearch/hnscrape Dataset contents No cleaning, validation or filtering was performed. The resulting data files are raw JSON API response dumps in zstd-compressed JSONL files. An example payload: { "by": "goldfish"… See the full description on the dataset page: https://huggingface.co/datasets/nixiesearch/hackernews-comments.tabular10M<n<100M3 likes403 downloads2y agoHugging Face09lszoszk /treaty-bodies-general-comments Treaty Bodies General Comments A paragraph-level dataset of General Comments and General Recommendations adopted by the nine UN human-rights Treaty Bodies, with concerned-group labels and document metadata. Companion to the UNHRD search interface. Licence The curated dataset (paragraph segmentation, label annotation, document metadata enrichment, footnote and section work) is released under Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International… See the full description on the dataset page: https://huggingface.co/datasets/lszoszk/treaty-bodies-general-comments.tabulartext-classification1K<n<10K0 likes392 downloads1mo agoHugging Face10RebuttalAgent /Comments_200K Comments-200k 📚 1. Introduction Raw reviews often intermix critical points with extraneous content, such as salutations and summaries. Directly feeding this unprocessed text into a model introduces significant noise and redundancy, which can compromise the precision of the generated rebuttal. Furthermore, due to diverse reviewer writing styles and varying conference formats, comments are typically presented in an unstructured manner. Therefore, to address these… See the full description on the dataset page: https://huggingface.co/datasets/RebuttalAgent/Comments_200K.2 likes318 downloads11mo agoHugging Face11touati-kamel /Algerian-Youtube-Comments Algerian YouTube Comments Dataset & Extraction Pipeline A large-scale, curated, and high-entropy NLP corpus of 55,365 authentic Algerian dialect comments collected from popular Algerian YouTube channels, cleaned through an 18-step NLP preprocessing engine, and classified using Google Gemini AI models to isolate genuine Algerian Darija (الدارجة الجزائرية) and Algerian Arabizi (العرابيزي). Hugging Face Hub: touati-kamel/Algerian-Youtube-Comments Primary Extraction Script:… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/Algerian-Youtube-Comments.texttext-classification10K<n<100K0 likes280 downloads13d agoHugging Face12MindCastSogang /Youtube_news_comments_analysis_ver2 YouTube 뉴스 댓글 감정 분석 — 최종 통합본 기존 분석과 2026-10-07 추가 분석을 댓글 생성일 기준으로 통합했습니다. 총 74,136,841개 댓글, 2014-04-21~2026-09-29, 관측 날짜 3,467일. 구조와 날짜 final/raw/json/YYYY/MM/01-10/news_comments.parquet final/raw/json/YYYY/MM/11-20/news_comments.parquet final/raw/json/YYYY/MM/21-말일/news_comments.parquet 경로와 파일 내 날짜 정렬 기준은 **dataset_date(댓글 작성일)**입니다. 윤년과 실제 말일을 반영합니다. news_date와 기존 호환용 date는 영상 게시일입니다. 댓글 시계열 집계에 사용하지 마세요. 신규 원본 comment_created_at, video_uploaded_at은 UTC 시각입니다.… See the full description on the dataset page: https://huggingface.co/datasets/MindCastSogang/Youtube_news_comments_analysis_ver2.tabulartext-classification10M<n<100M0 likes274 downloads15h agoHugging Face13shlomihod /civil-comments-wildsIn this dataset, given a textual dialogue i.e. an utterance along with two previous turns of context, the goal was to infer the underlying emotion of the utterance by choosing from four emotion classes - Happy, Sad, Angry and Others.text-classification100K<n<1M0 likes253 downloads3y agoHugging Face14alvanlii /reddit-comments-uwaterloo--- Generated Part of README Below --- Dataset Overview The goal is to have an open dataset of r/uwaterloo submissions, leveraging PRAW and the Reddit API to get downloads. Posts are here Comments are here Creation Details This dataset was created by alvanlii/dataset-creator-reddit-uwaterloo Update Frequency The dataset is updated custom with the most recent update being 2024-12-12 23:00:00 UTC+0000 where we added 72 new rows. Licensing… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/reddit-comments-uwaterloo.tabular1M<n<10M2 likes236 downloads2y agoHugging Face15AlexSham /Toxic_Russian_Commentshttps://www.kaggle.com/datasets/alexandersemiletov/toxic-russian-comments 0 - neutral user comments 1 - toxic user comments Toxic Russian Comments Dataset This dataset contains labelled comments from the popular Russian social network ok.ru. The data was used in a competition where participants had to automatically label each comment with at least one of the four predefined classes. The classes represent different levels of toxicity. The competition was held on the All Cups platform. Each… See the full description on the dataset page: https://huggingface.co/datasets/AlexSham/Toxic_Russian_Comments.texttext-classification100K<n<1M10 likes233 downloads3y agoHugging Face16andstor /smart_contract_code_commentstext100K<n<1M4 likes201 downloads3y agoHugging Face17algerian-nlp /Algerian-Youtube-Comments Algerian Youtube Comments 55,365 raw YouTube comments on Algeria-related videos for Darija social-text modeling, from the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/Algerian-Youtube-Comments: 55,365 train rows) and re-counted row-by-row with datasets streaming (load_dataset("algerian-nlp/Algerian-Youtube-Comments", split="train", streaming=True): 55,365 rows). The default config answers: how do Algerians actually write in… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/Algerian-Youtube-Comments.texttext-generation10K<n<100K0 likes197 downloads24d agoHugging Face18Jennyyahoo /RedditDataFor353_2026_COMMENTStabular1M<n<10M0 likes191 downloads2mo agoHugging Face19ro-h /regulatory_commentsUnited States governmental agencies often make proposed regulations open to the public for comment. Proposed regulations are organized into "dockets". This project will use Regulation.gov public API to aggregate and clean public comments for dockets that mention opioid use. Each example will consist of one docket, and include metadata such as docket id, docket title, etc. Each docket entry will also include information about the top 10 comments, including comment metadata and comment text.text-classificationn<1K46 likes178 downloads3y agoHugging Face20AISE-TUDelft /multilingual-code-comments-fixed-8 Multilingual code comments This dataset contains 500 source-code examples for each of Chinese, Dutch, English, Greek and Polish (2,500 examples total), with comments generated by five models and human correctness ratings and error annotations. Each language has a train split. Human ratings use Correct, Partial and Incorrect. Error fields contain comma-separated taxonomy codes. A generated comment can carry multiple error codes. Annotation interpretation Stored… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/multilingual-code-comments-fixed-8.text1K<n<10K0 likes163 downloads12h agoHugging Face21crawlora-net /tiktok-product-video-comments TikTok Product Video Comments 2026 999,930 TikTok comments and inline replies from 17,730 product videos across 12 shopping categories, collected September 10-13, 2026. This is the corpus behind Crawlora's study 17,774 TikTok 'Where Did You Get That?' Comments: Who Gets an Answer. Files File Rows Bytes SHA-256 data/videos.parquet 17,730 4,229,777 5aef621072ada8ff852007e810532c13069b3244ffae6c4fe47ff32549e83020 data/comments.parquet 999,930 81,410,979… See the full description on the dataset page: https://huggingface.co/datasets/crawlora-net/tiktok-product-video-comments.tabular1M<n<10M1 likes147 downloads27d agoHugging Face22Jalik /TrojanLoc-gpt-4.1-embeddings-wo-comments0 likes135 downloads11mo agoHugging Face23Heliosoph /Jigsaw-Toxic-Comments Jigsaw Toxic Comments — canonical 2017–2018 Kaggle release A verbatim mirror of the Jigsaw Toxic Comment Classification Challenge training and test data, repackaged as three gzip-compressed CSVs. No rows added, removed, or reordered relative to the upstream Kaggle release — only the hosting moved and each file was gzipped. Re-hosted under Heliosoph for ingestion-pipeline stability — the upstream files live behind Kaggle's competition-rules click-through and require an… See the full description on the dataset page: https://huggingface.co/datasets/Heliosoph/Jigsaw-Toxic-Comments.tabulartext-classification100K<n<1M0 likes133 downloads4mo agoHugging Face24IvanFed /russian-toxic-comments-multilabel Russian Toxic Comments Multi-label Dataset Dataset Description Этот датасет содержит размеченные комментарии на русском языке для задачи многозадачной (multi-task) и мультилейбл (multi-label) бинарной классификации токсичности. Цель Обучение модели для автоматического обнаружения трех типов токсичного контента: Profanity (ненормативная лексика) — мат, оскорбления, нецензурная брань Threat (угрозы) — явные или скрытые угрозы в адрес других людей… See the full description on the dataset page: https://huggingface.co/datasets/IvanFed/russian-toxic-comments-multilabel.tabulartext-classification100K<n<1M1 likes133 downloads3mo agoHugging Face25Lots-of-LoRAs /task1720_civil_comments_toxicity_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1720_civil_comments_toxicity_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1720_civil_comments_toxicity_classification.texttext-generationn<1K0 likes131 downloads2y agoHugging Face26Tungtom2004 /850_commentsgatedimage1K<n<10K0 likes126 downloads1mo agoHugging Face27abdulahad-dev /civil_comments Dataset Card for "civil_comments" Dataset Summary The comments in this dataset come from an archive of the Civil Comments platform, a commenting plugin for independent news sites. These public comments were created from 2015 - 2017 and appeared on approximately 50 English-language news sites across the world. When Civil Comments shut down in 2017, they chose to make the public comments available in a lasting open archive to enable future research. The original data… See the full description on the dataset page: https://huggingface.co/datasets/abdulahad-dev/civil_comments.tabulartext-classification1M<n<10M0 likes125 downloads5mo agoHugging Face28ro-h /regulatory_comments_apiUnited States governmental agencies often make proposed regulations open to the public for comment. Proposed regulations are organized into "dockets". This dataset will use Regulation.gov public API to aggregate and clean public comments for dockets that mention opioid use. Each example will consist of one docket, and include metadata such as docket id, docket title, etc. Each docket entry will also include information about the top 10 comments, including comment metadata and comment text.text-classificationn<1K41 likes116 downloads3y agoHugging Face29lighteval /civil_comments_helmtext100K<n<1M1 likes108 downloads1y agoHugging Face30gorpeliates /swe-bench-commentsIncludes the SWE-bench dataset along with the corresponding comments for each instance_id text1K<n<10K0 likes108 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.