Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google /civil_comments Dataset Card for "civil_comments" Dataset Summary The comments in this dataset come from an archive of the Civil Comments platform, a commenting plugin for independent news sites. These public comments were created from 2015 - 2017 and appeared on approximately 50 English-language news sites across the world. When Civil Comments shut down in 2017, they chose to make the public comments available in a lasting open archive to enable future research. The original data… See the full description on the dataset page: https://huggingface.co/datasets/google/civil_comments.tabulartext-classification1M<n<10M41 likes10k downloads3y agoHugging Face02thesofakillers /jigsaw-toxic-comment-classification-challenge Dataset Description You are provided with a large number of Wikipedia comments which have been labeled by human raters for toxic behavior. The types of toxicity are: toxic severe_toxic obscene threat insult identity_hate You must create a model which predicts a probability of each type of toxicity for each comment. File descriptions train.csv - the training set, contains comments with their binary labels test.csv - the test set, you must predict the toxicity… See the full description on the dataset page: https://huggingface.co/datasets/thesofakillers/jigsaw-toxic-comment-classification-challenge.tabular100K<n<1M13 likes9.7k downloads2y agoHugging Face03HuggingFaceGECLM /REDDIT_comments Dataset Card for "REDDIT_comments" Dataset Summary Comments of 50 high-quality subreddits, extracted from the REDDIT PushShift data dumps (from 2006 to Jan 2023). Supported Tasks These comments can be used for text generation and language modeling, as well as dialogue modeling. Dataset Structure Data Splits Each split corresponds to a specific subreddit in the following list: "tifu", "explainlikeimfive", "WritingPrompts", "changemyview"… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceGECLM/REDDIT_comments.texttext-generation100M<n<1B27 likes8.1k downloads4y agoHugging Face04fddemarco /pushshift-reddit-comments Dataset Card for "pushshift-reddit" More Information needed tabular1B<n<10B27 likes4.6k downloads3y agoHugging Face05Helsinki-NLP /news_commentary Dataset Card for OPUS News-Commentary Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/news_commentary.texttranslation1M<n<10M39 likes3.4k downloads3y agoHugging Face06uisp /pali-commentary-thai-script-siamrath-version Multi-File CSV Dataset คำอธิบาย อรรถกถาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๘ เล่ม ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์ 01/010001.csv: เล่ม 1 หน้า 1 01/010002.csv: เล่ม 1 หน้า 2 ... 02/020001.csv: เล่ม 2 หน้า 1 คำอธิบายของแต่ละเล่ม เล่ม ๑: วินยฏฺกถา (สมนฺตปาสาทิกา ๑) เล่ม ๒: วินยฏฺกถา (สมนฺตปาสาทิกา ๒) เล่ม ๓: วินยฏฺกถา (สมนฺตปาสาทิกา ๓) เล่ม ๔: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๑) เล่ม ๕: ทีฆนิกายฏฺกถา (สุมงฺคลวิลาสินี ๒) เล่ม ๖: ทีฆนิกายฏฺกถา… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-commentary-thai-script-siamrath-version.tabular100K<n<1M3 likes1.2k downloads2y agoHugging Face07AISE-TUDelft /leading-comments Leading Comments We release Leading Comments, a collection of opening comment blocks from source code in The Stack v1, GitHub-Code, CodeParrot, The Pile, and RedPajama. The dataset brings together file-level license notices, copyright statements, redistribution conditions, and other introductory comments to support research on code licensing and dataset curation. Paper · Preprint · Code and replication package Getting started Install the Hugging Face Datasets… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/leading-comments.text100M<n<1B0 likes1.1k downloads9d agoHugging Face08TheGreatRambler /mm2_level_comments Mario Maker 2 level comments Part of the Mario Maker 2 Dataset Collection Dataset Description The Mario Maker 2 level comment dataset consists of 31.9 million level comments from Nintendo's online service totaling around 20GB of data. The dataset was created using the self-hosted Mario Maker 2 api over the course of 1 month in February 2022. How to use it The Mario Maker 2 level comment dataset is a very large dataset so for most use cases it is recommended to… See the full description on the dataset page: https://huggingface.co/datasets/TheGreatRambler/mm2_level_comments.tabularother10M<n<100M3 likes1.1k downloads4y agoHugging Face09jiebi /ids-paragraph-commentary-corpusThe data was collected from GitHub using https://github.com/cheop-byeon/RFCRationaleBuilder. This is our first version of data for RFC Rationale. Dataset Citation If you find this dataset useful and include it in your studies, please cite our paper: @inproceedings{bian2024tell, title={Tell Me Why: Language Models Help Explain the Rationale Behind Internet Protocol Design}, author={Bian, Jie and Welzl, Michael and Kutuzov, Andrey and Arefyev, Nikolay}, booktitle={2024 IEEE… See the full description on the dataset page: https://huggingface.co/datasets/jiebi/ids-paragraph-commentary-corpus.text-classification0 likes943 downloads6mo agoHugging Face10Jkatzy /code-comments-small Comment Dataset Opening comments extracted from code datasets with CommentMiner and ML4SE-toolkit. Files are grouped as <dataset>/<language>/part-*.parquet. The Hugging Face dataset card declares one config per source dataset and one split-safe language name per language. Each row contains dataset, record_id, opening_comment, language, path, repo, extracted_at, and metadata. For Parquet exports, metadata is stored as a JSON string so every source dataset shares one stable… See the full description on the dataset page: https://huggingface.co/datasets/Jkatzy/code-comments-small.0 likes895 downloads3mo agoHugging Face11community-datasets /danish_political_comments Dataset Card for DanishPoliticalComments Dataset Summary The dataset consists of 9008 sentences that are labeled with fine-grained polarity in the range from -2 to 2 (negative to positive). The quality of the fine-grained is not cross-validated and is therefore subject to uncertainties; however, the simple polarity has been cross-validated and therefore is considered to be more correct. Supported Tasks and Leaderboards [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/danish_political_comments.texttext-classification1K<n<10K3 likes851 downloads2y agoHugging Face12sentence-transformers /parallel-sentences-news-commentary Dataset Card for Parallel Sentences - News Commentary This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website. In particular, this dataset contains the News-Commentary dataset. Related Datasets The following datasets are also a part of the Parallel Sentences collection: parallel-sentences-europarl parallel-sentences-global-voices… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-news-commentary.textfeature-extraction1M<n<10M2 likes771 downloads2y agoHugging Face13nixiesearch /hackernews-comments Hackernews Comments Dataset A dataset of all HN API items from id=0 till id=41422887 (so from 2006 till 02 Sep 2024). The dataset is build by scraping the HN API according to its official schema and docs. Scraper code is also available on github: nixiesearch/hnscrape Dataset contents No cleaning, validation or filtering was performed. The resulting data files are raw JSON API response dumps in zstd-compressed JSONL files. An example payload: { "by": "goldfish"… See the full description on the dataset page: https://huggingface.co/datasets/nixiesearch/hackernews-comments.tabular10M<n<100M3 likes403 downloads2y agoHugging Face14lszoszk /treaty-bodies-general-comments Treaty Bodies General Comments A paragraph-level dataset of General Comments and General Recommendations adopted by the nine UN human-rights Treaty Bodies, with concerned-group labels and document metadata. Companion to the UNHRD search interface. Licence The curated dataset (paragraph segmentation, label annotation, document metadata enrichment, footnote and section work) is released under Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International… See the full description on the dataset page: https://huggingface.co/datasets/lszoszk/treaty-bodies-general-comments.tabulartext-classification1K<n<10K0 likes392 downloads1mo agoHugging Face15wecover /OPUS_News-Commentarytext1M<n<10M0 likes357 downloads3y agoHugging Face16RebuttalAgent /Comments_200K Comments-200k 📚 1. Introduction Raw reviews often intermix critical points with extraneous content, such as salutations and summaries. Directly feeding this unprocessed text into a model introduces significant noise and redundancy, which can compromise the precision of the generated rebuttal. Furthermore, due to diverse reviewer writing styles and varying conference formats, comments are typically presented in an unstructured manner. Therefore, to address these… See the full description on the dataset page: https://huggingface.co/datasets/RebuttalAgent/Comments_200K.2 likes318 downloads11mo agoHugging Face17touati-kamel /Algerian-Youtube-Comments Algerian YouTube Comments Dataset & Extraction Pipeline A large-scale, curated, and high-entropy NLP corpus of 55,365 authentic Algerian dialect comments collected from popular Algerian YouTube channels, cleaned through an 18-step NLP preprocessing engine, and classified using Google Gemini AI models to isolate genuine Algerian Darija (الدارجة الجزائرية) and Algerian Arabizi (العرابيزي). Hugging Face Hub: touati-kamel/Algerian-Youtube-Comments Primary Extraction Script:… See the full description on the dataset page: https://huggingface.co/datasets/touati-kamel/Algerian-Youtube-Comments.texttext-classification10K<n<100K0 likes280 downloads12d agoHugging Face18MindCastSogang /Youtube_news_comments_analysis_ver2 YouTube 뉴스 댓글 감정 분석 — 최종 통합본 기존 분석과 2026-10-07 추가 분석을 댓글 생성일 기준으로 통합했습니다. 총 74,136,841개 댓글, 2014-04-21~2026-09-29, 관측 날짜 3,467일. 구조와 날짜 final/raw/json/YYYY/MM/01-10/news_comments.parquet final/raw/json/YYYY/MM/11-20/news_comments.parquet final/raw/json/YYYY/MM/21-말일/news_comments.parquet 경로와 파일 내 날짜 정렬 기준은 **dataset_date(댓글 작성일)**입니다. 윤년과 실제 말일을 반영합니다. news_date와 기존 호환용 date는 영상 게시일입니다. 댓글 시계열 집계에 사용하지 마세요. 신규 원본 comment_created_at, video_uploaded_at은 UTC 시각입니다.… See the full description on the dataset page: https://huggingface.co/datasets/MindCastSogang/Youtube_news_comments_analysis_ver2.tabulartext-classification10M<n<100M0 likes274 downloads11h agoHugging Face19shlomihod /civil-comments-wildsIn this dataset, given a textual dialogue i.e. an utterance along with two previous turns of context, the goal was to infer the underlying emotion of the utterance by choosing from four emotion classes - Happy, Sad, Angry and Others.text-classification100K<n<1M0 likes253 downloads3y agoHugging Face20alvanlii /reddit-comments-uwaterloo--- Generated Part of README Below --- Dataset Overview The goal is to have an open dataset of r/uwaterloo submissions, leveraging PRAW and the Reddit API to get downloads. Posts are here Comments are here Creation Details This dataset was created by alvanlii/dataset-creator-reddit-uwaterloo Update Frequency The dataset is updated custom with the most recent update being 2024-12-12 23:00:00 UTC+0000 where we added 72 new rows. Licensing… See the full description on the dataset page: https://huggingface.co/datasets/alvanlii/reddit-comments-uwaterloo.tabular1M<n<10M2 likes236 downloads2y agoHugging Face21AlexSham /Toxic_Russian_Commentshttps://www.kaggle.com/datasets/alexandersemiletov/toxic-russian-comments 0 - neutral user comments 1 - toxic user comments Toxic Russian Comments Dataset This dataset contains labelled comments from the popular Russian social network ok.ru. The data was used in a competition where participants had to automatically label each comment with at least one of the four predefined classes. The classes represent different levels of toxicity. The competition was held on the All Cups platform. Each… See the full description on the dataset page: https://huggingface.co/datasets/AlexSham/Toxic_Russian_Comments.texttext-classification100K<n<1M10 likes233 downloads3y agoHugging Face22SalamBadak /fufufafa-comment-datasetimage1K<n<10K0 likes220 downloads3mo agoHugging Face23andstor /smart_contract_code_commentstext100K<n<1M4 likes201 downloads3y agoHugging Face24algerian-nlp /Algerian-Youtube-Comments Algerian Youtube Comments 55,365 raw YouTube comments on Algeria-related videos for Darija social-text modeling, from the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/Algerian-Youtube-Comments: 55,365 train rows) and re-counted row-by-row with datasets streaming (load_dataset("algerian-nlp/Algerian-Youtube-Comments", split="train", streaming=True): 55,365 rows). The default config answers: how do Algerians actually write in… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/Algerian-Youtube-Comments.texttext-generation10K<n<100K0 likes197 downloads23d agoHugging Face25Jennyyahoo /RedditDataFor353_2026_COMMENTStabular1M<n<10M0 likes191 downloads2mo agoHugging Face26dirtycomputer /Toxic_Comment_Classification_Challenge5 likes186 downloads3y agoHugging Face27ro-h /regulatory_commentsUnited States governmental agencies often make proposed regulations open to the public for comment. Proposed regulations are organized into "dockets". This project will use Regulation.gov public API to aggregate and clean public comments for dockets that mention opioid use. Each example will consist of one docket, and include metadata such as docket id, docket title, etc. Each docket entry will also include information about the top 10 comments, including comment metadata and comment text.text-classificationn<1K46 likes178 downloads3y agoHugging Face28AmaanP314 /youtube-comment-sentiment YouTube Comments Sentiment Analysis Dataset (1M+ Labeled Comments) Overview This dataset comprises over one million YouTube comments, each annotated with sentiment labels—Positive, Neutral, or Negative. The comments span a diverse range of topics including programming, news, sports, politics and more, and are enriched with comprehensive metadata to facilitate various NLP and sentiment analysis tasks. How to use: import pandas as pd df =… See the full description on the dataset page: https://huggingface.co/datasets/AmaanP314/youtube-comment-sentiment.tabulartext-classification1M<n<10M5 likes175 downloads7mo agoHugging Face29AISE-TUDelft /multilingual-code-comments-fixed-8 Multilingual code comments This dataset contains 500 source-code examples for each of Chinese, Dutch, English, Greek and Polish (2,500 examples total), with comments generated by five models and human correctness ratings and error annotations. Each language has a train split. Human ratings use Correct, Partial and Incorrect. Error fields contain comma-separated taxonomy codes. A generated comment can carry multiple error codes. Annotation interpretation Stored… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/multilingual-code-comments-fixed-8.text1K<n<10K0 likes163 downloads8h agoHugging Face30affahrizain /jigsaw-toxic-comment Dataset Card for "jigsaw-toxic-comment" More Information needed text100K<n<1M4 likes152 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.