Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nixiesearch /hackernews-comments Hackernews Comments Dataset A dataset of all HN API items from id=0 till id=41422887 (so from 2006 till 02 Sep 2024). The dataset is build by scraping the HN API according to its official schema and docs. Scraper code is also available on github: nixiesearch/hnscrape Dataset contents No cleaning, validation or filtering was performed. The resulting data files are raw JSON API response dumps in zstd-compressed JSONL files. An example payload: { "by": "goldfish"… See the full description on the dataset page: https://huggingface.co/datasets/nixiesearch/hackernews-comments.tabular10M<n<100M3 likes403 downloads2y agoHugging Face02AlexSham /Toxic_Russian_Commentshttps://www.kaggle.com/datasets/alexandersemiletov/toxic-russian-comments 0 - neutral user comments 1 - toxic user comments Toxic Russian Comments Dataset This dataset contains labelled comments from the popular Russian social network ok.ru. The data was used in a competition where participants had to automatically label each comment with at least one of the four predefined classes. The classes represent different levels of toxicity. The competition was held on the All Cups platform. Each… See the full description on the dataset page: https://huggingface.co/datasets/AlexSham/Toxic_Russian_Comments.texttext-classification100K<n<1M10 likes233 downloads3y agoHugging Face03oaaoaa /game_commentary_sftgatedtextn<1K1 likes139 downloads7mo agoHugging Face04voilaj /swiss-law-commentary Swiss Law Commentary Dataset Open-access, AI-generated legal commentary on 8 Swiss federal laws. Source: openlegalcommentary.ch Laws: BV, ZGB, OR, ZPO, StGB, StPO, SchKG, VwVG Layers: Summary (B1 level), Doctrine (academic), Case Law (BGE digest) Languages: DE, FR, IT, EN Articles exported: 215 License: CC BY-SA 4.0 Built with ♥ by Jonas Hertner Files One JSONL file per law (e.g., or.jsonl, zgb.jsonl). Schema Field Type Description law… See the full description on the dataset page: https://huggingface.co/datasets/voilaj/swiss-law-commentary.tabular1K<n<10K1 likes126 downloads7mo agoHugging Face05LLM-SocialMedia /Korean-YouTube-Comment-Sentiment-Dataset Korean YouTube Comment Sentiment Dataset Data Overview Summary 본 데이터셋은 유튜브에서 수집된 한국어 댓글 5,482개와 이에 대응하는 감정 레이블(긍정, 부정, 중립, 불명확)로 구성된 감정 분류용 데이터셋입니다. 주요 레이블: 긍정, 부정, 중립, 불명확 Features 수집 대상: 요리, 뷰티, 게임, 여행, 쇼핑 등 분야의 10만 명 이상 구독자를 보유한 유튜브 채널 형식: JSON (id, text, label) 검수: 한국인 검수자에 의한 수작업 라벨링 및 교차 검토 본 데이터셋은 구어체, 이모지, 줄임말 등 실제 사용자 표현이 반영되어 있습니다. Dataset Structure Dataset Fields Field Type Description id string 각 댓글의… See the full description on the dataset page: https://huggingface.co/datasets/LLM-SocialMedia/Korean-YouTube-Comment-Sentiment-Dataset.tabulartext-classification10K<n<100K3 likes118 downloads1y agoHugging Face06Leoooooops /douban_movie_commentstext100K<n<1M1 likes106 downloads2y agoHugging Face07kattymandy /cricket-commentary-ball-datatabular100K<n<1M0 likes104 downloads4mo agoHugging Face08sunny-chokshi /commentinject CommentInject: Testing Whether Code Comments Mislead AI Security Reviewers A benchmark of adversarial code-comment injection against LLM-based code vulnerability detectors. The vulnerable code stays exactly the same; only a comment claiming the code is safe is added. 30 vulnerable Python functions across 14 CWE categories (samples). 12 adversarial comment strategies in four families (strategies). 3,360 recorded trials from four studies, run locally through Ollama at temperature… See the full description on the dataset page: https://huggingface.co/datasets/sunny-chokshi/commentinject.texttext-classification1K<n<10K0 likes92 downloads7d agoHugging Face09AccountVerify /swiss-law-commentary Swiss Law Commentary Dataset Open-access, AI-generated legal commentary on 8 Swiss federal laws. Source: openlegalcommentary.ch Laws: BV, ZGB, OR, ZPO, StGB, StPO, SchKG, VwVG Layers: Summary (B1 level), Doctrine (academic), Case Law (BGE digest) Languages: DE, FR, IT, EN Articles exported: 215 License: CC BY-SA 4.0 Built with ♥ by Jonas Hertner Files One JSONL file per law (e.g., or.jsonl, zgb.jsonl). Schema Field Type Description law… See the full description on the dataset page: https://huggingface.co/datasets/AccountVerify/swiss-law-commentary.tabular1K<n<10K0 likes76 downloads24d agoHugging Face10ifmain /comment-translation-01This dataset is based on Kaggle. This dataset includes translations of 69,000 Reddit comments into 17 languages (English to 16 languages): Belarusian, Czech, German, English, Spanish, Finnish, French, Italian, Japanese, Kazakh, Korean, Latvian, Polish, Russian, Swedish, Ukrainian, and Chinese. It contains 50% regular comments and 50% highly negative ones. Enjoy using it! text1M<n<10M1 likes70 downloads2y agoHugging Face11ImpulseLeap /ru-comments-classificationtext10K<n<100K1 likes66 downloads3mo agoHugging Face12hausmer /truexa-comments Truha audience comments 45,918 real audience comments from the public satirical news channel «Труха⚡️Україна» (June 30 – July 22, 2026), collected from posts and cleaned (ads, links, mentions, ultra-short fragments, duplicates and phone numbers dropped — 50,000 raw → 45,918 clean). The audience writes in both Ukrainian and Russian (a mix, not a clean split — a share of comments code-switch between the two), so the corpus carries language: [ru, uk]. These are the "taste anchor"… See the full description on the dataset page: https://huggingface.co/datasets/hausmer/truexa-comments.tabulartext-classification10K<n<100K0 likes65 downloads25d agoHugging Face13brianhliou /youtube-comment-exposure YouTube Comment Exposure Corpus v1 1,345,353 top-level YouTube comments from 633 fully-enumerated videos across 34 channels, with the field the existing public corpora leave out: when each comment was posted. Why this exists The two large public YouTube comment datasets, YT-30M and YTCommentVerse (~32M comments between them), carry like counts with no timestamps and no reply counts. A like count without a post time is close to unusable for anything causal, because… See the full description on the dataset page: https://huggingface.co/datasets/brianhliou/youtube-comment-exposure.tabular1M<n<10M0 likes63 downloads2mo agoHugging Face14coreywoo27 /xhs-comment-responsetextn<1K0 likes61 downloads21d agoHugging Face15jslin09 /news_commentary_tw本資料集是來自QingySi所搜集的中英對照新聞評論,一共有 252,776 對中英語翻譯的句子,是使用Alpaca的指令資料集格式製成。本資料集利用了OpenCC 進行簡轉繁。 texttranslation100K<n<1M3 likes46 downloads3y agoHugging Face16theblackcat102 /bilibili_comments_sharegpt 林亦LYi B站留言 sharegpt 格式 把 train-test-validation 全合并了,因为使用上是混合其他对话资料训练,没有 overfitting 问题。如果你只是单训练这一份资料,小心overfitting 资料清理也把 B站表情符号去掉了,本来想保留但是无法都放到 system prompt 里,所以还是下次吧 text10K<n<100K6 likes37 downloads2y agoHugging Face17AnandforU /youtube-comment-insights-chatml YouTube Comment Insights - ChatML Overview This dataset contains instruction-tuning samples for structured YouTube comment analysis. The dataset is formatted in ChatML conversational format and is intended for supervised fine-tuning (SFT), QLoRA, and instruction tuning of large language models. Each sample contains: sentiment tone pros cons Dataset Statistics ~20k training samples ~2k validation samples Multilingual YouTube comments Structured JSON… See the full description on the dataset page: https://huggingface.co/datasets/AnandforU/youtube-comment-insights-chatml.texttext-classification10K<n<100K0 likes34 downloads5mo agoHugging Face18fewshot-goes-multilingual /cs_facebook-comments Dataset Card for Czech Facebook comments Dataset Description The dataset contains user comments from Facebook. Each comment contains text, sentiment (positive/negative/neutral). The dataset has in total (train+validation+test) 6,600 reviews. The data is balanced. Dataset Features Each sample contains: comment_id: unique string identifier of the comment. sentiment_str: string representation of the rating - "pozitivní" / "neutrální" / "negativní" sentiment_int:… See the full description on the dataset page: https://huggingface.co/datasets/fewshot-goes-multilingual/cs_facebook-comments.texttext-classification1K<n<10K0 likes32 downloads4y agoHugging Face19davanstrien /test-sync-commentstextn<1K0 likes30 downloads3y agoHugging Face20clockwork7 /reddit_news_articles_commentstext100K<n<1M2 likes26 downloads1y agoHugging Face21MentionBroker /reddit-comment-generation-v1 🤖 Official MentionBroker Research: Legacy V1 Comment Generation Dataset 📌 Dataset Summary This is an official Legacy V1 Release from the MentionBroker Research Lab. This corpus represents our foundational work in mapping Community Linguistic Variance and Conversational Response Efficacy. This dataset is specifically curated for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs) to generate high-authority, authentic community responses. By releasing this V1… See the full description on the dataset page: https://huggingface.co/datasets/MentionBroker/reddit-comment-generation-v1.tabulartext-classification100K<n<1M2 likes26 downloads7mo agoHugging Face22PingVortex /Youtube_shorts_comments Fine-tuned distilgpt2 on this dataset The average amount of emojis in a YouTube short comment is 4.59 (Based on this dataset) 12 millions view 😂😂😂 good god 🤦‍♂️ text1M<n<10M1 likes24 downloads2mo agoHugging Face23RinaL /lemmy-world-commentsThis is a data set of lemmy.world comments. text100K<n<1M0 likes22 downloads3y agoHugging Face24stm32Master /commenttext1K<n<10K0 likes22 downloads2y agoHugging Face25sedayzc /trendyol-electronics-products-features-and-commentstextn<1K0 likes19 downloads3mo agoHugging Face26speech-uk /generated-news-commentsThis dataset has been generated by an LLM. Convert to a DuckDB # 1) Download the JSONL file: wget "https://huggingface.co/datasets/speech-uk/generated-news-comments/resolve/main/generated_comments.jsonl" # 2) Open a DuckDB session in the terminal, then import the JSONL file into a table: duckdb $ CREATE TABLE generated_comments AS SELECT * FROM read_json_auto('generated_comments.jsonl'); # 3) Export the data from memory to a file: $ ATTACH 'my_database.db'; $ COPY FROM… See the full description on the dataset page: https://huggingface.co/datasets/speech-uk/generated-news-comments.text1K<n<10K1 likes17 downloads1y agoHugging Face27Alpaczyk /synthetic-football-commentary-qwen Synthetic Passionate Football Commentary Dataset Summary This dataset contains synthetic conversational data designed to fine-tune large language models for creative writing and persona adoption. Specifically, it trains models to act as a passionate football commentator. The data pairs factual football match events with highly dramatic, emotional, and tactical commentary. Data Generation Base Data: The raw input features (Minute, Match, Team, Player, Action)… See the full description on the dataset page: https://huggingface.co/datasets/Alpaczyk/synthetic-football-commentary-qwen.texttext-generation1K<n<10K0 likes17 downloads5mo agoHugging Face28sanka85 /rstp_comments_2text1K<n<10K0 likes15 downloads3y agoHugging Face29ELiRF /UX-commentsgated Cross-Domain Polarity Models to Evaluate User eXperience in E-learning Abstract Virtual learning environments are growing in importance as fast as e-learning, which is becoming highly demanded by universities and students worldwide. This paper investigates how to automatically evaluate User eXperience in this domain using sentiment analysis techniques. For this purpose, a corpus has been built with the opinions of 583 users (107 English speakers and 476 Spanish… See the full description on the dataset page: https://huggingface.co/datasets/ELiRF/UX-comments.texttext-classification1K<n<10K1 likes15 downloads2y agoHugging Face30gabeorlanski /eval-corm_black_comments Ranking Evaluation Dataset tabular10K<n<100K0 likes15 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.