Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01abisee /cnn_dailymail Dataset Card for CNN Dailymail Dataset Dataset Summary The CNN / DailyMail Dataset is an English-language dataset containing just over 300k unique news articles as written by journalists at CNN and the Daily Mail. The current version supports both extractive and abstractive summarization, though the original version was created for machine reading and comprehension and abstractive question answering. Supported Tasks and Leaderboards 'summarization': Versions… See the full description on the dataset page: https://huggingface.co/datasets/abisee/cnn_dailymail.textsummarization100K<n<1M353 likes66k downloads3y agoHugging Face02ccdv /cnn_dailymailCNN/DailyMail non-anonymized summarization dataset. There are two features: - article: text of news article, used as the document to be summarized - highlights: joined text of highlights with <s> and </s> around each highlight, which is the target summarysummarization100K<n<1M35 likes29k downloads4y agoHugging Face03raymondt /cn_name0 likes21k downloads9mo agoHugging Face04sywang /CNNDetection6 likes5.1k downloads2y agoHugging Face05RyokoAI /CNNovel125K Dataset Card for CNNovel125K The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible. Dataset Summary CNNovel125K is a dataset composed of approximately 125,000 novels downloaded from the Chinese novel hosting site http://ibiquw.com. Supported Tasks and Leaderboards This dataset is primarily intended for unsupervised training of text generation models; however, it may be useful for other… See the full description on the dataset page: https://huggingface.co/datasets/RyokoAI/CNNovel125K.text-classification100K<n<1M30 likes1.7k downloads4y agoHugging Face06f3r21 /actas-cnn-datasetdocument0 likes847 downloads4mo agoHugging Face07notgoodkeeper /cnn-based-drowsiness-detection-data CNN-Based Drowsiness Detection - Dataset Preprocessed, auto-labeled face-crop images used to train the model in notgoodkeeper/cnn-based-drowsiness-detection. Code: https://github.com/not-good-keeper/cnn-based-drowsiness-detection Collection Frames were captured from a webcam, then run through: Haar Cascade face detection -> crop + pad + resize to 412x412 MediaPipe Selfie Segmentation -> background replaced with white CLAHE contrast normalization -> grayscale… See the full description on the dataset page: https://huggingface.co/datasets/notgoodkeeper/cnn-based-drowsiness-detection-data.imageimage-classification1K<n<10K1 likes791 downloads1mo agoHugging Face08dolphinteam /OpenWhistle-CNN OpenWhistle CNN Dataset dolphinteam/OpenWhistle-CNN is the public CNN dataset used for binary dolphin whistle detection. It contains audio windows, spectrogram images, and binary labels: noise (label=0) whistle (label=1) The main dataset is the complete session-disjoint dataset used for training and evaluation. A smaller deterministic review-sample config is also provided so reviewers can inspect representative examples quickly. Dataset contents Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/dolphinteam/OpenWhistle-CNN.audioaudio-classification10K<n<100K1 likes733 downloads11d agoHugging Face09botp /RyokoAI_CNNovel125K Dataset Card for CNNovel125K The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible. Dataset Summary CNNovel125K is a dataset composed of approximately 125,000 novels downloaded from the Chinese novel hosting site http://ibiquw.com. Supported Tasks and Leaderboards This dataset is primarily intended for unsupervised training of text generation models; however, it may be useful for other purposes.… See the full description on the dataset page: https://huggingface.co/datasets/botp/RyokoAI_CNNovel125K.texttext-classification1K<n<10K2 likes585 downloads3y agoHugging Face10qqceqqq /CNNovel125K Dataset Card for CNNovel125K The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible. Dataset Summary CNNovel125K is a dataset composed of approximately 125,000 novels downloaded from the Chinese novel hosting site http://ibiquw.com. Supported Tasks and Leaderboards This dataset is primarily intended for unsupervised training of text generation models; however, it may be useful for other purposes.… See the full description on the dataset page: https://huggingface.co/datasets/qqceqqq/CNNovel125K.texttext-classification10K<n<100K0 likes331 downloads6mo agoHugging Face11ttxy /cn_ner来源 https://github.com/liucongg/NLPDataSet 从网上收集数据,将CMeEE数据集、IMCS21_task1数据集、CCKS2017_task2数据集、CCKS2018_task1数据集、CCKS2019_task1数据集、CLUENER2020数据集、MSRA数据集、NLPCC2018_task4数据集、CCFBDCI数据集、MMC数据集、WanChuang数据集、PeopleDairy1998数据集、PeopleDairy2004数据集、GAIIC2022_task2数据集、WeiBo数据集、ECommerce数据集、FinanceSina数据集、BoSon数据集、Resume数据集、Bank数据集、FNED数据集和DLNER数据集等22个数据集进行整理清洗,构建一个较完善的中文NER数据集。 数据集清洗时,仅进行了简单地规则清洗,并将格式进行了统一化,标签为“BIO”。 处理后数据集详细信息,见数据集描述。 数据集由NJUST-TB一起整理。 由于部分数据包含嵌套实体的情况,所以转换成BIO标签时,长实体会覆盖短实体。 数据… See the full description on the dataset page: https://huggingface.co/datasets/ttxy/cn_ner.texttoken-classification100K<n<1M6 likes272 downloads3y agoHugging Face12imvladikon /he_cnn_dailymail Dataset Card for "he_cnn_dailymail" More Information needed textsummarization100K<n<1M0 likes271 downloads3y agoHugging Face13gigant /cnn_dailymail_oreo_jinacolbertv2_10ktext10K<n<100K0 likes244 downloads2y agoHugging Face14sheng2004 /CNNDetection-ProGan-4cls0 likes231 downloads11mo agoHugging Face15Subhav-K /cnn-dailymail-chunked-512-embeddingstext100K<n<1M2 likes219 downloads3mo agoHugging Face16Subhav-K /cnn-dailymail-nemotron-embeddingstext100K<n<1M1 likes201 downloads3mo agoHugging Face17ml6team /cnn_dailymail_nl This dataset is the CNN/Dailymail dataset translated to Dutch. This is the original dataset: ``` load_dataset("cnn_dailymail", '3.0.0') ``` And this is the HuggingFace translation pipeline: ``` pipeline( task='translation_en_to_nl', model='Helsinki-NLP/opus-mt-en-nl', tokenizer='Helsinki-NLP/opus-mt-en-nl') ```100K<n<1M14 likes179 downloads4y agoHugging Face18beiwoshuisheng /CNNovel125K Dataset Card for CNNovel125K The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible. Dataset Summary CNNovel125K is a dataset composed of approximately 125,000 novels downloaded from the Chinese novel hosting site http://ibiquw.com. Supported Tasks and Leaderboards This dataset is primarily intended for unsupervised training of text generation models; however, it may be useful for other purposes.… See the full description on the dataset page: https://huggingface.co/datasets/beiwoshuisheng/CNNovel125K.texttext-classification10K<n<100K1 likes150 downloads9mo agoHugging Face19S3IC /cnn_dailymail CNN_Dailymail This repository hosts a copy of the CNN_Dailymail dataset, a large-scale dataset designed for evaluating abstractive text summarization systems. CNN_Dailymail consists of news articles paired with human-written summaries, commonly used for training and evaluating models on summarization tasks. It contains articles from CNN and Daily Mail, covering a wide range of topics. Contents cnn_dailymail.jsonl (or your actual filename): the standard set of news… See the full description on the dataset page: https://huggingface.co/datasets/S3IC/cnn_dailymail.textsummarizationn<1K0 likes147 downloads9mo agoHugging Face20extraordinarylab /cnn-dailymailtext100K<n<1M0 likes146 downloads1y agoHugging Face21whu9 /cnn_dailymail_ngrams_1_to_5 Dataset Card for "cnn_dailymail_ngrams_1_to_5" More Information needed text100M<n<1B0 likes140 downloads4y agoHugging Face22GlazJ /cn-news-impact-scores Chinese News Impact Scores 2024-2025 This dataset pairs complete Chinese financial-news collections for 2024 and 2025 with event-level market, industry/board, and stock impact scores. Data is stored in monthly Parquet shards. Dataset Structure raw_news: every collected news occurrence, including full text, a unique occurrence_id, and a stable news_id. impact_scores: one row per event-target pair with routing metadata and 16 impact dimensions. event_clusters:… See the full description on the dataset page: https://huggingface.co/datasets/GlazJ/cn-news-impact-scores.tabulartext-classification1M<n<10M0 likes140 downloads3mo agoHugging Face23marco-willi /cnnspot-smallimage100K<n<1M0 likes137 downloads9mo agoHugging Face24cestwc /cnn_dailymail_test_outputstext100K<n<1M0 likes115 downloads4y agoHugging Face25tristantanjh /gtzan-multi-cnnimage0 likes102 downloads6mo agoHugging Face26zerrougi /SVHN_CNN_Specialist_Zoo0 likes100 downloads23d agoHugging Face27CxsGHost /CNNSum CNNSum: Exploring Long-Context Summarization with Large Language Models in Chinese Novels Paper     GitHub [2025.5] - Accepted to Findings of ACL 2025 [2025.1] - Add inference script [2024.12] - CNNSum Dataset Release We are excited to announce the release of the CNNSum dataset! As outlined in Section 3.1 and Appendix E of our paper, we have conducted a final round of manual cleaning to address any possible omissions. This process affects only a… See the full description on the dataset page: https://huggingface.co/datasets/CxsGHost/CNNSum.textsummarizationn<1K5 likes97 downloads1y agoHugging Face28Gabriel /cnn_daily_swe Dataset Card for Swedish CNN Dailymail Dataset The Swedish CNN/DailyMail dataset has only been machine-translated to improve downstream fine-tuning on Swedish summarization tasks. Dataset Summary Read about the full details at original English version: https://huggingface.co/datasets/cnn_dailymail Data Fields id: a string containing the heximal formated SHA1 hash of the url where the story was retrieved from article: a string containing the body of the news… See the full description on the dataset page: https://huggingface.co/datasets/Gabriel/cnn_daily_swe.textsummarization100K<n<1M0 likes94 downloads4y agoHugging Face29ereverter /cnn_dailymail_extractive Data Card for Extractive CNN/DailyMail Dataset Overview This is an extractive version of the CNN/Dailymail dataset. The structure of this dataset is identical to the original except for a minor modification in the data representation and the introduction of labels to denote the extractive summary. The labels are generated following a greedy algorithm, as proposed by Liu (2019). The curation process can be found in the bertsum-hf repository. I am uploading it in case… See the full description on the dataset page: https://huggingface.co/datasets/ereverter/cnn_dailymail_extractive.textsummarization100K<n<1M6 likes91 downloads3y agoHugging Face30celsowm /cnn_news_ptbr Dataset Card for "cnn_news_ptbr" More Information needed texttext-classification1K<n<10K3 likes90 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.