Team Ai
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ccdv /cnn_dailymailCNN/DailyMail non-anonymized summarization dataset. There are two features: - article: text of news article, used as the document to be summarized - highlights: joined text of highlights with <s> and </s> around each highlight, which is the target summarysummarization100K<n<1M35 likes29k downloads4y agoHugging Face02RyokoAI /CNNovel125K Dataset Card for CNNovel125K The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible. Dataset Summary CNNovel125K is a dataset composed of approximately 125,000 novels downloaded from the Chinese novel hosting site http://ibiquw.com. Supported Tasks and Leaderboards This dataset is primarily intended for unsupervised training of text generation models; however, it may be useful for other… See the full description on the dataset page: https://huggingface.co/datasets/RyokoAI/CNNovel125K.text-classification100K<n<1M30 likes1.7k downloads4y agoHugging Face03botp /RyokoAI_CNNovel125K Dataset Card for CNNovel125K The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible. Dataset Summary CNNovel125K is a dataset composed of approximately 125,000 novels downloaded from the Chinese novel hosting site http://ibiquw.com. Supported Tasks and Leaderboards This dataset is primarily intended for unsupervised training of text generation models; however, it may be useful for other purposes.… See the full description on the dataset page: https://huggingface.co/datasets/botp/RyokoAI_CNNovel125K.texttext-classification1K<n<10K2 likes585 downloads3y agoHugging Face04qqceqqq /CNNovel125K Dataset Card for CNNovel125K The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible. Dataset Summary CNNovel125K is a dataset composed of approximately 125,000 novels downloaded from the Chinese novel hosting site http://ibiquw.com. Supported Tasks and Leaderboards This dataset is primarily intended for unsupervised training of text generation models; however, it may be useful for other purposes.… See the full description on the dataset page: https://huggingface.co/datasets/qqceqqq/CNNovel125K.texttext-classification10K<n<100K0 likes331 downloads6mo agoHugging Face05beiwoshuisheng /CNNovel125K Dataset Card for CNNovel125K The BigKnow2022 dataset and its subsets are not yet complete. Not all information here may be accurate or accessible. Dataset Summary CNNovel125K is a dataset composed of approximately 125,000 novels downloaded from the Chinese novel hosting site http://ibiquw.com. Supported Tasks and Leaderboards This dataset is primarily intended for unsupervised training of text generation models; however, it may be useful for other purposes.… See the full description on the dataset page: https://huggingface.co/datasets/beiwoshuisheng/CNNovel125K.texttext-classification10K<n<100K1 likes150 downloads9mo agoHugging Face06ilyasoulk /ai-vs-human-meta-llama-Llama-3.1-8B-Instruct-CNN AI vs Human dataset on the CNN DailyNews Dataset Description This dataset showcases pairs of truncated text and their respective completions, crafted either by humans or an AI language model. Each article was randomly truncated between 25% and 50% of its length. The language model was then tasked with generating a completion that mirrored the characters count of the original human-written continuation. Data Fields 'human': The original human-authored… See the full description on the dataset page: https://huggingface.co/datasets/ilyasoulk/ai-vs-human-meta-llama-Llama-3.1-8B-Instruct-CNN.texttext-classification1K<n<10K1 likes38 downloads2y agoHugging Face07PursuitOfDataScience /cnn-dailymail-llama4-maverick-summary CNN/DailyMail Summary Dataset (Llama-4-Maverick-17B-128E-Instruct-FP8) Dataset Description This dataset contains high-quality summaries of CNN and DailyMail news articles generated using the Llama-4-Maverick-17B-128E-Instruct-FP8 model. Each summary provides a concise, accurate overview of the main story while preserving key facts and context. Dataset Features High-quality summaries: Generated using Llama-4-Maverick-17B-128E-Instruct-FP8 model Comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/cnn-dailymail-llama4-maverick-summary.textsummarization100K<n<1M1 likes29 downloads1y agoHugging Face08Saleh11623 /CnnDailymailtexttext-classification100K<n<1M0 likes27 downloads2y agoHugging Face09ChamaraVishwajithRajapaksha /cnn-dailymail-sinhala-continuous-pretrain CNN DailyMail Sinhala Continuous Pretraining Dataset Dataset Description This dataset is designed for continuous pretraining of Sinhala Small Language Models (SLMs) and Large Language Models (LLMs). The dataset was created by processing the original Sinhala news articles from: CNN Daily Mail Sinhala Dataset The article_sinhala field from the original dataset was extracted, cleaned, and concatenated into larger continuous text blocks suitable for language model… See the full description on the dataset page: https://huggingface.co/datasets/ChamaraVishwajithRajapaksha/cnn-dailymail-sinhala-continuous-pretrain.texttext-generationn<1K0 likes19 downloads5mo agoHugging Face10keerti1104 /cnn_dailymailCNN/DailyMail non-anonymized summarization dataset. There are two features: - article: text of news article, used as the document to be summarized - highlights: joined text of highlights with <s> and </s> around each highlight, which is the target summarysummarization100K<n<1M0 likes15 downloads9mo agoHugging Face11llamafactory /cnn_dailymail_tinyThis dataset is a subset of https://huggingface.co/datasets/cnn_dailymail. The training set is composed of 2,000 examples of the original training set and the test set is composed of 1,000 examples of the original validation set. We use the version 1.0.0 of the CNN/DailyMail dataset. textsummarization1K<n<10K0 likes12 downloads3y agoHugging Face12ZhongshengWang /Alpaca-cnn-dailymail Data Summary Data set Alpaca-cnn-dailymail is a data set version format changed by ccdv/cnn_dailymail to meet Alpaca fine-tuning Llama2. Only versions 3.0.0 and 2.0.0 were used for merging and as a key data set for the summary extraction task. Licensing Information The Alpaca-cnn-dailymail dataset version 1.0.0 is released under the Apache-2.0 License. Citation Information @inproceedings{see-etal-2017-get, title = "Get To The Point: Summarization with… See the full description on the dataset page: https://huggingface.co/datasets/ZhongshengWang/Alpaca-cnn-dailymail.textsummarization100K<n<1M0 likes10 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.