datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
naver-news-summarization-ko
Naver-News-KO: A Korean News Summarization Dataset
A Korean news summarization dataset of 27,400 (document, summary) pairs, crawled from
Naver News over a ten-day window in July 2022. It was originally built for a
Korean NLP hands-on lab and has been publicly hosted on the Hugging Face Hub since January 2023.
A technical report documenting the collection protocol, corpus statistics, contamination analysis, and
reproducible baselines is available on arXiv: arXiv:2607.20442.… See the full description on the dataset page: https://huggingface.co/datasets/daekeun-ml/naver-news-summarization-ko.summarization-polish-summaries-corpussummarization-allegro-articlesgovreport-summarization-8192
GovReport Summarization - 8192 tokens
ccdv/govreport-summarization with the changes of:
data cleaned with the clean-text python package
total tokens for each column computed and added in new columns according to the long-t5 tokenizer (done after cleaning)
train info
RangeIndex: 8200 entries, 0 to 8199
Data columns (total 4 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 report 8200 non-null… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/govreport-summarization-8192.ro-text-summarizationentity-query-summarizationEDGAR-CORPUS-Financial-Summarization
EDGAR-CORPUS : 10K Financial Report Summarization
Extracted from SEC EDGAR filings (1993-2020). This dataset enhances financial report summarization by leveraging a hybrid AI model strategy.
Using:
ChatGPT-3.5 Turbo(~70%),
Claude 3.5 (~30% to generate structured, accurate, and concise summaries)
Dataset Composition
Summaries in this dataset are generated using a hybrid AI model strategy, balancing quality and efficiency:ChatGPT-3.5 Turbo (~70%) – Used for structured… See the full description on the dataset page: https://huggingface.co/datasets/kritsadaK/EDGAR-CORPUS-Financial-Summarization.vi-summarization
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/shannonnonshan/vi-summarization.hindi-article-summarization
Summary
hindi-article-summarization is an open source dataset of instruct-style records generated from the Hindi Text Short and Large Summarization dataset. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the CC BY-SA 4.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Hindi Version: 1.0
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/ganeshjcs/hindi-article-summarization.vietnamese_summarization_vr_vrp_resources
Vietnamese Summarization VR/VRP Resources
This repository consolidates the experimental resources associated with the paper:
Reinforcement Learning With Verifier Guidance and Penalty Shaping for Vietnamese Summarization Using Small Language Models
It contains:
CSV exports for Hugging Face Data Viewer,
Links to the released best checkpoints,
The link to the frozen evaluator MultiEvalSumViet2.
Representative Source Code
Dataset files used in the paper
Split… See the full description on the dataset page: https://huggingface.co/datasets/phuongntc/vietnamese_summarization_vr_vrp_resources.telegram-financial-sentiment-summarizationAMI-Corpus-Text-SummarizationMarathi_summarizationsinhala-summarization-dataset
Sinhala Text Summarization Dataset
Dataset Description
This dataset is a Sinhala text summarization dataset created for research in low-resource language summarization. The dataset contains 2,493 Sinhala article-summary pairs collected from diverse publicly accessible Sinhala online sources.
This repository contains a Sinhala article-summary dataset introduced in the following IEEE conference publication:
Sinhala Automatic Text Summarization: Dataset Creation and… See the full description on the dataset page: https://huggingface.co/datasets/hans1k/sinhala-summarization-dataset.summarizationMIMIC_IV_Summarization_SampleMIMIC_IV_summarization_shortautonlp-data-Ita-Summarizationcnn-summarization
cnn-summarization
1k rows randomly extracted from the cnn_dailymail dataset with reasoning traces generated by Qwen3-14b
simple-benchmark-arabic-summarizationnepali-summarization-datasetThis dataset was intended to to be used for finetuning the nepali text summerization task.
Feel free to contribute to this readme to add any information
three_line_summarization_for_japanese_news_articlesライブドアニュースコーパスの3行要約データセットです。
Llama v2向けのプロンプトを追加して成形してあります。
学習に利用する際は、 [R_START] [R_END] をspecial tokenとして追加することを推奨します。
Number of rows: 3,907
Datasetは以下のリポジトリを利用してscrapeしました。
git@github.com:KodairaTomonori/ThreeLineSummaryDataset.git
telugu-summarization-generation
Summary
aya-telugu-news-articles is an open source dataset of instruct-style records generated by webscraping a Telugu news articles website. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Telugu Version: 1.0
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/1-800-SHARED-TASKS/telugu-summarization-generation.synthetic-text-summarization-dataset-v1
Tanaos Text Summarization Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate text summarization systems — models that generate a concise, abstractive summary of a longer input text. It can be used to build summarization models for various applications, such as news summarization, document condensation, and content digestion.
Our flagship text summarization model… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-text-summarization-dataset-v1.rdf-summarization-dsummarization_evalmimic_note_summarizationro-text-summarizationko_summarization_linkbricks_single_dataset_with_prompt_text_huggingfaceFNS_Summarization
