Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01alexandrainst /nordjylland-news-summarization Dataset Card for "nordjylland-news-summarization" Dataset Summary This dataset consists of pairs containing text and corresponding summaries extracted from the Danish newspaper TV2 Nord. Supported Tasks and Leaderboards Summarization is the intended task for this dataset. No leaderboard is active at this point. Languages The dataset is available in Danish (da). Dataset Structure An example from the dataset looks as… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/nordjylland-news-summarization.tabularsummarization100K<n<1M2 likes485 downloads4mo agoHugging Face02jordiclive /scored_summarization_datasets Dataset Card for "Scored-Summarization-datasets" A collection of Text summarization datasets geared towards training a multi-purpose text summarizer. Each dataset is a parquet file with the following features. default text: a string feature. The source document summary: a string feature. The summary of the document provenance: a string feature. Information about the sub dataset. t5_text_token_count: a int64 feature. The number of tokens the text is encoded in.… See the full description on the dataset page: https://huggingface.co/datasets/jordiclive/scored_summarization_datasets.tabular1M<n<10M9 likes262 downloads4y agoHugging Face03whu9 /arxiv_summarization_postprocess Dataset Card for "arxiv_summarization_postprocess" More Information needed tabular100K<n<1M2 likes217 downloads3y agoHugging Face04YuvrajSingh9886 /reddit-posts-summarization-grpo GRPO Summarization Eval Rollouts Evaluation artifacts for all GRPO summarization checkpoints from smolcluster — a distributed GRPO training framework for Apple Silicon Mac clusters. Two base models were fine-tuned across two training strategies and six reward configurations each, then evaluated on 200 examples from the mlabonne/smoltldr test split. Judge: gpt-5-mini-2025-08-07 · Framework: DeepEval G-Eval · Rounds: 5 averaged · Metrics (each 0–1): Faithfulness · Coverage ·… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/reddit-posts-summarization-grpo.tabularsummarizationn<1K1 likes206 downloads3d agoHugging Face05pszemraj /govreport-summarization-8192 GovReport Summarization - 8192 tokens ccdv/govreport-summarization with the changes of: data cleaned with the clean-text python package total tokens for each column computed and added in new columns according to the long-t5 tokenizer (done after cleaning) train info RangeIndex: 8200 entries, 0 to 8199 Data columns (total 4 columns): # Column Non-Null Count Dtype --- ------ -------------- ----- 0 report 8200 non-null… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/govreport-summarization-8192.tabularsummarization10K<n<100K3 likes101 downloads10mo agoHugging Face06gsasikiran /collabllm-multiturn-scientific-papers-summarizationtabular10K<n<100K0 likes49 downloads10mo agoHugging Face07vltruong01 /amazon-all-beauty-led-summarization Amazon All_Beauty LED Summarization This dataset contains 2,000 product-level long-context review inputs with synthetic abstractive reference summaries for multi-review summarization. It is derived from the Amazon Reviews 2023 All_Beauty category and extends the source-only dataset: vltruong01/amazon-all-beauty-led-reviews Each example contains multiple selected customer reviews for one Amazon product, one synthetic teacher-generated consensus summary, and quality-control fields… See the full description on the dataset page: https://huggingface.co/datasets/vltruong01/amazon-all-beauty-led-summarization.tabularsummarization1K<n<10K0 likes45 downloads1mo agoHugging Face08bubbles01 /legal-document-summarization# Legal Document Summarization Dataset This dataset was prepared for a Legal Document Summarization project using Indian legal judgments. Dataset The dataset contains training and test records created from legal documents. Each record consists of a legal text chunk and its corresponding aligned summary. Split Records Train 1,287 Test 675 Total 1,962 Fields doc_id — Document identifier chunk_id — Chunk identifier within the document section —… See the full description on the dataset page: https://huggingface.co/datasets/bubbles01/legal-document-summarization.tabular1K<n<10K1 likes43 downloads10d agoHugging Face09mxlcw /telegram-financial-sentiment-summarizationtabulartext-classification10K<n<100K5 likes42 downloads2y agoHugging Face10shayekh /kazakh-news-summarization-20k-adapted This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. kazakh_news_summarization This dataset contains pairs of Kazakh language prompts and completions focused on summarizing news articles from sources like BAQ.KZ. The content covers diverse topics including social issues, legal cases, government initiatives, and international events within Kazakhstan and abroad. Each entry consists of a standard instruction to summarize text… See the full description on the dataset page: https://huggingface.co/datasets/shayekh/kazakh-news-summarization-20k-adapted.tabular10K<n<100K0 likes42 downloads4mo agoHugging Face11SahmBenchmark /financial-reports-extractive-summarization_eval Financial Reports Extractive Summarization Evaluation Dataset Validation and test splits for evaluating models on Arabic financial reports extractive summarization. Dataset Structure Format: Simple prompt-answer pairs Validation: ~20 examples (10%) Test: ~20 examples (10%) Language: Arabic Domain: Financial reports and market news Fields id: Unique identifier prompt: The summarization prompt full_text: Complete financial report answer: Ground… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/financial-reports-extractive-summarization_eval.tabularsummarizationn<1K0 likes32 downloads9mo agoHugging Face12hans1k /sinhala-summarization-dataset Sinhala Text Summarization Dataset Dataset Description This dataset is a Sinhala text summarization dataset created for research in low-resource language summarization. The dataset contains 2,493 Sinhala article-summary pairs collected from diverse publicly accessible Sinhala online sources. This repository contains a Sinhala article-summary dataset introduced in the following IEEE conference publication: Sinhala Automatic Text Summarization: Dataset Creation and… See the full description on the dataset page: https://huggingface.co/datasets/hans1k/sinhala-summarization-dataset.tabularsummarization1K<n<10K0 likes29 downloads5mo agoHugging Face13phuongntc /reference-free-rl-summarization-data Reference-free RL Summarization Experimental Data This repository contains experimental data splits, metadata, and processed subsets used for a study on verifier-composable penalty-shaped reinforcement learning for reference-free summarization. Configs vnexpress: Vietnamese VnExpress train/validation/test split used in the study. Unless explicit redistribution permission is available, this config releases metadata and split information only. cnn_dailymail_subset:… See the full description on the dataset page: https://huggingface.co/datasets/phuongntc/reference-free-rl-summarization-data.tabularsummarization10K<n<100K0 likes29 downloads4mo agoHugging Face14sorenmulli /nordjylland-news-summarization-subset [WIP] Dataset Card for "nordjylland-news-summarization-subset" Please note that this dataset and dataset card both are works in progress. For now refer to the related thesis for all details tabularn<1K0 likes27 downloads3y agoHugging Face15Lakshan2003 /Qwen3-4B-customerservice-context-summarization-llm-judge-data customer-service-context-summarization-evaluation-data Qwen3-4B-customerservice-context-summarization-llm-judge-data Dataset updated with context summarization evaluation columns. This README refresh triggers Hugging Face metadata re-index. tabular1K<n<10K0 likes25 downloads7mo agoHugging Face16SahmBenchmark /financial-reports-extractive-summarization_train Financial Reports Extractive Summarization Training Dataset Training split of the Arabic financial reports extractive summarization dataset in conversational format. Dataset Structure Format: Conversational (human-agent pairs) Size: ~160 training examples (80% of total) Language: Arabic Domain: Financial reports and market news Features id: Unique identifier conversations: Human prompt and agent summary report_type: Type of financial report… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/financial-reports-extractive-summarization_train.tabularsummarizationn<1K0 likes24 downloads10mo agoHugging Face17ibm-research /Wish-Summarization-Falcon Dataset Card for "Wish-Summarization-Falcon" More Information needed tabular10K<n<100K0 likes23 downloads3y agoHugging Face18vanya-robot /russian_summarizationtabularsummarization100K<n<1M4 likes20 downloads1y agoHugging Face19tBiski /Openai_Summarization_Preference_Geminitabular10K<n<100K0 likes20 downloads1y agoHugging Face20jbeiroa /resume-summarization-dataset Resume Summarization Dataset This dataset contains machine-generated summaries of 14,505 resumes using gpt-4o-mini. Each entry includes the original resume and a markdown-formatted summary divided into 5 sections. Structure Each row is a JSON object with: resume: The original resume text summary: The structured markdown summary input_tokens and output_tokens: (optional) token usage info License Some portions of this dataset are derived from public sources… See the full description on the dataset page: https://huggingface.co/datasets/jbeiroa/resume-summarization-dataset.tabularsummarization10K<n<100K1 likes19 downloads1y agoHugging Face21Lakshan2003 /gemini-2.5-flash-customerservice-context-summarization-llm-judge-data customer-service-context-summarization-evaluation-data Lakshan2003/gemini-2.5-flash-customerservice-context-summarization-llm-judge-data Dataset updated with context summarization evaluation columns. This README refresh triggers Hugging Face metadata re-index. tabular1K<n<10K0 likes19 downloads7mo agoHugging Face22divyajot5005 /plora-clean-summarization PLoRA Clean Multilingual Summarization This dataset is a cached, length-filtered training bundle for the local PLoRA notebook. It contains prompt/answer records whose full Qwen chat-formatted sequence length is at most 4096 tokens. Languages: hin_Deva (Hindi), fra_Latn (French), cmn_Hans (Chinese), urd_Arab (Urdu), eng_Latn (English), nld_Latn (Dutch), pol_Latn (Polish), snd_Arab (Sindhi), ben_Beng (Bengali), mar_Deva (Marathi) Counts: Train records: 100,000 Validation records: 8… See the full description on the dataset page: https://huggingface.co/datasets/divyajot5005/plora-clean-summarization.tabularsummarization100K<n<1M0 likes18 downloads5mo agoHugging Face23ibm-research /Wish-Summarization-Llama Dataset Card for "Wish-Summarization-Llama" More Information needed tabular10K<n<100K0 likes17 downloads3y agoHugging Face24punamvkhandar /multi_domain_summarization_datasettabular10K<n<100K0 likes16 downloads7mo agoHugging Face25Lakshan2003 /Llama3.2-3B-instruct-customerservice-context-summarization-llm-judge-data customer-service-context-summarization-evaluation-data Llama3.2-3B-instruct-customerservice-context-summarization-llm-judge-data Dataset updated with context summarization evaluation columns. This README refresh triggers Hugging Face metadata re-index. tabular1K<n<10K0 likes16 downloads7mo agoHugging Face26c00k1ez /summarization Dataset Card for "summarization" More Information needed tabularn<1K0 likes15 downloads4y agoHugging Face27iwasbinod /nepali-news-summarization-scraped-news-setopati-for-college-projecttabular1K<n<10K0 likes15 downloads2mo agoHugging Face28aduarte1 /openai_summarization_preference_gemini_justifiedtabularn<1K0 likes12 downloads1y agoHugging Face29Lakshan2003 /Llama3.1-8b-instruct-customerservice-context-summarization-llm-judge-data customer-service-context-summarization-evaluation-data Lakshan2003/Llama3.1-8b-instruct-customerservice-context-summarization-llm-judge-data Dataset updated with context summarization evaluation columns. This README refresh triggers Hugging Face metadata re-index. tabular1K<n<10K0 likes12 downloads7mo agoHugging Face30open-source-metrics /summarization-checkpoint-downloadstabularn<1K0 likes10 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.