summary
Datasets
All datasets matching “summary”news-summary
Dataset Card for "news-summary"
Dataset Summary
Officially it was supposed to be used for classification but, can you use this data set to summarize news articles?
Languages
english
Citation Information
Acknowledgements
Ahmed H, Traore I, Saad S. “Detecting opinion spams and fake news using text classification”, Journal of Security and Privacy, Volume 1, Issue 1, Wiley, January/February 2018.
Ahmed H, Traore I, Saad S. (2017) “Detection of Online… See the full description on the dataset page: https://huggingface.co/datasets/argilla/news-summary.hf_paper_summary
Paper Espresso Dataset
This dataset repository contains structured metadata, summaries, and topical analysis for trending AI research papers, as presented in the paper Paper Espresso: From Paper Overload to Research Insight.
Paper Espresso is an open-source platform designed to automatically discover, summarize, and analyze trending research papers from arXiv. The system uses large language models (LLMs) to generate structured summaries, topical labels, and keywords. Over 35… See the full description on the dataset page: https://huggingface.co/datasets/Elfsong/hf_paper_summary.asta-summary-citation-counts
Dataset Summary
This dataset tracks which scientific papers are most often cited by Asta, an agentic research platform that uses retrieval-augmented generation (RAG) to answer scientific questions. Each record is a paper cited by Asta's Summarize Literature tool, ranked by the number of times the system cited that paper. Across more than 113,000 user queries, we track 4M citations to over 2M distinct papers. By making this data public, we aim to create a transparent, trackable… See the full description on the dataset page: https://huggingface.co/datasets/allenai/asta-summary-citation-counts.medical-dialogue-to-soap-summary
Dataset Card for Synthetic Medical Dialogues and SOAP Summaries
Dataset Description
Abstract
This dataset consists of 10,000 synthetic dialogues between a patient and clinician, created using the GPT-4 dataset from NoteChat, based on PubMed Central (PMC) case-reports. Accompanying these dialogues are SOAP summaries generated through GPT-4. The dataset is split into 9250 training, 500 validation, and 250 test entries, each containing a dialogue column, a SOAP… See the full description on the dataset page: https://huggingface.co/datasets/omi-health/medical-dialogue-to-soap-summary.scientific-summary-distillation-Ornith-1.5-9B-736-20261004
Ornith 1.5 9B: 736 scientific paper generator summaries and actual reasoning
Frozen on 2026-10-04 at the user request. These are 736 successful raw generator
outputs from the same source cohort as the 865-paper Qwen distillation collection.
The remaining 129 papers were stopped; there is no claim that all 865 were completed.
Ten scientific domains; all records use the train split. No evaluation MCQs or
gold answers are included. The external 97-paper/970-MCQ evaluation is… See the full description on the dataset page: https://huggingface.co/datasets/ChristophSchuhmann/scientific-summary-distillation-Ornith-1.5-9B-736-20261004.distilled_sciinstruct_with_summary
