Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SuryaKrishna02 /aya-telugu-news-articles Summary aya-telugu-news-articles is an open source dataset of instruct-style records generated by webscraping a Telugu news articles website. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Telugu Version: 1.0 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-news-articles.texttext-generation100K<n<1M6 likes224 downloads3y agoHugging Face02SahandNZ /cryptonews-articles-with-price-momentum-labels Dataset Card for Cryptonews articles with price momentum labels Dataset Summary The dataset was gathered from two prominent sources in the cryptocurrency industry: Cryptonews.com and Binance.com. The aim of the dataset was to evaluate the impact of news on crypto price movements. As we know, news events such as regulatory changes, technological advancements, and major partnerships can have a significant impact on the price of cryptocurrencies. By analyzing the data… See the full description on the dataset page: https://huggingface.co/datasets/SahandNZ/cryptonews-articles-with-price-momentum-labels.texttext-classification100K<n<1M25 likes221 downloads3y agoHugging Face03allegro /summarization-allegro-articlestext100K<n<1M6 likes171 downloads5y agoHugging Face04vida-nyu /pmc-articles-dataset-mentions-snippets PMC Articles Dataset Mentions Snippets Text snippets from PubMed Central articles paired with structured dataset citations. Designed for training models to extract dataset references from scientific literature. Description Task: Extract structured dataset info (identifier, repository, webpage) from article text Source: PMC open-access articles Format: Text snippet → JSON output Examples: Positive (with datasets) and negative (no datasets) Fields… See the full description on the dataset page: https://huggingface.co/datasets/vida-nyu/pmc-articles-dataset-mentions-snippets.tabular1K<n<10K0 likes158 downloads3mo agoHugging Face05pierre-loic /climate-news-articles 🌍 Jeu de données d'articles de presse française labellisés comme traitant ou non des sujets liés au climat 🇬🇧 / 🇺🇸 : as this data set is based only on French data, all explanations are written in French in this repository. The goal of the dataset is to train a model to classify titles of French newspapers in two categories : if it's about climate or not. 🗺️ Le contexte Ce jeu de données de classification de titres d'article de presse française a été réalisé pour… See the full description on the dataset page: https://huggingface.co/datasets/pierre-loic/climate-news-articles.texttext-classification1K<n<10K1 likes114 downloads3y agoHugging Face06MIT-WAL /ai-jobs-news-articles-abstracts News articles and research abstracts on AI, labor, and jobs Dataset summary This file is a standalone CSV of news articles (full scraped text) and scholarly paper abstracts curated for research on artificial intelligence, work, and labor markets. Each row is one document: a stable id, publication date, normalized title and main text, and a small metadata dictionary. Rows: 53,526 document_class Rows Approx. date range (date column) news 29,857 Jan. 2025… See the full description on the dataset page: https://huggingface.co/datasets/MIT-WAL/ai-jobs-news-articles-abstracts.text10K<n<100K1 likes100 downloads2mo agoHugging Face07etrent17 /irs-articlestext1K<n<10K1 likes92 downloads4y agoHugging Face08haoxianc /gandpablo_news-articles-for-political-bias-classification News articles for political bias classification Mirror of the Kaggle dataset gandpablo/news-articles-for-political-bias-classification by Pablo Gandia, released under MIT. All credit goes to the original author; please cite and link the Kaggle page when using this data. Text from 10k+ English news articles classified by political bias License MIT License, Copyright (c) Pablo Gandia. The full license text is in LICENSE and applies to all files in this repository.… See the full description on the dataset page: https://huggingface.co/datasets/haoxianc/gandpablo_news-articles-for-political-bias-classification.text10K<n<100K0 likes91 downloads4d agoHugging Face09valurank /News_Articles_Categorization Dataset Card for News_Articles_Categorization Dataset Description 3722 News Articles classified into different categories namely: World, Politics, Tech, Entertainment, Sport, Business, Health, and Science Languages The text in the dataset is in English Dataset Structure The dataset consists of two columns namely Text and Category. The Text column consists of the news article and the Category column consists of the class each article belongs to… See the full description on the dataset page: https://huggingface.co/datasets/valurank/News_Articles_Categorization.texttext-classification1K<n<10K5 likes83 downloads3y agoHugging Face10fdaudens /ai-jobs-news-articles Dataset Summary This dataset brings together 1,000 English-language news articles all about the impact of artificial intelligence on jobs and the workforce. From automation to new tech-driven opportunities, these articles cover a wide range of perspectives and industries. It’s a great resource for anyone interested in how AI is shaping the future of work. Source Data The articles were collected from various reputable news outlets, focusing on recent developments and trends at the… See the full description on the dataset page: https://huggingface.co/datasets/fdaudens/ai-jobs-news-articles.tabulartext-classification1K<n<10K1 likes82 downloads1y agoHugging Face11AyoubChLin /CNN_News_Articles_2011-2022 CNN News Articles 2011-2022 Dataset Introduction This dataset contains CNN News Articles from 2011 to 2022 after basic cleaning. The dataset includes the following information: Category Full text The data was downloaded from Kaggle at this URL: https://www.kaggle.com/datasets/hadasu92/cnn-articles-after-basic-cleaning. The dataset was split into two sets: Train set with 32,218 examples Test set with 5,686 examples Usage This dataset can be used for… See the full description on the dataset page: https://huggingface.co/datasets/AyoubChLin/CNN_News_Articles_2011-2022.texttext-classification10K<n<100K11 likes79 downloads4y agoHugging Face12BASF-AI /ai4chem-clir-jrc-acquis-articles-lunatext10K<n<100K0 likes63 downloads8d agoHugging Face13hugginglearners /russia-ukraine-conflict-articles Dataset Card for Russia Ukraine Conflict Dataset Summary ###Context On 24 February 2022, Russia invaded Ukraine in a major escalation of the Russo-Ukrainian War that began in 2014. The invasion caused Europe's largest refugee crisis since World War II, with more than 6.3 million Ukrainians fleeing the country and a third of the population displaced (Source: Wikipedia). ###Content This dataset is a collection of 407 news articles from NYT and Guardians related to ongoing… See the full description on the dataset page: https://huggingface.co/datasets/hugginglearners/russia-ukraine-conflict-articles.textn<1K4 likes61 downloads4y agoHugging Face14Shwetasss /HinduTamil-News-Articles-Dataset HinduTamil News Articles Dataset Overview This dataset contains news articles in Tamil language scraped from the Hindu Tamil news website. Each article includes its title, author, city, published date, and text. Motivation This dataset was created to provide a comprehensive collection of Tamil news articles for research and analysis purposes. Data Sources and collection method The data in this dataset was collected from the Hindu Tamil news website… See the full description on the dataset page: https://huggingface.co/datasets/Shwetasss/HinduTamil-News-Articles-Dataset.texttext-classification10K<n<100K1 likes59 downloads3y agoHugging Face15hpe-ai /demo-articles-and-summarytextn<1K0 likes44 downloads3y agoHugging Face16prithivMLmods /Content-Articles Content-Articles Dataset Overview The Content-Articles dataset is a collection of academic articles and research papers across various subjects, including Computer Science, Physics, and Mathematics. This dataset is designed to facilitate research and analysis in these fields by providing structured data on article titles, abstracts, and subject classifications. Dataset Details Modalities Tabular: The dataset is structured in a tabular format. Text:… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Content-Articles.tabulartext-generation10K<n<100K3 likes41 downloads2y agoHugging Face17Navanjana /sinhala-articles Sinhala Articles Dataset A large-scale, high-quality Sinhala text corpus curated from diverse sources including news articles, Wikipedia entries, and general web content. This dataset is designed to support a wide range of Sinhala Natural Language Processing (NLP) tasks. 📊 Dataset Overview Name: Navanjana/sinhala-articles Total Samples: 2,148,688 Languages: Sinhala (si) Features: text: A single column containing Sinhala text passages. Size: Approximately 1M < n <… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/sinhala-articles.texttext-generation1M<n<10M1 likes39 downloads1y agoHugging Face18crawlfeeds /Medium-Articles-Corpus Medium Articles Corpus (10K Sample) The Medium Articles Corpus is a massive, clean dataset of articles scraped from Medium.com. This sample version contains 10,000 articles + and is designed to showcase the quality and structure of the full corpus for researchers and developers. This is the subset from the large dataset https://crawlfeeds.com/websites/medium/text_data/medium_articles Dataset Features This dataset includes the following key features, provided in a… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medium-Articles-Corpus.imagetext-classification10K<n<100K2 likes39 downloads1y agoHugging Face19crawlfeeds /Medical-Health-QA-Articles-Dataset Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development. Dataset Overview Field Details Sources iCliniq, HealthTap, WebMD Total Records 1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.imagetext-classification1K<n<10K0 likes37 downloads6mo agoHugging Face20BrightData /Wikipedia-Articles Dataset Card for "BrightData/Wikipedia-Articles" Dataset Summary Explore a collection of millions of Wikipedia articles with the Wikipedia dataset, comprising over 1.23M structured records and 10 data fields updated and refreshed regularly. Each entry includes all major data points such as timestamp, URLs, article titles, raw and cataloged text, images, "see also" references, external links, and a structured table of contents. For a complete list of data points, please… See the full description on the dataset page: https://huggingface.co/datasets/BrightData/Wikipedia-Articles.texttext-classification100K<n<1M7 likes36 downloads2y agoHugging Face21EmanuelNovelo /guardian_articles_full_contenttext1K<n<10K0 likes36 downloads2y agoHugging Face22kaengreg /wikifacts-articles-qrelstext1K<n<10K0 likes36 downloads2y agoHugging Face23freococo /moi-news-articles-dataset MOI News & Article Dataset 🇲🇲 This dataset contains over 16,000 cleaned news articles and feature stories extracted from the official website of the Ministry of Information (MOI) of Myanmar: moi.gov.mm. It is intended for use in news title generation, text classification, and Myanmar NLP research. The dataset is shared in the spirit of supporting freedom of information, language preservation, and the development of AI tools for the Burmese language (မြန်မာဘာသာ). 🗂️… See the full description on the dataset page: https://huggingface.co/datasets/freococo/moi-news-articles-dataset.texttext-classification10K<n<100K0 likes35 downloads1y agoHugging Face24lkarjun /Malayalam-Articlestext10K<n<100K1 likes34 downloads5y agoHugging Face25Johny201 /gdpr-articlestabularn<1K1 likes34 downloads4y agoHugging Face26haoxianc /hadasu92_cnn-articles-after-basic-cleaning CNN News Articles from 2011 to 2022 Mirror of the Kaggle dataset hadasu92/cnn-articles-after-basic-cleaning by Hadas Unger, released under CC0: Public Domain. All credit goes to the original author; please cite and link the Kaggle page when using this data. CNN News Articles from 2011 to 2022 after basic cleaning Original description (from Kaggle) Context This dataset contains around 38000 lines of articles from CNN news from the year 2011 to 2022.&nbsp;The data… See the full description on the dataset page: https://huggingface.co/datasets/haoxianc/hadasu92_cnn-articles-after-basic-cleaning.text10K<n<100K0 likes34 downloads5d agoHugging Face27HiTZ /ALIA_scientic_articles Dataset Description This dataset contains a high-quality, real (non-synthetic) parallel corpus in the scientific and academic domain, covering Basque (eu), Spanish (es), and English (en). Unlike synthetically generated datasets, this corpus is built from genuine, human-authored trilingual abstracts extracted from ADDI (Archivo Digital de Docencia e Investigación), the institutional digital repository of the University of the Basque Country (UPV/EHU). It is specifically designed… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/ALIA_scientic_articles.text1K<n<10K0 likes31 downloads4mo agoHugging Face28haoxianc /balabaskar_elon-musk-news-articles-corpora Elon Musk - News articles text corpora Mirror of the Kaggle dataset balabaskar/elon-musk-news-articles-corpora by Bala Baskar, released under CC0: Public Domain. All credit goes to the original author; please cite and link the Kaggle page when using this data. Text corpora of 5k news articles about Elon Musk Original description (from Kaggle) Introduction Elon Reeve Musk FRS is a business magnate and investor. He is the founder, CEO, and chief… See the full description on the dataset page: https://huggingface.co/datasets/haoxianc/balabaskar_elon-musk-news-articles-corpora.image1K<n<10K0 likes31 downloads5d agoHugging Face29bushra1dajam /news_articles News Articles Classification Dataset This dataset consists of news articles labeled with corresponding categories for classification tasks. Overview The news articles classification dataset is a collection of articles sourced from various news outlets, each labeled with a specific category. The dataset is designed for tasks such as text classification, topic modeling, and sentiment analysis. Dataset Information Name: News Articles Classification Dataset… See the full description on the dataset page: https://huggingface.co/datasets/bushra1dajam/news_articles.texttext-classification1K<n<10K2 likes30 downloads2y agoHugging Face30david-sprague /Medical-Health-QA-Articles-Dataset Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development. Dataset Overview Field Details Sources iCliniq, HealthTap, WebMD Total Records 1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/david-sprague/Medical-Health-QA-Articles-Dataset.imagetext-classification1K<n<10K0 likes29 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.