datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tech-articles-live持续收集国内各个科技公众号的文章并提供简易分析工具,目前工具位于新智元公众号文件夹下。
使用的自己修改的微信公众号文章导出器,项目链接:我的GitHub仓库链接
android-times-articles
Android Times — Articles Dataset Archive
Private dataset containing synthesized and processed article archives, multi-language transcripts, metadata, and editorial assets for Android Times.
Dataset Structure
articles/
├── en-US/ # English (United States) localized articles & scripts
├── ja-JP/ # Japanese localized articles & scripts
├── en-AU/ # Australian localized articles
├── en-CN/ # China localized English… See the full description on the dataset page: https://huggingface.co/datasets/aoiandroid/android-times-articles.filtered_articles_by_year
Dataset Card for Filtered Articles by Year
Dataset Summary
The Filtered Articles by Year dataset contains yearly-segmented web articles from the FineWeb dataset, specifically filtered and processed for temporal language analysis and Word2Vec model training. This dataset spans 21 years (2005-2025) and serves as the foundation for research into semantic change, concept emergence, and language evolution over time.
Supported Tasks and Leaderboards
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/adameubanks/filtered_articles_by_year.sdear-articlespolitical_bias_in_news_articlesBangla_Financial_news_articles_Dataset
Bangla-Financial-news-articles-Dataset
A Comprehensive Resource for Analyzing Sentiments in over 7600+ Bangla News.
Downloads
🔴 Download the "💥Bangla_fin_news.zip" file for all "7,695" news and extract it.
About Dataset
Welcome to our Bengali Financial News Sentiment Analysis dataset! This collection comprises 7,695 financial news articles extracted, covering the period from March 3, 2014, to December 29, 2021. Utilizing the powerful web scraping tool… See the full description on the dataset page: https://huggingface.co/datasets/ashtrayAI/Bangla_Financial_news_articles_Dataset.kurdish-wikipedia-articles
Summary
Extracted from the wikidump. There are summaries and categories available for each article. Will look into adding them later.
Usage
from datasets import load_dataset
ds = load_dataset("nazimali/kurdish-wikipedia-articles", split="train")
ds
Dataset({
features: ['id', 'url', 'title', 'text'],
num_rows: 63076
})
ticker_analysis_articlesfinancial-news-articles
Dataset Card for "financial-news-articles"
More Information needed
The data was obtained from here
german-wikipedia-articlesnews-articles-ptbr-dataset
Dataset Card for "news-articles-ptbr-dataset"
More Information needed
pubmed_articles_domain10x_20250227medium-articles
Data source
This data has been collected through a standard scraping process from the Medium website, looking for published articles.
Data description
Each row in the data is a different article published on Medium. For each article, you have the following features:
title [string]: The title of the article.
text [string]: The text content of the article.
url [string]: The URL associated to the article.
authors [list of strings]: The article authors.
timestamp [string]:… See the full description on the dataset page: https://huggingface.co/datasets/fabiochiu/medium-articles.ai-tech-articles
AI/Tech Dataset
This dataset is a collection of AI/tech articles scraped from the web: the 2023
corpus collected by the original Haskell scraper plus everything the Rust
crawler has added since (2023 onwards, with full publication dates).
It's hosted on HuggingFace Datasets, so it is easier to load in and work with.
To load the dataset
1. Install HuggingFace Datasets
pip install datasets
2. Load the dataset
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/siavava/ai-tech-articles.medium-articles-posts-with-content
Medium Articles Dataset Generator
This project combines multiple datasets from Kaggle and Hugging Face to create a comprehensive collection of Medium articles. The combined dataset is available on Hugging Face Hub.
Dataset Description
This dataset is a unique compilation that not only combines multiple sources but also ensures data quality through normalization and deduplication. A key feature is that all entries in the text column are unique - there are no duplicate… See the full description on the dataset page: https://huggingface.co/datasets/Alaamer/medium-articles-posts-with-content.note-articles
note articles
note.com の公開ページから抽出した本文と IPADIC による MeCab 解析結果です。
各行は text と、構造化されたトークン列 mecab を持ちます。有料記事は公開されている範囲のみです。
mades-articlesde-gender-case-articles
German Gender Case Articles
Dataset Summary
aieng-lab/de-gender-case-articles is a collection of German sentences containing definite singular articles in controlled gender × case settings. Each example provides an original sentence (text) and a version where all occurrences of the targeted article form (e.g., nominative-male der ) are replaced by a placeholder token ([MASK]). The gold label (label) is the grammatically licensed article form in uppercase (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/aieng-lab/de-gender-case-articles.pubmed_articles_edu3_20250227EnWiki-Pretrain-Articlesaya-telugu-news-articles
Summary
aya-telugu-news-articles is an open source dataset of instruct-style records generated by webscraping a Telugu news articles website. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Telugu Version: 1.0
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-news-articles.europe_pmc_articles_part_1
Dataset Card for "europe_pmc_articles_part_1"
More Information needed
cryptonews-articles-with-price-momentum-labels
Dataset Card for Cryptonews articles with price momentum labels
Dataset Summary
The dataset was gathered from two prominent sources in the cryptocurrency industry: Cryptonews.com and Binance.com. The aim of the dataset was to evaluate the impact of news on crypto price movements.
As we know, news events such as regulatory changes, technological advancements, and major partnerships can have a significant impact on the price of cryptocurrencies. By analyzing the data… See the full description on the dataset page: https://huggingface.co/datasets/SahandNZ/cryptonews-articles-with-price-momentum-labels.vital-articles-da-wiki
Vital Articles Danish Wikipedia Dataset
Overview
Total articles: 28,006
Files: 29 Parquet shards
Language: Danish (da)
Contents
Each row is one article with these fields:
en_title: English Wikipedia title
da_title: Danish Wikipedia title
da_url: Danish Wikipedia article URL
markdown: Article content in markdown format
markdown_chars: Character count of markdown
source_lang: Source language code (da)
fetched_at_utc: UTC timestamp when the article was fetched
sofc_materials_articlesThe SOFC-Exp corpus consists of 45 open-access scholarly articles annotated by domain experts.
A corpus and an inter-annotator agreement study demonstrate the complexity of the suggested
named entity recognition and slot filling tasks as well as high annotation quality is presented
in the accompanying paper.medium_articles
Dataset Card for "medium_articles"
More Information needed
europe_pmc_articles_part_2
Dataset Card for "europe_pmc_articles_part_2"
More Information needed
mteb-nl-news-articles-clsThis dataset contains Dutch news articles along with their corresponding categories, sourced from the Nederlandse Oproep Stichting.
Citation Information
If you find our paper, benchmark or models helpful, please consider cite as follows:
@misc{banar2025mtebnle5nlembeddingbenchmark,
title={MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch},
author={Nikolay Banar and Ehsan Lotfi and Jens Van Nooten and Cristina Arhiliuc and Marija Kliocaite and Walter Daelemans}… See the full description on the dataset page: https://huggingface.co/datasets/clips/mteb-nl-news-articles-cls.aim-technical-articles
Analytics India Magazine Technical Articles Dataset 🚀
Dataset Description
This comprehensive dataset contains 25,685 high-quality technical articles from Analytics India Magazine, one of India's leading publications covering artificial intelligence, machine learning, data science, and emerging technologies.
✨ Dataset Highlights
📚 Comprehensive Coverage: Latest AI models, frameworks, and tools
🔬 Technical Depth: Extracted keywords and complexity scoring
🏭… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/aim-technical-articles.summarization-allegro-articles
