datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
android-times-articles
Android Times — Articles Dataset Archive
Private dataset containing synthesized and processed article archives, multi-language transcripts, metadata, and editorial assets for Android Times.
Dataset Structure
articles/
├── en-US/ # English (United States) localized articles & scripts
├── ja-JP/ # Japanese localized articles & scripts
├── en-AU/ # Australian localized articles
├── en-CN/ # China localized English… See the full description on the dataset page: https://huggingface.co/datasets/aoiandroid/android-times-articles.filtered_articles_by_year
Dataset Card for Filtered Articles by Year
Dataset Summary
The Filtered Articles by Year dataset contains yearly-segmented web articles from the FineWeb dataset, specifically filtered and processed for temporal language analysis and Word2Vec model training. This dataset spans 21 years (2005-2025) and serves as the foundation for research into semantic change, concept emergence, and language evolution over time.
Supported Tasks and Leaderboards
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/adameubanks/filtered_articles_by_year.political_bias_in_news_articleskurdish-wikipedia-articles
Summary
Extracted from the wikidump. There are summaries and categories available for each article. Will look into adding them later.
Usage
from datasets import load_dataset
ds = load_dataset("nazimali/kurdish-wikipedia-articles", split="train")
ds
Dataset({
features: ['id', 'url', 'title', 'text'],
num_rows: 63076
})
ticker_analysis_articlesfinancial-news-articles
Dataset Card for "financial-news-articles"
More Information needed
The data was obtained from here
german-wikipedia-articlesnews-articles-ptbr-dataset
Dataset Card for "news-articles-ptbr-dataset"
More Information needed
pubmed_articles_domain10x_20250227ai-tech-articles
AI/Tech Dataset
This dataset is a collection of AI/tech articles scraped from the web: the 2023
corpus collected by the original Haskell scraper plus everything the Rust
crawler has added since (2023 onwards, with full publication dates).
It's hosted on HuggingFace Datasets, so it is easier to load in and work with.
To load the dataset
1. Install HuggingFace Datasets
pip install datasets
2. Load the dataset
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/siavava/ai-tech-articles.medium-articles-posts-with-content
Medium Articles Dataset Generator
This project combines multiple datasets from Kaggle and Hugging Face to create a comprehensive collection of Medium articles. The combined dataset is available on Hugging Face Hub.
Dataset Description
This dataset is a unique compilation that not only combines multiple sources but also ensures data quality through normalization and deduplication. A key feature is that all entries in the text column are unique - there are no duplicate… See the full description on the dataset page: https://huggingface.co/datasets/Alaamer/medium-articles-posts-with-content.note-articles
note articles
note.com の公開ページから抽出した本文と IPADIC による MeCab 解析結果です。
各行は text と、構造化されたトークン列 mecab を持ちます。有料記事は公開されている範囲のみです。
de-gender-case-articles
German Gender Case Articles
Dataset Summary
aieng-lab/de-gender-case-articles is a collection of German sentences containing definite singular articles in controlled gender × case settings. Each example provides an original sentence (text) and a version where all occurrences of the targeted article form (e.g., nominative-male der ) are replaced by a placeholder token ([MASK]). The gold label (label) is the grammatically licensed article form in uppercase (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/aieng-lab/de-gender-case-articles.pubmed_articles_edu3_20250227EnWiki-Pretrain-Articlesaya-telugu-news-articles
Summary
aya-telugu-news-articles is an open source dataset of instruct-style records generated by webscraping a Telugu news articles website. This was created as part of Aya Open Science Initiative from Cohere For AI.
This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License.
Supported Tasks:
Training LLMs
Synthetic Data Generation
Data Augmentation
Languages: Telugu Version: 1.0
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-news-articles.europe_pmc_articles_part_1
Dataset Card for "europe_pmc_articles_part_1"
More Information needed
cryptonews-articles-with-price-momentum-labels
Dataset Card for Cryptonews articles with price momentum labels
Dataset Summary
The dataset was gathered from two prominent sources in the cryptocurrency industry: Cryptonews.com and Binance.com. The aim of the dataset was to evaluate the impact of news on crypto price movements.
As we know, news events such as regulatory changes, technological advancements, and major partnerships can have a significant impact on the price of cryptocurrencies. By analyzing the data… See the full description on the dataset page: https://huggingface.co/datasets/SahandNZ/cryptonews-articles-with-price-momentum-labels.vital-articles-da-wiki
Vital Articles Danish Wikipedia Dataset
Overview
Total articles: 28,006
Files: 29 Parquet shards
Language: Danish (da)
Contents
Each row is one article with these fields:
en_title: English Wikipedia title
da_title: Danish Wikipedia title
da_url: Danish Wikipedia article URL
markdown: Article content in markdown format
markdown_chars: Character count of markdown
source_lang: Source language code (da)
fetched_at_utc: UTC timestamp when the article was fetched
medium_articles
Dataset Card for "medium_articles"
More Information needed
europe_pmc_articles_part_2
Dataset Card for "europe_pmc_articles_part_2"
More Information needed
mteb-nl-news-articles-clsThis dataset contains Dutch news articles along with their corresponding categories, sourced from the Nederlandse Oproep Stichting.
Citation Information
If you find our paper, benchmark or models helpful, please consider cite as follows:
@misc{banar2025mtebnle5nlembeddingbenchmark,
title={MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch},
author={Nikolay Banar and Ehsan Lotfi and Jens Van Nooten and Cristina Arhiliuc and Marija Kliocaite and Walter Daelemans}… See the full description on the dataset page: https://huggingface.co/datasets/clips/mteb-nl-news-articles-cls.aim-technical-articles
Analytics India Magazine Technical Articles Dataset 🚀
Dataset Description
This comprehensive dataset contains 25,685 high-quality technical articles from Analytics India Magazine, one of India's leading publications covering artificial intelligence, machine learning, data science, and emerging technologies.
✨ Dataset Highlights
📚 Comprehensive Coverage: Latest AI models, frameworks, and tools
🔬 Technical Depth: Extracted keywords and complexity scoring
🏭… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/aim-technical-articles.summarization-allegro-articlespmc-articles-dataset-mentions-snippets
PMC Articles Dataset Mentions Snippets
Text snippets from PubMed Central articles paired with structured dataset citations. Designed for training models to extract dataset references from scientific literature.
Description
Task: Extract structured dataset info (identifier, repository, webpage) from article text
Source: PMC open-access articles
Format: Text snippet → JSON output
Examples: Positive (with datasets) and negative (no datasets)
Fields… See the full description on the dataset page: https://huggingface.co/datasets/vida-nyu/pmc-articles-dataset-mentions-snippets.mteb-nl-news-articles-retThis dataset contains Dutch news articles, sourced from the Nederlandse Oproep Stichting.
Citation Information
If you find our paper, benchmark or models helpful, please consider cite as follows:
@misc{banar2025mtebnle5nlembeddingbenchmark,
title={MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch},
author={Nikolay Banar and Ehsan Lotfi and Jens Van Nooten and Cristina Arhiliuc and Marija Kliocaite and Walter Daelemans},
year={2025}… See the full description on the dataset page: https://huggingface.co/datasets/clips/mteb-nl-news-articles-ret.enwiki-articles-markdown-cleaned-2410wikipedia-articles
Wikipedia articles (plain text)
10,418 English Wikipedia articles as plain text, one JSON object per line in wikipedia_articles.jsonl:
{"title": "Xcode", "text": "Xcode is a suite of developer tools ..."}
Source: the English Wikipedia (https://en.wikipedia.org), downloaded as plain text with the MediaWiki API in September 2026.
Images, tables and infoboxes are not included. Section headings appear as == Heading == lines.
Selection: the articles were picked from keyword lists… See the full description on the dataset page: https://huggingface.co/datasets/Luishae0705/wikipedia-articles.research-articles
Threadbaire Research Articles
Independent research on AI economics and infrastructure: where the value in the AI economy is created, who keeps it, and who ends up paying for it. By Lida Liberopoulou (about · threadbaire.com).
Structural analysis of AI industry dynamics, software value collapse, and open infrastructure. Thirteen articles published between January and September 2026, available as raw markdown for analysis, citation, and AI-readable ingestion.
About… See the full description on the dataset page: https://huggingface.co/datasets/Threadbaire/research-articles.pubmed_articles
