datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ticker_analysis_articlesai-tech-articles
AI/Tech Dataset
This dataset is a collection of AI/tech articles scraped from the web: the 2023
corpus collected by the original Haskell scraper plus everything the Rust
crawler has added since (2023 onwards, with full publication dates).
It's hosted on HuggingFace Datasets, so it is easier to load in and work with.
To load the dataset
1. Install HuggingFace Datasets
pip install datasets
2. Load the dataset
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/siavava/ai-tech-articles.medium-articles-posts-with-content
Medium Articles Dataset Generator
This project combines multiple datasets from Kaggle and Hugging Face to create a comprehensive collection of Medium articles. The combined dataset is available on Hugging Face Hub.
Dataset Description
This dataset is a unique compilation that not only combines multiple sources but also ensures data quality through normalization and deduplication. A key feature is that all entries in the text column are unique - there are no duplicate… See the full description on the dataset page: https://huggingface.co/datasets/Alaamer/medium-articles-posts-with-content.note-articles
note articles
note.com の公開ページから抽出した本文と IPADIC による MeCab 解析結果です。
各行は text と、構造化されたトークン列 mecab を持ちます。有料記事は公開されている範囲のみです。
aim-technical-articles
Analytics India Magazine Technical Articles Dataset 🚀
Dataset Description
This comprehensive dataset contains 25,685 high-quality technical articles from Analytics India Magazine, one of India's leading publications covering artificial intelligence, machine learning, data science, and emerging technologies.
✨ Dataset Highlights
📚 Comprehensive Coverage: Latest AI models, frameworks, and tools
🔬 Technical Depth: Extracted keywords and complexity scoring
🏭… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/aim-technical-articles.pmc-articles-dataset-mentions-snippets
PMC Articles Dataset Mentions Snippets
Text snippets from PubMed Central articles paired with structured dataset citations. Designed for training models to extract dataset references from scientific literature.
Description
Task: Extract structured dataset info (identifier, repository, webpage) from article text
Source: PMC open-access articles
Format: Text snippet → JSON output
Examples: Positive (with datasets) and negative (no datasets)
Fields… See the full description on the dataset page: https://huggingface.co/datasets/vida-nyu/pmc-articles-dataset-mentions-snippets.mteb-nl-news-articles-retThis dataset contains Dutch news articles, sourced from the Nederlandse Oproep Stichting.
Citation Information
If you find our paper, benchmark or models helpful, please consider cite as follows:
@misc{banar2025mtebnle5nlembeddingbenchmark,
title={MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch},
author={Nikolay Banar and Ehsan Lotfi and Jens Van Nooten and Cristina Arhiliuc and Marija Kliocaite and Walter Daelemans},
year={2025}… See the full description on the dataset page: https://huggingface.co/datasets/clips/mteb-nl-news-articles-ret.enwiki-articles-markdown-cleaned-2410justicio-BOE-A-1978-31229-constitucion-by-articles-qa-multilingual-e5-large-groq_llama3_70b-sas
Dataset summary
It is an end-to-end evaluation dataset (using SAS metric) for Justicio.
Domain: Legal, Law, Spanish Constitution
Language: Spanish
SAS summary
The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth.
Justicio summary
Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-multilingual-e5-large-groq_llama3_70b-sas.justicio-BOE-A-1978-31229-constitucion-by-articles-qa-qa-groq_llama3_70b_8192-sas
Dataset summary
It is an end-to-end evaluation dataset (using SAS metric) for Justicio.
Domain: Legal, Law, Spanish Constitution
Language: Spanish
SAS summary
The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth.
Justicio summary
Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-qa-groq_llama3_70b_8192-sas.structured-stern-neon-articles
Structured Stern NEON Community Articles
This repository contains approximately 20k user written texts,
articles, and poetry pulled from archives of the Stern NEON website.
Stern NEON was a community platform where users could write and publish their own articles.
Many of the articles are personal stories, poems, or opinion pieces.
The articles are structured in a way that they can be used for further analysis.
Dataset Details
Uses
This dataset can… See the full description on the dataset page: https://huggingface.co/datasets/dotwee/structured-stern-neon-articles.all-the-news-2-pythia-tfidf-topic-stratified-v1-articlessamorzad-gov-pl-articles
Artykuły z platformy samorzad.gov.pl
Wersja: v0.2
Zbiór zawiera 81 418 artykułów opublikowanych na stronach 247 instytucji korzystających ze wspólnej platformy samorzad.gov.pl. Są to między innymi urzędy gmin i powiatów, szkoły, instytucje pomocy społecznej i instytucje kultury.
W wersji v0.2 usunięto dane osobowe i kontaktowe z pól tekstowych przeznaczonych dla odbiorcy. Usunięte wartości zastąpiono jednoznacznymi znacznikami, zachowując układ i znaczenie pozostałej treści.… See the full description on the dataset page: https://huggingface.co/datasets/dawidmajewski/samorzad-gov-pl-articles.v4_nuclear_power_articles
Dataset Card for Nuclear News V4 Dataset
Dataset Summary
The Nuclear News V4 Dataset is a multilingual dataset consisting of 33,104 unique news articles sourced from 12 online news platforms across the Visegrád Group (V4) countries — Poland, Czech Republic, Slovakia, and Hungary — published between 1998 and 2025.
The goal of the dataset is to analyze media narratives surrounding nuclear energy in Central Europe.
While the dataset does not contain human-annotated (golden)… See the full description on the dataset page: https://huggingface.co/datasets/eoplumbum/v4_nuclear_power_articles.ai-jobs-news-articles
Dataset Summary
This dataset brings together 1,000 English-language news articles all about the impact of artificial intelligence on jobs and the workforce. From automation to new tech-driven opportunities, these articles cover a wide range of perspectives and industries. It’s a great resource for anyone interested in how AI is shaping the future of work.
Source Data
The articles were collected from various reputable news outlets, focusing on recent developments and trends at the… See the full description on the dataset page: https://huggingface.co/datasets/fdaudens/ai-jobs-news-articles.justicio-BOE-A-1978-31229-constitucion-by-articles-qa-bge-m3-groq_llama3_70b_8192-sas
Dataset summary
It is an end-to-end evaluation dataset (using SAS metric) for Justicio.
Domain: Legal, Law, Spanish Constitution
Language: Spanish
SAS summary
The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth.
Justicio summary
Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-bge-m3-groq_llama3_70b_8192-sas.enwiki-articles-html-2410justicio-BOE-A-1978-31229-constitucion-by-articles-qa-bge-m3-5k-chunks-groq_llama3_70b_8192-sassubstack-newsletter-articles-metadata-3.5m
Substack Articles & Newsletter Metadata Dataset (3.5M+ Records)
Dataset Summary
This dataset contains 3,513,881 unique records collected from public Substack newsletters and RSS indices. It features cleaned metadata including post titles, subtitles/summaries, author details, standardized publish timestamps, title length metrics, and rule-based thematic categories across 10 topics.
Mendeley DOI: 10.17632/btvzfpd33f.1
Primary Format: Apache Parquet (.parquet)
Total… See the full description on the dataset page: https://huggingface.co/datasets/reja273/substack-newsletter-articles-metadata-3.5m.dzen-russian-articles
Dzen Russian Articles Dataset
Русскоязычные статьи с dzen.ru.
Датасет в активном сборе — новые статьи добавляются регулярно, объём постоянно растёт.
Как устроен парсинг
Статьи собираются с dzen.ru и обрабатываются через Gemini в качестве движка извлечения: модель очищает текст, разбивает на абзацы, определяет категорию, теги, ключевые слова и тональность. Поле extractor содержит название используемой модели (gemini/gemini-3.5-flash-lite).
Структура… See the full description on the dataset page: https://huggingface.co/datasets/ru-dataset/dzen-russian-articles.shipping_news_articles_lsaMULTI_VALUE_qqp_definite_for_indefinite_articles
Dataset Card for "MULTI_VALUE_qqp_definite_for_indefinite_articles"
More Information needed
all-the-news-2-pythia-tfidf-invfreq-topic-stratified-v1-articlese-mordovia-articles-2024
"e-mordovia-articles-2024": a parallel news dataset for Russian, Erzya and Moksha
This is a semi-aligned dataset of Russian, Erzya and Moksha news articles, crawled from https://www.e-mordovia.ru.
Dataset Description
Dataset Summary
This is a dataset of news articles collected from https://www.e-mordovia.ru, the official portal of the state authorities of the Republic of Mordovia.
The articles have been paired by the following algorithm:
Calculate similarities… See the full description on the dataset page: https://huggingface.co/datasets/slone/e-mordovia-articles-2024.Detect-Egyptian-Wikipedia-Articles Detect Egyptian Wikipedia Template-translated Articles
Dataset Description:
We release the heuristically filtered, manually processed, and automatically classified Egyptian Arabic Wikipedia articles dataset. This dataset was used to develop a web-based detection system to automatically identify the template-translated articles on the Egyptian Arabic Wikipedia edition. The system is called Egyptian Arabic Wikipedia Scanner and is hosted on Hugging Face Spaces, here:… See the full description on the dataset page: https://huggingface.co/datasets/SaiedAlshahrani/Detect-Egyptian-Wikipedia-Articles.Content-Articles
Content-Articles Dataset
Overview
The Content-Articles dataset is a collection of academic articles and research papers across various subjects, including Computer Science, Physics, and Mathematics. This dataset is designed to facilitate research and analysis in these fields by providing structured data on article titles, abstracts, and subject classifications.
Dataset Details
Modalities
Tabular: The dataset is structured in a tabular format.
Text:… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Content-Articles.Medium-Articles-Corpus
Medium Articles Corpus (10K Sample)
The Medium Articles Corpus is a massive, clean dataset of articles scraped from Medium.com. This sample version contains 10,000 articles + and is designed to showcase the quality and structure of the full corpus for researchers and developers.
This is the subset from the large dataset https://crawlfeeds.com/websites/medium/text_data/medium_articles
Dataset Features
This dataset includes the following key features, provided in a… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medium-Articles-Corpus.trwiki-articles-turkish
Turkish Wikipedia articles (trwiki_articles) - content namespace
Encyclopedic Turkish articles (Vikipedi, namespace 0, non-redirect) from
the pinned trwiki pages-meta-current dump (2026-10-01), latest revision
per page, wikitext stripped to plain text.
Source pages: 701,597 retained from 3,430,902 dump pages (all namespaces)
Retained records: 326,685 (46.6% of content pages)
Measured cl100k_base proxy tokens: 410,371,495
Characters: 1,043,169,578
Time span (revision dates):… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/trwiki-articles-turkish.cswiki-articles-czech
Czech Wikipedia articles (cswiki_articles) - content namespace
Encyclopedic Czech articles (Wikipedie, namespace 0, non-redirect) from
the pinned cswiki pages-meta-current dump (2026-10-01), latest revision
per page, wikitext stripped to plain text.
Source pages: 599,166 retained from 1,666,311 dump pages (all namespaces)
Retained records: 499,957 (83.4% of content pages)
Measured cl100k_base proxy tokens: 729,882,310
Characters: 1,643,148,657
Time span (revision dates):… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/cswiki-articles-czech.Medical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.
