Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01adameubanks /filtered_articles_by_year Dataset Card for Filtered Articles by Year Dataset Summary The Filtered Articles by Year dataset contains yearly-segmented web articles from the FineWeb dataset, specifically filtered and processed for temporal language analysis and Word2Vec model training. This dataset spans 21 years (2005-2025) and serves as the foundation for research into semantic change, concept emergence, and language evolution over time. Supported Tasks and Leaderboards This dataset… See the full description on the dataset page: https://huggingface.co/datasets/adameubanks/filtered_articles_by_year.texttext-generation10M<n<100M1 likes3.2k downloads1y agoHugging Face02Luishae0705 /wikipedia-articles Wikipedia articles (plain text) 10,418 English Wikipedia articles as plain text, one JSON object per line in wikipedia_articles.jsonl: {"title": "Xcode", "text": "Xcode is a suite of developer tools ..."} Source: the English Wikipedia (https://en.wikipedia.org), downloaded as plain text with the MediaWiki API in September 2026. Images, tables and infoboxes are not included. Section headings appear as == Heading == lines. Selection: the articles were picked from keyword lists… See the full description on the dataset page: https://huggingface.co/datasets/Luishae0705/wikipedia-articles.text10K<n<100K1 likes145 downloads15d agoHugging Face03Threadbaire /research-articles Threadbaire Research Articles Independent research on AI economics and infrastructure: where the value in the AI economy is created, who keeps it, and who ends up paying for it. By Lida Liberopoulou (about · threadbaire.com). Structural analysis of AI industry dynamics, software value collapse, and open infrastructure. Thirteen articles published between January and September 2026, available as raw markdown for analysis, citation, and AI-readable ingestion. About… See the full description on the dataset page: https://huggingface.co/datasets/Threadbaire/research-articles.texttext-classificationn<1K0 likes135 downloads12d agoHugging Face04MongoDB /devcenter-articles Overview This dataset consists of ~600 articles from the MongoDB Developer Center. Dataset Structure The dataset consists of the following fields: sourceName: The source of the article. This value is devcenter for the entire dataset. url: Link to the article action: Action taken on the article. This value is created for the entire dataset. body: Content of the article in Markdown format format: Format of the content. This value is md for all articles. metadata: Metadata… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/devcenter-articles.textquestion-answeringn<1K0 likes115 downloads2y agoHugging Face05KrossKinetic /SP500-Financial-News-Articles-Time-SeriesTextual Time Series Dataset for finetuning / pretraining. Json version of original dataset. Original Dataset : https://www.kaggle.com/datasets/skywalker290/financial-news-article-and-stock-trend-dataset?select=stock_data_articles.csv text1K<n<10K6 likes114 downloads2y agoHugging Face06dotwee /structured-stern-neon-articles Structured Stern NEON Community Articles This repository contains approximately 20k user written texts, articles, and poetry pulled from archives of the Stern NEON website. Stern NEON was a community platform where users could write and publish their own articles. Many of the articles are personal stories, poems, or opinion pieces. The articles are structured in a way that they can be used for further analysis. Dataset Details Uses This dataset can… See the full description on the dataset page: https://huggingface.co/datasets/dotwee/structured-stern-neon-articles.tabulartext-classification10K<n<100K0 likes87 downloads4d agoHugging Face07eoplumbum /v4_nuclear_power_articles Dataset Card for Nuclear News V4 Dataset Dataset Summary The Nuclear News V4 Dataset is a multilingual dataset consisting of 33,104 unique news articles sourced from 12 online news platforms across the Visegrád Group (V4) countries — Poland, Czech Republic, Slovakia, and Hungary — published between 1998 and 2025. The goal of the dataset is to analyze media narratives surrounding nuclear energy in Central Europe. While the dataset does not contain human-annotated (golden)… See the full description on the dataset page: https://huggingface.co/datasets/eoplumbum/v4_nuclear_power_articles.tabulartext-classification10K<n<100K1 likes82 downloads1y agoHugging Face08nisaar /Articles_Constitution_3300_Instruction_SetDataset Card for Indian Constitutional Law Instruction-Response Dataset Dataset Summary The dataset contains instruction-input-output pairs on Indian Constitutional Law, specifically addressing Articles 12, 14, 19, 21, and 15. It's designed to assist AI models, researchers, and learners in understanding and generating responses to complex legal questions related to the Indian Constitution. Supported Tasks This dataset supports tasks such as question answering, text comprehension, language… See the full description on the dataset page: https://huggingface.co/datasets/nisaar/Articles_Constitution_3300_Instruction_Set.text1K<n<10K6 likes71 downloads3y agoHugging Face09ru-dataset /dzen-russian-articles Dzen Russian Articles Dataset Русскоязычные статьи с dzen.ru. Датасет в активном сборе — новые статьи добавляются регулярно, объём постоянно растёт. Как устроен парсинг Статьи собираются с dzen.ru и обрабатываются через Gemini в качестве движка извлечения: модель очищает текст, разбивает на абзацы, определяет категорию, теги, ключевые слова и тональность. Поле extractor содержит название используемой модели (gemini/gemini-3.5-flash-lite). Структура… See the full description on the dataset page: https://huggingface.co/datasets/ru-dataset/dzen-russian-articles.tabulartext-classificationn<1K1 likes66 downloads2mo agoHugging Face10toni5rovic /bcms-fake-news-articlestexttext-classification10K<n<100K0 likes65 downloads1y agoHugging Face11kaengreg /wikifacts-articles_v0text1K<n<10K0 likes49 downloads2y agoHugging Face12kaengreg /wikifacts-articlestext10K<n<100K0 likes49 downloads2y agoHugging Face13psychologie-et-serenite /articles-metadata Psychology Articles Metadata (FR-EN Bilingual) A bilingual (French / English) metadata catalog of clinical psychology articles published on psychologieetserenite.com, authored by Gildas Garrec (CBT psychopractitioner). Each article is paired across the two languages with canonical URLs, themes, keywords, word counts and timestamps. The dataset is designed for: Translation alignment research (FR↔EN parallel article metadata) Multilingual text classification (psychology themes)… See the full description on the dataset page: https://huggingface.co/datasets/psychologie-et-serenite/articles-metadata.tabulartext-classificationn<1K0 likes35 downloads6mo agoHugging Face14anhnon /vietnamese-corporate-legal-articles-fsm Lexora Knowledge - Vietnamese Legal Documents Dataset Summary A structured Vietnamese legal knowledge base crawled from vbpl.vn (CSDL Quốc gia về Pháp luật - Vietnam's National Legal Database), published as 4 linked subsets: full documents, individual articles (Điều), the citation graph between documents/articles, and domain-concept tags. Load a specific subset with datasets.load_dataset("anhnon/vietnamese-corporate-legal-articles-fsm", "articles") etc. Intended… See the full description on the dataset page: https://huggingface.co/datasets/anhnon/vietnamese-corporate-legal-articles-fsm.texttext-generation100K<n<1M1 likes34 downloads2mo agoHugging Face15hybridfree /phoronix-articles Phoronix Articles Dataset: The Archive of Open-Source Computing Journalism The definitive dataset of Phoronix - your gateway to years of open-source hardware/software evolution, performance analysis, and Linux ecosystem journalism. 🚀 What's Inside? This dataset contains the complete archive of Phoronix articles - from bleeding-edge hardware launches to deep-dive Linux kernel analysis. Perfect for researchers, developers, and AI enthusiasts who need high-quality technical… See the full description on the dataset page: https://huggingface.co/datasets/hybridfree/phoronix-articles.tabulartext-generation10K<n<100K0 likes30 downloads8mo agoHugging Face16Mindofmachine /paul_graham_and_sam_altman_articlestextn<1K2 likes29 downloads3y agoHugging Face17Svngoku /french-colonial-articles This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. french_colonial_articles This dataset contains instruction-response pairs focused on African and French colonial history, featuring detailed historical accounts of events such as the Algerian War, the Thiaroye massacre, and post-colonial diplomatic tensions. The content is structured as conversations where an expert historian provides objective, fact-based answers with… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/french-colonial-articles.textn<1K1 likes29 downloads6mo agoHugging Face18clockwork7 /reddit_news_articles_commentstext100K<n<1M2 likes26 downloads1y agoHugging Face19khaihernlow /bitcoin-news-articles-text-corporaimage1K<n<10K0 likes25 downloads2y agoHugging Face20TawasulAI /egyptian-law-articlestext1K<n<10K0 likes22 downloads1y agoHugging Face21welyjesch /hiligaynon_news_articlestext1K<n<10K1 likes22 downloads6mo agoHugging Face22HoundyWoundy /Scientific-dataset-on-articles-small-thinkScientific Dataset on Articles (Small) Это специализированный русскоязычный датасет небольшого объема (~1.2 тыс. строк), содержащий текстовую информацию, извлеченную из научных журналов по биологии (энтомология, арахнология, палеонтология), биографических очерков ученых-исследователей, а также дополненную тематическими материалами из открытых источников интернета. Описание датасета Датасет спроектирован для задач извлечения знаний (Information Extraction), ответов на вопросы по научным текстам… See the full description on the dataset page: https://huggingface.co/datasets/HoundyWoundy/Scientific-dataset-on-articles-small-think.text1K<n<10K0 likes20 downloads2mo agoHugging Face23ayoubelmhamdi /prompts-simplify-articles10+3 prompts to fine-tune Llm to simplify Articles texts. textn<1K0 likes19 downloads3y agoHugging Face24googlair /news_articlestext10K<n<100K0 likes17 downloads3y agoHugging Face25Tobander /combined_articles_hftext10K<n<100K0 likes17 downloads2y agoHugging Face26MongoDB /devcenter-articles-embedded Overview This dataset consists of chunked and embedded versions of a subset of articles from the MongoDB Developer Center. Dataset Structure The dataset consists of the following fields: sourceName: The source of the article. This value is devcenter for the entire dataset. url: Link to the article action: Action taken on the article. This value is created for the entire dataset. body: Content of the chunk in Markdown format format: Format of the content. This value is… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/devcenter-articles-embedded.textquestion-answeringn<1K0 likes17 downloads2y agoHugging Face27ibunescu /gdpr-articles-dataset-traintextn<1K3 likes14 downloads3y agoHugging Face28abdullah1027 /maleeha-lodhi-dawn-articles Maleeha Lodhi Opinion Articles Dataset This dataset contains a curated collection of opinion articles originally published in Dawn, one of Pakistan’s leading English-language newspapers, over the past six years.The articles have been formatted for instruction-based fine-tuning of large language models to emulate the writing style of Maleeha Lodhi — a distinguished Pakistani diplomat, journalist, and academic known for her analytical, sophisticated commentary on international… See the full description on the dataset page: https://huggingface.co/datasets/abdullah1027/maleeha-lodhi-dawn-articles.texttext-generationn<1K0 likes14 downloads1y agoHugging Face29SaifullahBinYusuf /banglatribune_news_articles Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/SaifullahBinYusuf/banglatribune_news_articles.text100K<n<1M0 likes13 downloads2y agoHugging Face30Omarrran /BBC_Eng_News_Articles_datasetgated BBC News Articles Dataset Dataset Description A collection of 2,225 news articles from BBC, suitable for text classification, summarization, and NLP tasks. Dataset Summary Metric Value Total Articles 2,225 Unique Articles 2,092 Columns filename, article_text Language English Source BBC News Dataset Structure Data Fields Field Type Description filename string Unique identifier/filename for each… See the full description on the dataset page: https://huggingface.co/datasets/Omarrran/BBC_Eng_News_Articles_dataset.texttext-classification1K<n<10K0 likes13 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.