Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01aoiandroid /android-times-articles Android Times — Articles Dataset Archive Private dataset containing synthesized and processed article archives, multi-language transcripts, metadata, and editorial assets for Android Times. Dataset Structure articles/ ├── en-US/ # English (United States) localized articles & scripts ├── ja-JP/ # Japanese localized articles & scripts ├── en-AU/ # Australian localized articles ├── en-CN/ # China localized English… See the full description on the dataset page: https://huggingface.co/datasets/aoiandroid/android-times-articles.imagetext-generation1K<n<10K0 likes9.1k downloads11d agoHugging Face02adameubanks /filtered_articles_by_year Dataset Card for Filtered Articles by Year Dataset Summary The Filtered Articles by Year dataset contains yearly-segmented web articles from the FineWeb dataset, specifically filtered and processed for temporal language analysis and Word2Vec model training. This dataset spans 21 years (2005-2025) and serves as the foundation for research into semantic change, concept emergence, and language evolution over time. Supported Tasks and Leaderboards This dataset… See the full description on the dataset page: https://huggingface.co/datasets/adameubanks/filtered_articles_by_year.texttext-generation10M<n<100M1 likes3.2k downloads1y agoHugging Face03Aditya11096 /political_bias_in_news_articlestext10K<n<100K1 likes1.5k downloads1y agoHugging Face04nazimali /kurdish-wikipedia-articles Summary Extracted from the wikidump. There are summaries and categories available for each article. Will look into adding them later. Usage from datasets import load_dataset ds = load_dataset("nazimali/kurdish-wikipedia-articles", split="train") ds Dataset({ features: ['id', 'url', 'title', 'text'], num_rows: 63076 }) texttext-classification10K<n<100K0 likes438 downloads2y agoHugging Face05jtviegas /ticker_analysis_articlestabular10K<n<100K0 likes432 downloads21h agoHugging Face06ashraq /financial-news-articles Dataset Card for "financial-news-articles" More Information needed The data was obtained from here text100K<n<1M21 likes399 downloads4y agoHugging Face07jonas-is-coding /german-wikipedia-articlestext1M<n<10M2 likes385 downloads2y agoHugging Face08iara-project /news-articles-ptbr-dataset Dataset Card for "news-articles-ptbr-dataset" More Information needed text100K<n<1M4 likes358 downloads3y agoHugging Face09rntc /pubmed_articles_domain10x_20250227text1M<n<10M0 likes351 downloads2y agoHugging Face10siavava /ai-tech-articles AI/Tech Dataset This dataset is a collection of AI/tech articles scraped from the web: the 2023 corpus collected by the original Haskell scraper plus everything the Rust crawler has added since (2023 onwards, with full publication dates). It's hosted on HuggingFace Datasets, so it is easier to load in and work with. To load the dataset 1. Install HuggingFace Datasets pip install datasets 2. Load the dataset from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/siavava/ai-tech-articles.tabulartext-generation10K<n<100K8 likes325 downloads7h agoHugging Face11Alaamer /medium-articles-posts-with-content Medium Articles Dataset Generator This project combines multiple datasets from Kaggle and Hugging Face to create a comprehensive collection of Medium articles. The combined dataset is available on Hugging Face Hub. Dataset Description This dataset is a unique compilation that not only combines multiple sources but also ensures data quality through normalization and deduplication. A key feature is that all entries in the text column are unique - there are no duplicate… See the full description on the dataset page: https://huggingface.co/datasets/Alaamer/medium-articles-posts-with-content.tabulartext-classification100K<n<1M3 likes306 downloads2y agoHugging Face12nakasyou /note-articles note articles note.com の公開ページから抽出した本文と IPADIC による MeCab 解析結果です。 各行は text と、構造化されたトークン列 mecab を持ちます。有料記事は公開されている範囲のみです。 tabular100K<n<1M1 likes279 downloads2mo agoHugging Face13aieng-lab /de-gender-case-articles German Gender Case Articles Dataset Summary aieng-lab/de-gender-case-articles is a collection of German sentences containing definite singular articles in controlled gender × case settings. Each example provides an original sentence (text) and a version where all occurrences of the targeted article form (e.g., nominative-male der ) are replaced by a placeholder token ([MASK]). The gold label (label) is the grammatically licensed article form in uppercase (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/aieng-lab/de-gender-case-articles.textfill-mask100K<n<1M1 likes253 downloads3mo agoHugging Face14rntc /pubmed_articles_edu3_20250227text1M<n<10M0 likes246 downloads2y agoHugging Face15MemoryAsModality /EnWiki-Pretrain-Articlestext10M<n<100M0 likes233 downloads6mo agoHugging Face16SuryaKrishna02 /aya-telugu-news-articles Summary aya-telugu-news-articles is an open source dataset of instruct-style records generated by webscraping a Telugu news articles website. This was created as part of Aya Open Science Initiative from Cohere For AI. This dataset can be used for any purpose, whether academic or commercial, under the terms of the Apache 2.0 License. Supported Tasks: Training LLMs Synthetic Data Generation Data Augmentation Languages: Telugu Version: 1.0 Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/SuryaKrishna02/aya-telugu-news-articles.texttext-generation100K<n<1M6 likes224 downloads3y agoHugging Face17Hack90 /europe_pmc_articles_part_1 Dataset Card for "europe_pmc_articles_part_1" More Information needed text100K<n<1M2 likes223 downloads3y agoHugging Face18SahandNZ /cryptonews-articles-with-price-momentum-labels Dataset Card for Cryptonews articles with price momentum labels Dataset Summary The dataset was gathered from two prominent sources in the cryptocurrency industry: Cryptonews.com and Binance.com. The aim of the dataset was to evaluate the impact of news on crypto price movements. As we know, news events such as regulatory changes, technological advancements, and major partnerships can have a significant impact on the price of cryptocurrencies. By analyzing the data… See the full description on the dataset page: https://huggingface.co/datasets/SahandNZ/cryptonews-articles-with-price-momentum-labels.texttext-classification100K<n<1M25 likes221 downloads3y agoHugging Face19V4ldeLund /vital-articles-da-wiki Vital Articles Danish Wikipedia Dataset Overview Total articles: 28,006 Files: 29 Parquet shards Language: Danish (da) Contents Each row is one article with these fields: en_title: English Wikipedia title da_title: Danish Wikipedia title da_url: Danish Wikipedia article URL markdown: Article content in markdown format markdown_chars: Character count of markdown source_lang: Source language code (da) fetched_at_utc: UTC timestamp when the article was fetched texttext-generation10K<n<100K0 likes199 downloads8mo agoHugging Face20Kamaljp /medium_articles Dataset Card for "medium_articles" More Information needed text100K<n<1M6 likes187 downloads3y agoHugging Face21Hack90 /europe_pmc_articles_part_2 Dataset Card for "europe_pmc_articles_part_2" More Information needed text1M<n<10M1 likes186 downloads3y agoHugging Face22clips /mteb-nl-news-articles-clsThis dataset contains Dutch news articles along with their corresponding categories, sourced from the Nederlandse Oproep Stichting. Citation Information If you find our paper, benchmark or models helpful, please consider cite as follows: @misc{banar2025mtebnle5nlembeddingbenchmark, title={MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch}, author={Nikolay Banar and Ehsan Lotfi and Jens Van Nooten and Cristina Arhiliuc and Marija Kliocaite and Walter Daelemans}… See the full description on the dataset page: https://huggingface.co/datasets/clips/mteb-nl-news-articles-cls.text1K<n<10K0 likes183 downloads1y agoHugging Face23abhilash88 /aim-technical-articles Analytics India Magazine Technical Articles Dataset 🚀 Dataset Description This comprehensive dataset contains 25,685 high-quality technical articles from Analytics India Magazine, one of India's leading publications covering artificial intelligence, machine learning, data science, and emerging technologies. ✨ Dataset Highlights 📚 Comprehensive Coverage: Latest AI models, frameworks, and tools 🔬 Technical Depth: Extracted keywords and complexity scoring 🏭… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/aim-technical-articles.tabulartext-classification10K<n<100K2 likes182 downloads1y agoHugging Face24allegro /summarization-allegro-articlestext100K<n<1M6 likes171 downloads5y agoHugging Face25vida-nyu /pmc-articles-dataset-mentions-snippets PMC Articles Dataset Mentions Snippets Text snippets from PubMed Central articles paired with structured dataset citations. Designed for training models to extract dataset references from scientific literature. Description Task: Extract structured dataset info (identifier, repository, webpage) from article text Source: PMC open-access articles Format: Text snippet → JSON output Examples: Positive (with datasets) and negative (no datasets) Fields… See the full description on the dataset page: https://huggingface.co/datasets/vida-nyu/pmc-articles-dataset-mentions-snippets.tabular1K<n<10K0 likes158 downloads3mo agoHugging Face26clips /mteb-nl-news-articles-retThis dataset contains Dutch news articles, sourced from the Nederlandse Oproep Stichting. Citation Information If you find our paper, benchmark or models helpful, please consider cite as follows: @misc{banar2025mtebnle5nlembeddingbenchmark, title={MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch}, author={Nikolay Banar and Ehsan Lotfi and Jens Van Nooten and Cristina Arhiliuc and Marija Kliocaite and Walter Daelemans}, year={2025}… See the full description on the dataset page: https://huggingface.co/datasets/clips/mteb-nl-news-articles-ret.tabular100K<n<1M0 likes148 downloads1y agoHugging Face27tom-010 /enwiki-articles-markdown-cleaned-2410tabular1M<n<10M0 likes145 downloads2y agoHugging Face28Luishae0705 /wikipedia-articles Wikipedia articles (plain text) 10,418 English Wikipedia articles as plain text, one JSON object per line in wikipedia_articles.jsonl: {"title": "Xcode", "text": "Xcode is a suite of developer tools ..."} Source: the English Wikipedia (https://en.wikipedia.org), downloaded as plain text with the MediaWiki API in September 2026. Images, tables and infoboxes are not included. Section headings appear as == Heading == lines. Selection: the articles were picked from keyword lists… See the full description on the dataset page: https://huggingface.co/datasets/Luishae0705/wikipedia-articles.text10K<n<100K1 likes145 downloads15d agoHugging Face29Threadbaire /research-articles Threadbaire Research Articles Independent research on AI economics and infrastructure: where the value in the AI economy is created, who keeps it, and who ends up paying for it. By Lida Liberopoulou (about · threadbaire.com). Structural analysis of AI industry dynamics, software value collapse, and open infrastructure. Thirteen articles published between January and September 2026, available as raw markdown for analysis, citation, and AI-readable ingestion. About… See the full description on the dataset page: https://huggingface.co/datasets/Threadbaire/research-articles.texttext-classificationn<1K0 likes135 downloads12d agoHugging Face30JJinho /pubmed_articlestext10M<n<100M5 likes131 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.