Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jtviegas /ticker_analysis_articlestabular10K<n<100K0 likes432 downloads22h agoHugging Face02siavava /ai-tech-articles AI/Tech Dataset This dataset is a collection of AI/tech articles scraped from the web: the 2023 corpus collected by the original Haskell scraper plus everything the Rust crawler has added since (2023 onwards, with full publication dates). It's hosted on HuggingFace Datasets, so it is easier to load in and work with. To load the dataset 1. Install HuggingFace Datasets pip install datasets 2. Load the dataset from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/siavava/ai-tech-articles.tabulartext-generation10K<n<100K8 likes325 downloads8h agoHugging Face03Alaamer /medium-articles-posts-with-content Medium Articles Dataset Generator This project combines multiple datasets from Kaggle and Hugging Face to create a comprehensive collection of Medium articles. The combined dataset is available on Hugging Face Hub. Dataset Description This dataset is a unique compilation that not only combines multiple sources but also ensures data quality through normalization and deduplication. A key feature is that all entries in the text column are unique - there are no duplicate… See the full description on the dataset page: https://huggingface.co/datasets/Alaamer/medium-articles-posts-with-content.tabulartext-classification100K<n<1M3 likes306 downloads2y agoHugging Face04nakasyou /note-articles note articles note.com の公開ページから抽出した本文と IPADIC による MeCab 解析結果です。 各行は text と、構造化されたトークン列 mecab を持ちます。有料記事は公開されている範囲のみです。 tabular100K<n<1M1 likes279 downloads2mo agoHugging Face05abhilash88 /aim-technical-articles Analytics India Magazine Technical Articles Dataset 🚀 Dataset Description This comprehensive dataset contains 25,685 high-quality technical articles from Analytics India Magazine, one of India's leading publications covering artificial intelligence, machine learning, data science, and emerging technologies. ✨ Dataset Highlights 📚 Comprehensive Coverage: Latest AI models, frameworks, and tools 🔬 Technical Depth: Extracted keywords and complexity scoring 🏭… See the full description on the dataset page: https://huggingface.co/datasets/abhilash88/aim-technical-articles.tabulartext-classification10K<n<100K2 likes182 downloads1y agoHugging Face06vida-nyu /pmc-articles-dataset-mentions-snippets PMC Articles Dataset Mentions Snippets Text snippets from PubMed Central articles paired with structured dataset citations. Designed for training models to extract dataset references from scientific literature. Description Task: Extract structured dataset info (identifier, repository, webpage) from article text Source: PMC open-access articles Format: Text snippet → JSON output Examples: Positive (with datasets) and negative (no datasets) Fields… See the full description on the dataset page: https://huggingface.co/datasets/vida-nyu/pmc-articles-dataset-mentions-snippets.tabular1K<n<10K0 likes158 downloads3mo agoHugging Face07clips /mteb-nl-news-articles-retThis dataset contains Dutch news articles, sourced from the Nederlandse Oproep Stichting. Citation Information If you find our paper, benchmark or models helpful, please consider cite as follows: @misc{banar2025mtebnle5nlembeddingbenchmark, title={MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch}, author={Nikolay Banar and Ehsan Lotfi and Jens Van Nooten and Cristina Arhiliuc and Marija Kliocaite and Walter Daelemans}, year={2025}… See the full description on the dataset page: https://huggingface.co/datasets/clips/mteb-nl-news-articles-ret.tabular100K<n<1M0 likes148 downloads1y agoHugging Face08tom-010 /enwiki-articles-markdown-cleaned-2410tabular1M<n<10M0 likes145 downloads2y agoHugging Face09dariolopez /justicio-BOE-A-1978-31229-constitucion-by-articles-qa-multilingual-e5-large-groq_llama3_70b-sas Dataset summary It is an end-to-end evaluation dataset (using SAS metric) for Justicio. Domain: Legal, Law, Spanish Constitution Language: Spanish SAS summary The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth. Justicio summary Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-multilingual-e5-large-groq_llama3_70b-sas.tabularquestion-answeringn<1K0 likes105 downloads2y agoHugging Face10dariolopez /justicio-BOE-A-1978-31229-constitucion-by-articles-qa-qa-groq_llama3_70b_8192-sas Dataset summary It is an end-to-end evaluation dataset (using SAS metric) for Justicio. Domain: Legal, Law, Spanish Constitution Language: Spanish SAS summary The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth. Justicio summary Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-qa-groq_llama3_70b_8192-sas.tabularquestion-answeringn<1K1 likes97 downloads2y agoHugging Face11dotwee /structured-stern-neon-articles Structured Stern NEON Community Articles This repository contains approximately 20k user written texts, articles, and poetry pulled from archives of the Stern NEON website. Stern NEON was a community platform where users could write and publish their own articles. Many of the articles are personal stories, poems, or opinion pieces. The articles are structured in a way that they can be used for further analysis. Dataset Details Uses This dataset can… See the full description on the dataset page: https://huggingface.co/datasets/dotwee/structured-stern-neon-articles.tabulartext-classification10K<n<100K0 likes87 downloads4d agoHugging Face12MA-tokenweights /all-the-news-2-pythia-tfidf-topic-stratified-v1-articlestabular100K<n<1M0 likes86 downloads2mo agoHugging Face13dawidmajewski /samorzad-gov-pl-articles Artykuły z platformy samorzad.gov.pl Wersja: v0.2 Zbiór zawiera 81 418 artykułów opublikowanych na stronach 247 instytucji korzystających ze wspólnej platformy samorzad.gov.pl. Są to między innymi urzędy gmin i powiatów, szkoły, instytucje pomocy społecznej i instytucje kultury. W wersji v0.2 usunięto dane osobowe i kontaktowe z pól tekstowych przeznaczonych dla odbiorcy. Usunięte wartości zastąpiono jednoznacznymi znacznikami, zachowując układ i znaczenie pozostałej treści.… See the full description on the dataset page: https://huggingface.co/datasets/dawidmajewski/samorzad-gov-pl-articles.tabulartext-generation10K<n<100K0 likes83 downloads2mo agoHugging Face14eoplumbum /v4_nuclear_power_articles Dataset Card for Nuclear News V4 Dataset Dataset Summary The Nuclear News V4 Dataset is a multilingual dataset consisting of 33,104 unique news articles sourced from 12 online news platforms across the Visegrád Group (V4) countries — Poland, Czech Republic, Slovakia, and Hungary — published between 1998 and 2025. The goal of the dataset is to analyze media narratives surrounding nuclear energy in Central Europe. While the dataset does not contain human-annotated (golden)… See the full description on the dataset page: https://huggingface.co/datasets/eoplumbum/v4_nuclear_power_articles.tabulartext-classification10K<n<100K1 likes82 downloads1y agoHugging Face15fdaudens /ai-jobs-news-articles Dataset Summary This dataset brings together 1,000 English-language news articles all about the impact of artificial intelligence on jobs and the workforce. From automation to new tech-driven opportunities, these articles cover a wide range of perspectives and industries. It’s a great resource for anyone interested in how AI is shaping the future of work. Source Data The articles were collected from various reputable news outlets, focusing on recent developments and trends at the… See the full description on the dataset page: https://huggingface.co/datasets/fdaudens/ai-jobs-news-articles.tabulartext-classification1K<n<10K1 likes82 downloads1y agoHugging Face16dariolopez /justicio-BOE-A-1978-31229-constitucion-by-articles-qa-bge-m3-groq_llama3_70b_8192-sas Dataset summary It is an end-to-end evaluation dataset (using SAS metric) for Justicio. Domain: Legal, Law, Spanish Constitution Language: Spanish SAS summary The concept of Semantic Answer Similarity refers to the evaluation of the semantic similarity between the generated answer and the ground truth. The score ranges from 0 to 1. A higher score indicates a better match between the generated answer and the ground truth. Justicio summary Justicio is a… See the full description on the dataset page: https://huggingface.co/datasets/dariolopez/justicio-BOE-A-1978-31229-constitucion-by-articles-qa-bge-m3-groq_llama3_70b_8192-sas.tabularquestion-answeringn<1K0 likes78 downloads2y agoHugging Face17tom-010 /enwiki-articles-html-2410tabular1M<n<10M0 likes71 downloads2y agoHugging Face18dariolopez /justicio-BOE-A-1978-31229-constitucion-by-articles-qa-bge-m3-5k-chunks-groq_llama3_70b_8192-sastabularn<1K0 likes67 downloads2y agoHugging Face19reja273 /substack-newsletter-articles-metadata-3.5m Substack Articles & Newsletter Metadata Dataset (3.5M+ Records) Dataset Summary This dataset contains 3,513,881 unique records collected from public Substack newsletters and RSS indices. It features cleaned metadata including post titles, subtitles/summaries, author details, standardized publish timestamps, title length metrics, and rule-based thematic categories across 10 topics. Mendeley DOI: 10.17632/btvzfpd33f.1 Primary Format: Apache Parquet (.parquet) Total… See the full description on the dataset page: https://huggingface.co/datasets/reja273/substack-newsletter-articles-metadata-3.5m.tabulartext-classification1M<n<10M0 likes67 downloads10d agoHugging Face20ru-dataset /dzen-russian-articles Dzen Russian Articles Dataset Русскоязычные статьи с dzen.ru. Датасет в активном сборе — новые статьи добавляются регулярно, объём постоянно растёт. Как устроен парсинг Статьи собираются с dzen.ru и обрабатываются через Gemini в качестве движка извлечения: модель очищает текст, разбивает на абзацы, определяет категорию, теги, ключевые слова и тональность. Поле extractor содержит название используемой модели (gemini/gemini-3.5-flash-lite). Структура… See the full description on the dataset page: https://huggingface.co/datasets/ru-dataset/dzen-russian-articles.tabulartext-classificationn<1K1 likes66 downloads2mo agoHugging Face21Ktzoras /shipping_news_articles_lsatabular10K<n<100K0 likes64 downloads1y agoHugging Face22liuyanchen1015 /MULTI_VALUE_qqp_definite_for_indefinite_articles Dataset Card for "MULTI_VALUE_qqp_definite_for_indefinite_articles" More Information needed tabular100K<n<1M0 likes62 downloads4y agoHugging Face23MA-tokenweights /all-the-news-2-pythia-tfidf-invfreq-topic-stratified-v1-articlestabular100K<n<1M0 likes61 downloads2mo agoHugging Face24slone /e-mordovia-articles-2024 "e-mordovia-articles-2024": a parallel news dataset for Russian, Erzya and Moksha This is a semi-aligned dataset of Russian, Erzya and Moksha news articles, crawled from https://www.e-mordovia.ru. Dataset Description Dataset Summary This is a dataset of news articles collected from https://www.e-mordovia.ru, the official portal of the state authorities of the Republic of Mordovia. The articles have been paired by the following algorithm: Calculate similarities… See the full description on the dataset page: https://huggingface.co/datasets/slone/e-mordovia-articles-2024.tabulartranslation100K<n<1M3 likes55 downloads2y agoHugging Face25SaiedAlshahrani /Detect-Egyptian-Wikipedia-Articles Detect Egyptian Wikipedia Template-translated Articles Dataset Description: We release the heuristically filtered, manually processed, and automatically classified Egyptian Arabic Wikipedia articles dataset. This dataset was used to develop a web-based detection system to automatically identify the template-translated articles on the Egyptian Arabic Wikipedia edition. The system is called Egyptian Arabic Wikipedia Scanner and is hosted on Hugging Face Spaces, here:… See the full description on the dataset page: https://huggingface.co/datasets/SaiedAlshahrani/Detect-Egyptian-Wikipedia-Articles.tabulartext-classification100K<n<1M1 likes42 downloads2y agoHugging Face26prithivMLmods /Content-Articles Content-Articles Dataset Overview The Content-Articles dataset is a collection of academic articles and research papers across various subjects, including Computer Science, Physics, and Mathematics. This dataset is designed to facilitate research and analysis in these fields by providing structured data on article titles, abstracts, and subject classifications. Dataset Details Modalities Tabular: The dataset is structured in a tabular format. Text:… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Content-Articles.tabulartext-generation10K<n<100K3 likes41 downloads2y agoHugging Face27crawlfeeds /Medium-Articles-Corpus Medium Articles Corpus (10K Sample) The Medium Articles Corpus is a massive, clean dataset of articles scraped from Medium.com. This sample version contains 10,000 articles + and is designed to showcase the quality and structure of the full corpus for researchers and developers. This is the subset from the large dataset https://crawlfeeds.com/websites/medium/text_data/medium_articles Dataset Features This dataset includes the following key features, provided in a… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medium-Articles-Corpus.imagetext-classification10K<n<100K2 likes39 downloads1y agoHugging Face28PiotrSty /trwiki-articles-turkish Turkish Wikipedia articles (trwiki_articles) - content namespace Encyclopedic Turkish articles (Vikipedi, namespace 0, non-redirect) from the pinned trwiki pages-meta-current dump (2026-10-01), latest revision per page, wikitext stripped to plain text. Source pages: 701,597 retained from 3,430,902 dump pages (all namespaces) Retained records: 326,685 (46.6% of content pages) Measured cl100k_base proxy tokens: 410,371,495 Characters: 1,043,169,578 Time span (revision dates):… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/trwiki-articles-turkish.tabulartext-generation100K<n<1M0 likes38 downloads3d agoHugging Face29PiotrSty /cswiki-articles-czech Czech Wikipedia articles (cswiki_articles) - content namespace Encyclopedic Czech articles (Wikipedie, namespace 0, non-redirect) from the pinned cswiki pages-meta-current dump (2026-10-01), latest revision per page, wikitext stripped to plain text. Source pages: 599,166 retained from 1,666,311 dump pages (all namespaces) Retained records: 499,957 (83.4% of content pages) Measured cl100k_base proxy tokens: 729,882,310 Characters: 1,643,148,657 Time span (revision dates):… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/cswiki-articles-czech.tabulartext-generation100K<n<1M0 likes38 downloads3d agoHugging Face30crawlfeeds /Medical-Health-QA-Articles-Dataset Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development. Dataset Overview Field Details Sources iCliniq, HealthTap, WebMD Total Records 1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.imagetext-classification1K<n<10K0 likes37 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.