Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01acmc /beamit-annotated-full-texts-dataset Dataset Card for "beamit-annotated-full-texts-dataset" More Information needed tabular10K<n<100K0 likes195 downloads3y agoHugging Face02Madras1 /rag-qa-fulltext-ptbr RAG QA Full-Text PT-BR Mistral A large-scale dataset of Brazilian Portuguese RAG-style question-answer pairs with grounded evidence spans, generated from Madras1/corpus-ptbr-v1 documents using Mistral models. Every answer is anchored to literal quotations from the source text, making this dataset suitable for training and evaluating retrieval-augmented generation systems, extractive QA models, and reading comprehension benchmarks in Portuguese. Two configurations are available:… See the full description on the dataset page: https://huggingface.co/datasets/Madras1/rag-qa-fulltext-ptbr.tabularquestion-answering1M<n<10M0 likes52 downloads5mo agoHugging Face03SkyWhal3 /stxbp1-pubmed-central-fulltext source_datasets: - PubMed Central STXBP1 PubMed Central Full-Text Dataset v2 A comprehensive collection of 31,786 full-text scientific articles from PubMed Central related to STXBP1, synaptic function, and neurological research. 🆕 Version 2 Updates (December 2025) Complete re-extraction with improved HTML parsing Full main text with proper section headers Enhanced metadata extraction 99.7% figure-image matching (see companion multimodal dataset)… See the full description on the dataset page: https://huggingface.co/datasets/SkyWhal3/stxbp1-pubmed-central-fulltext.tabulartext-generation10K<n<100K0 likes51 downloads10mo agoHugging Face04metehan777 /cc-aeo-geo-fulltext-CC-MAIN-2026-21tabular100K<n<1M0 likes42 downloads4mo agoHugging Face05metehan777 /cc-turkish-fulltext-CC-MAIN-2026-21 CC-MAIN-2026-21 Turkish URLs 31.4M URLs · 538K domains from Common Crawl columnar index (content_languages contains tur, HTTP 200). Explorer Browse with pagination and domain search: CC Turkish Explorer Space Files File Description turkish_CC-MAIN-2026-21_urls.parquet 31.4M URLs (1 GB) turkish_CC-MAIN-2026-21_domain_leaderboard.parquet 538K domains turkish_CC-MAIN-2026-21_domain_leaderboard_webgraph.parquet + HC/PR/tier… See the full description on the dataset page: https://huggingface.co/datasets/metehan777/cc-turkish-fulltext-CC-MAIN-2026-21.tabular10M<n<100M0 likes42 downloads4mo agoHugging Face06acmc /beamit-annotated_full_texts_dataset Dataset Card for "beamit-annotated_full_texts_dataset" More Information needed tabular1K<n<10K0 likes29 downloads3y agoHugging Face07albertge /ni-20-clustered-fulltext-modernbert-sweep-20250107tabular10K<n<100K0 likes29 downloads2y agoHugging Face08taesiri /ArXivSignals-FullText ArXivSignals FullText — arXiv Papers OCR'd to Markdown + Layout A continuously-updated, day-partitioned dataset of arXiv papers converted to clean full text by a vision OCR pipeline: each paper's PDF is rendered to Markdown (headings, paragraphs, tables as HTML, math as LaTeX) plus a structured layout JSON (typed, bounding-boxed blocks). It is the full-text companion to taesiri/ArXivSignals (metadata + LLM signal & summaries) and joins it on paper_id. How it's made… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/ArXivSignals-FullText.tabulartext-generation1K<n<10K0 likes27 downloads3mo agoHugging Face09albertge /dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p0.0-20250109tabular10K<n<100K0 likes20 downloads2y agoHugging Face10Mikimi /ru-wikipedia-100k-full-text-daily-stats-10-years 📚 Russian Wikipedia Top 100K: Full Text with Daily Pageviews **Крупнейший открытый датасет русскоязычной Википедии с полными текстами статей и ежедневной статистикой просмотров за 10 лет. МГУ, ОТиПЛ, 2025** 📖 Описание Этот датасет содержит 99,348 самых популярных статей русскоязычной Википедии, отобранных по совокупному количеству просмотров за последние 10 лет. Для каждой статьи собраны полный текст с сохранением структуры, метаданные и детальная ежедневная… See the full description on the dataset page: https://huggingface.co/datasets/Mikimi/ru-wikipedia-100k-full-text-daily-stats-10-years.tabulartext-generation10K<n<100K1 likes20 downloads10mo agoHugging Face11albertge /dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p0.6-20250109tabular10K<n<100K0 likes18 downloads2y agoHugging Face12albertge /dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p1.0-20250109tabular10K<n<100K0 likes17 downloads2y agoHugging Face13nibzard /narodne-novine-full-text-markdown Narodne Novine Full Text Markdown HTML-to-Markdown text extraction snapshot derived from the NN archive. Coverage Extracted acts: 96855 Failed acts: 157 Missing HTML embodiments: 150 Files texts.parquet failures.parquet metadata.json Notes Extraction prefers /hrv/printhtml, then falls back to /hrv/html. Conversion method: markitdown_html This snapshot does not mirror PDFs. tabulartext-generation10K<n<100K0 likes17 downloads6mo agoHugging Face14albertge /dolly-15k-clustered-fulltext-modernbert-sweep-20250106tabular10K<n<100K0 likes13 downloads2y agoHugging Face15AdrienB134 /MOSEL-fr-tts-text-tags-full-v1tabular1M<n<10M0 likes12 downloads2y agoHugging Face16albertge /dolly-15k-clustered-fulltext-less-sweep-20250106tabular10K<n<100K0 likes12 downloads2y agoHugging Face17albertge /dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p0.4-20250109tabular10K<n<100K0 likes12 downloads2y agoHugging Face18albertge /dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p0.8-20250109tabular10K<n<100K0 likes12 downloads2y agoHugging Face19rchu233 /ni-20-clustered-fulltext-modernbert-sweep-20250107-modernbert-split-kmeans-dim768-20250130tabular10K<n<100K0 likes12 downloads2y agoHugging Face20AdrienB134 /fr-tts-text-tags-full-v1tabular100K<n<1M0 likes9 downloads2y agoHugging Face21albertge /dolly-15k-clustered-less-fulltext-20250106tabular10K<n<100K0 likes9 downloads2y agoHugging Face22SF-Corpus /EF_Full_Texts Dataset Card for SF Nexus Extracted Features: Full Texts Dataset Summary The SF Nexus Extracted Features Full Texts dataset contains text and metadata from 403 mid-twentieth century science fiction books, originally digitized from Temple University Libraries' Paskow Science Fiction Collection. After digitization, the books were cleaned using Abbyy FineReader. Because this is a collection of copyrighted fiction, the books have been disaggregated. Each row of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/SF-Corpus/EF_Full_Texts.tabularn<1K1 likes8 downloads3y agoHugging Face23albertge /dolly-15k-clustered-fulltext-agglomerative-16-20250103tabular10K<n<100K0 likes8 downloads2y agoHugging Face24Mikimi /ru-wikipedia-daily-pageviews-full-texttabular10K<n<100K2 likes7 downloads10mo agoHugging Face25AdrienB134 /Emilia-fr-tts-text-tags-full-v1tabular100K<n<1M0 likes6 downloads2y agoHugging Face26albertge /dolly-15k-clustered-fulltext-full_dim_768-20250103tabular10K<n<100K0 likes6 downloads2y agoHugging Face27amang1802 /synthetic_data_qna_fulltext_conditioned_L3.3_70Btabular10K<n<100K0 likes6 downloads2y agoHugging Face28amang1802 /synthetic_data_fulltext_conditioned_L3.3_70Btabular10K<n<100K0 likes5 downloads2y agoHugging Face29albertge /dolly-15k-clustered-fulltext-20250102tabular10K<n<100K0 likes5 downloads2y agoHugging Face30Mikimi /ru-wikipedia-top-100k-full-texttabular10K<n<100K0 likes5 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.