Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01acmc /beamit-full-texts-dataset Dataset Card for "beamit-full-texts-dataset" More Information needed text10K<n<100K0 likes3.1k downloads3y agoHugging Face02Felldude /Gradients_Gradients_and_Text_Full_Logic_Captionsimage1K<n<10K2 likes1.8k downloads1mo agoHugging Face03amrachraf /arXiv-full-text-chunked Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]:… See the full description on the dataset page: https://huggingface.co/datasets/amrachraf/arXiv-full-text-chunked.texttext-generation100K<n<1M1 likes1.1k downloads2y agoHugging Face04MathematicianNLPer /hamela_books_text_full_oktext1M<n<10M1 likes306 downloads8mo agoHugging Face05acmc /beamit-annotated-full-texts-dataset Dataset Card for "beamit-annotated-full-texts-dataset" More Information needed tabular10K<n<100K0 likes195 downloads3y agoHugging Face06MoMonir /shamela_books_text_full Shamela_Books_Text_Full This dataset contains the full text content of Islamic Arabic books from the Shamela Library, organized by category, book, volume, and page, with footnotes stored separately. It is designed to support Arabic NLP, digital humanities, and bibliographic analysis. 🔗 This dataset is linked to the companion metadata dataset: 👉 Shamela_Books_info via the book_id field. Update : The dataset includes the original raw files as well as a single… See the full description on the dataset page: https://huggingface.co/datasets/MoMonir/shamela_books_text_full.text10M<n<100M2 likes185 downloads4mo agoHugging Face07laion /Pes2oX-fulltextIntroducing Pes2oX Full Text, a transformed dataset derived from the original Allen AI's Pes2o dataset. Our focus in this dataset was to restructure and reorganize the original Pes2o dataset. This was done to make it more accessible to research groups in terms of using it for training Artificial Intelligence models and fine-tuning for specific tasks within a particular domain. Why was restructuring necessary? After examining the original Pes2o dataset's structure, we found it necessary to… See the full description on the dataset page: https://huggingface.co/datasets/laion/Pes2oX-fulltext.text1M<n<10M1 likes179 downloads2y agoHugging Face08ghananlpcommunity /ewe-tts-bible-full-audio-text This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Ewe Tts Bible Full Audio Text audio1K<n<10K0 likes164 downloads4mo agoHugging Face09kastan /stormfront-full-textonly Dataset Card for "stormfront-full-textonly" More Information needed text10M<n<100M0 likes154 downloads4y agoHugging Face10distilabel-internal-testing /alvarobartt-improving-text-embeddings-with-llms-full Dataset Card for alvarobartt-improving-text-embeddings-with-llms-full This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/distilabel-internal-testing/alvarobartt-improving-text-embeddings-with-llms-full/raw/main/pipeline.yaml" or explore the… See the full description on the dataset page: https://huggingface.co/datasets/distilabel-internal-testing/alvarobartt-improving-text-embeddings-with-llms-full.textn<1K0 likes87 downloads2y agoHugging Face11pritamdeka /cord-19-fulltext Dataset Card for [pritamdeka/cord-19-fulltext] Dataset Description Dataset Summary This is a modified cord19 dataset which contains only the fulltext field. This can be used directly for language modelling tasks. Languages English Citation Information @article{Wang2020CORD19TC, title={CORD-19: The Covid-19 Open Research Dataset}, author={Lucy Lu Wang and Kyle Lo and Yoganand Chandrasekhar and Russell Reas and Jiangjiang Yang and Darrin… See the full description on the dataset page: https://huggingface.co/datasets/pritamdeka/cord-19-fulltext.text100K<n<1M2 likes81 downloads5y agoHugging Face12Lots-of-LoRAs /task1292_yelp_review_full_text_categorization Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1292_yelp_review_full_text_categorization Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1292_yelp_review_full_text_categorization.texttext-generation1K<n<10K0 likes76 downloads2y agoHugging Face13mncai /orpo-text-pairs-full ORPO Text Preference Pairs (Full) This dataset contains two versions of preference pairs for training language models using ORPO, DPO, or similar preference-based alignment methods. Dataset Description File Rows Description orpo_pairs.jsonl 8,249 Refined/filtered pairs (recommended) orpo_pairs_all.jsonl 14,214 Full dataset before filtering Format: JSONL Language: English Task: Text-only preference learning (no images) Schema Each row… See the full description on the dataset page: https://huggingface.co/datasets/mncai/orpo-text-pairs-full.text10K<n<100K0 likes71 downloads8mo agoHugging Face14sonnetechnology /license-plate-text-recognition-full Dataset Card for "license-plate-text-recognition-full" Background Information This dataset is generated from keremberke/license-plate-object-detection dataset. What we have done is: Get the Bounding Boxes for each plate in an image, Crop the image to make the plate only visible, Run it through the microsoft/trocr-large-printed model to extract the written information. Structure of the Dataset It has the same structure as the… See the full description on the dataset page: https://huggingface.co/datasets/sonnetechnology/license-plate-text-recognition-full.imageimage-to-text1K<n<10K3 likes64 downloads3y agoHugging Face15PranavHarshan /sharegpt_formatted_cord19_fulltexttext100K<n<1M0 likes62 downloads2y agoHugging Face16OTAR3088 /CeLLaTe-benchMark-set-fulltext Dataset Card for CeLLaTe FullText Benchmark Dataset Dataset summary The CeLLaTe FullText Benchmark Dataset is a curated collection of biomedical fulltext created as a benchmark set for evaluating CeLLaTe named entity recognition (NER) models. It was extracted from Europe PMC article XML sources, with a focus on open-access articles. The dataset is intended to support model development and testing by providing an independent evaluation set for demo production runs… See the full description on the dataset page: https://huggingface.co/datasets/OTAR3088/CeLLaTe-benchMark-set-fulltext.text10K<n<100K0 likes55 downloads2mo agoHugging Face17PranavHarshan /sharegpt_formatted_cord19_fulltext_convotext100K<n<1M0 likes52 downloads2y agoHugging Face18Madras1 /rag-qa-fulltext-ptbr RAG QA Full-Text PT-BR Mistral A large-scale dataset of Brazilian Portuguese RAG-style question-answer pairs with grounded evidence spans, generated from Madras1/corpus-ptbr-v1 documents using Mistral models. Every answer is anchored to literal quotations from the source text, making this dataset suitable for training and evaluating retrieval-augmented generation systems, extractive QA models, and reading comprehension benchmarks in Portuguese. Two configurations are available:… See the full description on the dataset page: https://huggingface.co/datasets/Madras1/rag-qa-fulltext-ptbr.tabularquestion-answering1M<n<10M0 likes52 downloads5mo agoHugging Face19SkyWhal3 /stxbp1-pubmed-central-fulltext source_datasets: - PubMed Central STXBP1 PubMed Central Full-Text Dataset v2 A comprehensive collection of 31,786 full-text scientific articles from PubMed Central related to STXBP1, synaptic function, and neurological research. 🆕 Version 2 Updates (December 2025) Complete re-extraction with improved HTML parsing Full main text with proper section headers Enhanced metadata extraction 99.7% figure-image matching (see companion multimodal dataset)… See the full description on the dataset page: https://huggingface.co/datasets/SkyWhal3/stxbp1-pubmed-central-fulltext.tabulartext-generation10K<n<100K0 likes51 downloads10mo agoHugging Face20SIRIS-Lab /scilake-fulltext-corpus SciLake Fulltext Corpus The SciLake Fulltext Corpus is a collection of scientific papers parsed and segmented by section, primarily designed for research in the development and evaluation of NLP models. This dataset contains 1,000 full-text papers from various scientific domains, including Neuroscience, Cancer, Transport, and Energy, along with an additional 5,000 random papers from general scientific domains. All papers have been curated with licenses that allow for legal usage… See the full description on the dataset page: https://huggingface.co/datasets/SIRIS-Lab/scilake-fulltext-corpus.text1K<n<10K1 likes45 downloads6mo agoHugging Face21crawlfeeds /Curated-Fox-News-Headlines-and-Full-Text Curated Fox News Headlines and Full Text This dataset contains a clean, curated collection of Fox News articles, including both headlines and full article text. It is designed for use in natural language processing (NLP) tasks such as sentiment analysis, summarization, topic classification, and media analysis. 📁 Dataset Format Format: CSV Encoding: UTF-8 Fields: headline: The article title or headline publish_date: Date the article was published (YYYY-MM-DD) content:… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Curated-Fox-News-Headlines-and-Full-Text.imagetext-classification1K<n<10K2 likes43 downloads1y agoHugging Face22metehan777 /cc-aeo-geo-fulltext-CC-MAIN-2026-21tabular100K<n<1M0 likes42 downloads4mo agoHugging Face23metehan777 /cc-turkish-fulltext-CC-MAIN-2026-21 CC-MAIN-2026-21 Turkish URLs 31.4M URLs · 538K domains from Common Crawl columnar index (content_languages contains tur, HTTP 200). Explorer Browse with pagination and domain search: CC Turkish Explorer Space Files File Description turkish_CC-MAIN-2026-21_urls.parquet 31.4M URLs (1 GB) turkish_CC-MAIN-2026-21_domain_leaderboard.parquet 538K domains turkish_CC-MAIN-2026-21_domain_leaderboard_webgraph.parquet + HC/PR/tier… See the full description on the dataset page: https://huggingface.co/datasets/metehan777/cc-turkish-fulltext-CC-MAIN-2026-21.tabular10M<n<100M0 likes42 downloads4mo agoHugging Face24Menlo /Instruction-text-only-fulltext10K<n<100K3 likes38 downloads1y agoHugging Face25EtMmohammedHafsati /darija_speech_to_text_metadata_fullaudio1K<n<10K0 likes37 downloads6mo agoHugging Face26buseskorkmaz /scigen_enriched_with_full_body_texttextn<1K0 likes34 downloads3y agoHugging Face27amrachraf /arXiv-full-text-chunked-testtextn<1K0 likes34 downloads2y agoHugging Face28amrachraf /arXiv-full-text-chunked-qatext100K<n<1M0 likes33 downloads2y agoHugging Face29datapointai /text-2-image-dpo-human-preferences-fullgated Text-2-Image DPO Human Preferences (Full) The complete human preference dataset for text-to-image generation. 416,360 pairwise judgments from ~20,000 annotators comparing AI-generated images across two evaluation dimensions: prompt alignment and overall preference. This is the full, unfiltered version with uniform vote weights. For quality-filtered subsets with calibrated annotator weighting, see: datapointai/text-2-image-dpo-human-preferences (5,000 pairs, trust-weighted)… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-dpo-human-preferences-full.imageimage-classification10K<n<100K1 likes30 downloads7mo agoHugging Face30acmc /beamit-annotated_full_texts_dataset Dataset Card for "beamit-annotated_full_texts_dataset" More Information needed tabular1K<n<10K0 likes29 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.