Team Ai
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mlfoundations /MINT-1T-PDF-CC-2023-14 πŸƒ MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens πŸƒ MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. πŸƒ MINT-1T is designed to facilitate research in multimodal pretraining. πŸƒ MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-14.imageimage-to-text1M<n<10M6 likes17k downloads2y agoHugging Face02mlfoundations /MINT-1T-PDF-CC-2024-10 πŸƒ MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens πŸƒ MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. πŸƒ MINT-1T is designed to facilitate research in multimodal pretraining. πŸƒ MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.imageimage-to-text1M<n<10M5 likes8.7k downloads2y agoHugging Face03mlfoundations /MINT-1T-PDF-CC-2023-23 πŸƒ MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens πŸƒ MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. πŸƒ MINT-1T is designed to facilitate research in multimodal pretraining. πŸƒ MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-23.imageimage-to-text1M<n<10M10 likes6.7k downloads2y agoHugging Face04pixparse /pdfa-eng-wds Dataset Card for PDF Association dataset (PDFA) Dataset Summary PDFA dataset is a document dataset filtered from the SafeDocs corpus, aka CC-MAIN-2021-31-PDF-UNTRUNCATED. The original purpose of that corpus is for comprehensive pdf documents analysis. The purpose of that subset differs in that regard, as focus has been done on making the dataset machine learning-ready for vision-language models. An example page of one pdf document, with added bounding boxes… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/pdfa-eng-wds.textimage-to-text1K<n<10K161 likes5k downloads3y agoHugging Face05mlfoundations /MINT-1T-PDF-CC-2023-50 πŸƒ MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens πŸƒ MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. πŸƒ MINT-1T is designed to facilitate research in multimodal pretraining. πŸƒ MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2023-50.imageimage-to-text1M<n<10M14 likes2.5k downloads2y agoHugging Face06laion /CS-Arxiv-PDFs-08-25text10M<n<100M2 likes2.3k downloads1y agoHugging Face07LLMDH /marianne_pdf_7text10K<n<100K0 likes1.3k downloads2y agoHugging Face08LLMDH /marianne_pdf_5text100K<n<1M0 likes1.1k downloads2y agoHugging Face09PleIAs /WTO-PDFtext100K<n<1M8 likes1k downloads2y agoHugging Face10LLMDH /marianne_pdf_9text100K<n<1M0 likes996 downloads2y agoHugging Face11LLMDH /marianne_pdf_3text100K<n<1M0 likes975 downloads2y agoHugging Face12LLMDH /marianne_pdf_4text10K<n<100K0 likes795 downloads2y agoHugging Face13PleIAs /AMF-PDFtext100K<n<1M7 likes772 downloads2y agoHugging Face14LLMDH /hal_pdf_extratext10K<n<100K0 likes673 downloads2y agoHugging Face15LLMDH /marianne_pdf_8text100K<n<1M0 likes672 downloads2y agoHugging Face16LLMDH /marianne_pdf_10text100K<n<1M0 likes559 downloads2y agoHugging Face17PleIAs /Math-PDFtext100K<n<1M2 likes458 downloads2y agoHugging Face18davanstrien /chrono-no-pdfimage10K<n<100K0 likes9 downloads1y agoHugging Face19smatiush /pdfa-eng-wds Dataset Card for PDF Association dataset (PDFA) Dataset Summary PDFA dataset is a document dataset filtered from the SafeDocs corpus, aka CC-MAIN-2021-31-PDF-UNTRUNCATED. The original purpose of that corpus is for comprehensive pdf documents analysis. The purpose of that subset differs in that regard, as focus has been done on making the dataset machine learning-ready for vision-language models. An example page of one pdf document, with added bounding boxes… See the full description on the dataset page: https://huggingface.co/datasets/smatiush/pdfa-eng-wds.textimage-to-text1M<n<10M0 likes8 downloads5mo agoHugging Face20ahmetklnc /pdf-indirilmeyentext10K<n<100K0 likes6 downloads6mo agoHugging Face21thinking-bio-lab /selected_cns_pdfgatedtext100K<n<1M0 likes3 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.