Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AbstractPhil /conceptual-captions-12m-webdataset-bertstext10M<n<100M1 likes6.6k downloads2mo agoHugging Face02styletts2-community /multilingual-pl-bertAttribution: Wikipedia.org text100K<n<1M19 likes5.5k downloads3y agoHugging Face03Bertievidgen /SimpleSafetyTeststexttext-generationn<1K12 likes3.1k downloads3y agoHugging Face04upup-ashton-wang /temp-bert-train-tokenizedtextn<1K0 likes2.3k downloads5mo agoHugging Face05bertram-gilfoyle /CC-MAIN-2021-49-rawtext10M<n<100M0 likes2k downloads3y agoHugging Face06bertram-gilfoyle /CC-MAIN-2022-40-rawtext10M<n<100M0 likes2k downloads3y agoHugging Face07gsgoncalves /bert_pretrain Dataset Card for "bert_pretrain" More Information needed text10M<n<100M0 likes1.9k downloads3y agoHugging Face08bertram-gilfoyle /CC-MAIN-2022-27-rawtext10M<n<100M0 likes1.6k downloads3y agoHugging Face09bertram-gilfoyle /CC-MAIN-2022-21-rawtext10M<n<100M0 likes1.6k downloads3y agoHugging Face10bertram-gilfoyle /CC-MAIN-2022-05-rawtext10M<n<100M0 likes1.5k downloads3y agoHugging Face11LakoreAI /bert-mlm-experiments-en Unified English MLM Pre-training Corpus (80M Rows) This dataset is a massive, diverse, multi-domain English text corpus explicitly engineered for pre-training and domain-adaptation of BERT-style models via Masked Language Modeling (MLM). It aggregates over 80 million rows of text, completely stripped of auxiliary metadata, labels, and identifiers to expose purely raw text strings. Dataset Details Repository ID: 8Opt/bert-mlm-experiments-en Total Rows: 80,489,226… See the full description on the dataset page: https://huggingface.co/datasets/LakoreAI/bert-mlm-experiments-en.textfill-mask10M<n<100M1 likes1.5k downloads4mo agoHugging Face12bertin-project /mc4-es-sampled50 million documents in Spanish extracted from mC4 applying perplexity sampling via mc4-sampling: "https://huggingface.co/datasets/bertin-project/mc4-sampling". Please, refer to BERTIN Project. The original dataset is the Multlingual Colossal, Cleaned version of Common Crawl's web crawl corpus (mC4), based on the Common Crawl dataset: "https://commoncrawl.org", and processed by AllenAI.texttext-generation1M<n<10M2 likes1.3k downloads4y agoHugging Face13bertram-gilfoyle /CC-MAIN-2023-06-rawtext10M<n<100M0 likes1.3k downloads3y agoHugging Face14nthngdy /bert_dataset_202203 Dataset Card for "bert_dataset_202203" More Information needed texttext-generation100M<n<1B0 likes1.3k downloads4y agoHugging Face15bertram-gilfoyle /CC-MAIN-2021-25-rawtext10M<n<100M0 likes1.1k downloads3y agoHugging Face16bertram-gilfoyle /CC-MAIN-2023-14-rawtext10M<n<100M0 likes756 downloads3y agoHugging Face17JackBAI /bert_pretrain_datasets Dataset Card for "bert_pretrain_datasets" This dataset is essentially a concatenation of the training set of the English Wikipedia (wikipedia.20220301.en.train) and the Book Corpus (bookcorpus.train). This is exactly how I get this dataset: from datasets import load_dataset, concatenate_datasets, load_from_disk cache_dir = "/data/haob2/cache/" # book corpus bookcorpus = load_dataset("bookcorpus", split="train", cache_dir=cache_dir) # english wikipedia wiki =… See the full description on the dataset page: https://huggingface.co/datasets/JackBAI/bert_pretrain_datasets.text10M<n<100M1 likes691 downloads3y agoHugging Face18bertram-gilfoyle /CC-MAIN-2022-33-rawtext10M<n<100M0 likes639 downloads3y agoHugging Face19bertram-gilfoyle /CC-MAIN-2021-10-rawtext10M<n<100M0 likes601 downloads3y agoHugging Face20bertram-gilfoyle /CC-MAIN-2023-40-rawtext10M<n<100M0 likes572 downloads3y agoHugging Face21bertin-project /alpaca-spanish BERTIN Alpaca Spanish This dataset is a translation to Spanish of alpaca_data_cleaned.json, a clean version of the Alpaca dataset made at Stanford. An earlier version used Facebook's NLLB 1.3B model, but the current version uses OpenAI's gpt-3.5-turbo, hence this dataset cannot be used to create models that compete in any way against OpenAI. texttext-generation10K<n<100K36 likes547 downloads4y agoHugging Face22bertram-gilfoyle /CC-MAIN-2022-49-rawtext10M<n<100M0 likes499 downloads3y agoHugging Face23surogate /fineweb2-ro-bert FineWeb2-Ro-BERT FineWeb2-Ro-BERT is a large-scale pretraining dataset in the Romanian language. The data is derived from FineWeb2 and annotated using a bert architecture for signals such as educational quality or topic. More details can be found here. Key Features Massive Scale: Contains approximately 54.1M rows (documents or sequences), providing comprehensive linguistic coverage for training robust Romanian embeddings and encoders. Usage You… See the full description on the dataset page: https://huggingface.co/datasets/surogate/fineweb2-ro-bert.tabular10M<n<100M0 likes466 downloads1mo agoHugging Face24HiTZ /BertaQA Dataset Card for BertaQA BertaQA is a trivia dataset comprising 4,756 multiple-choice trivia questions, with one single correct answer and 2 additional distractors. Crucially, questions are distributed between local and global topics. Whereas answering questions in the latter group requires general world knowledge, local questions require specific knowledge about the Basque Country and its culture. Additionally, questions are classified into eight categories, namely Basque and… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/BertaQA.tabularquestion-answering10K<n<100K1 likes465 downloads2y agoHugging Face25bertram-gilfoyle /CC-MAIN-2020-29-rawtext10M<n<100M0 likes450 downloads3y agoHugging Face26bertram-gilfoyle /CC-MAIN-2023-50-rawtext10M<n<100M0 likes433 downloads3y agoHugging Face27m-a-p /FineFineWeb-bert-seeddata FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-bert-seeddata.texttext-classification1M<n<10M2 likes399 downloads2y agoHugging Face28gmongaras /BERT_Base_Cased_512_DatasetDataset using the bert-cased tokenizer, cutoff sentences to 512 length (not sentence pairs), all sentence pairs extracted. Original datasets: https://huggingface.co/datasets/bookcorpus https://huggingface.co/datasets/wikipedia Variant: 20220301.en text100M<n<1B0 likes397 downloads3y agoHugging Face29bertram-gilfoyle /CC-MAIN-2020-50text1M<n<10M0 likes348 downloads3y agoHugging Face30bertram-gilfoyle /CC-MAIN-2023-23-rawtext10M<n<100M0 likes340 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.