Team Ai
Datasetpublic

JackBAI/bert_pretrain_datasets

Dataset Card for "bert_pretrain_datasets" This dataset is essentially a concatenation of the training set of the English Wikipedia (wikipedia.20220301.en.train) and the Book Corpus (bookcorpus.train). This is exactly how I get this dataset: from datasets import load_dataset, concatenate_datasets, load_from_disk cache_dir = "/data/haob2/cache/" # book corpus bookcorpus = load_dataset("bookcorpus", split="train", cache_dir=cache_dir) # english wikipedia wiki =… See the full description on the dataset page: https://huggingface.co/datasets/JackBAI/bert_pretrain_datasets.

sourceHugging Faceupdated 3y agoView on Hugging Face
1likes568downloads
settings

This repository belongs to JackBAI on Hugging Face.

Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namebert_pretrain_datasets
visibilitypublic
licencenot set
gatedno
ownerJackBAI
Account settings
JackBAI/bert_pretrain_datasets · Team Ai