JackBAI/bert_pretrain_datasets
Dataset Card for "bert_pretrain_datasets" This dataset is essentially a concatenation of the training set of the English Wikipedia (wikipedia.20220301.en.train) and the Book Corpus (bookcorpus.train). This is exactly how I get this dataset: from datasets import load_dataset, concatenate_datasets, load_from_disk cache_dir = "/data/haob2/cache/" # book corpus bookcorpus = load_dataset("bookcorpus", split="train", cache_dir=cache_dir) # english wikipedia wiki =… See the full description on the dataset page: https://huggingface.co/datasets/JackBAI/bert_pretrain_datasets.
This repository belongs to JackBAI on Hugging Face.
Team Ai never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
