datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
conceptual-captions-12m-webdataset-bertsmultilingual-pl-bertAttribution: Wikipedia.org
bert-128-grouped
Dataset Card for "bert-128-grouped"
More Information needed
SimpleSafetyTeststemp-bert-train-tokenizedBERTJudge-Dataset
BERTJudge-Dataset
Dataset Description
BERTJudge-Dataset is the training dataset used for developing BERTJudge models, as introduced in the paper BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation. It comprises question–candidate–reference triplets generated by 36 recent open-weight, instruction-tuned models across 7 established tasks, and synthetically annotated using nvidia/Llama-3_3-Nemotron-Super-49B-v1_5.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/BERTJudge-Dataset.bert-base-uncased-refined-web-segment0
Dataset Card for "bert-base-uncased-refined-web-segment0"
More Information needed
CC-MAIN-2023-50CC-MAIN-2021-49-rawCC-MAIN-2022-40-rawbert_pretrain
Dataset Card for "bert_pretrain"
More Information needed
CC-MAIN-2022-27-rawCC-MAIN-2022-21-rawCC-MAIN-2022-05-rawbert-mlm-experiments-en
Unified English MLM Pre-training Corpus (80M Rows)
This dataset is a massive, diverse, multi-domain English text corpus explicitly engineered for pre-training and domain-adaptation of BERT-style models via Masked Language Modeling (MLM). It aggregates over 80 million rows of text, completely stripped of auxiliary metadata, labels, and identifiers to expose purely raw text strings.
Dataset Details
Repository ID: 8Opt/bert-mlm-experiments-en
Total Rows: 80,489,226… See the full description on the dataset page: https://huggingface.co/datasets/LakoreAI/bert-mlm-experiments-en.mc4-es-sampled50 million documents in Spanish extracted from mC4 applying perplexity sampling via mc4-sampling: "https://huggingface.co/datasets/bertin-project/mc4-sampling". Please, refer to BERTIN Project. The original dataset is the Multlingual Colossal, Cleaned version of Common Crawl's web crawl corpus (mC4), based on the Common Crawl dataset: "https://commoncrawl.org", and processed by AllenAI.CC-MAIN-2023-06-rawbert_dataset_202203
Dataset Card for "bert_dataset_202203"
More Information needed
bert_cot_em
Can you tell a model is about to misbehave by reading its reasoning?
Short answer: no — but you can change what it does by writing its reasoning for it.
This repo studies a large language model that has been deliberately made
misaligned, and asks whether its chain-of-thought (the "thinking out loud" it
does before answering) gives away that a harmful answer is coming.
The setup in plain terms
Researchers found that fine-tuning a model on bad medical advice makes… See the full description on the dataset page: https://huggingface.co/datasets/mild-rgb/bert_cot_em.CC-MAIN-2021-25-rawnomic-bert-2048-pretraining-data
Dataset Card for "bert-pretokenized-2048-wiki-2023"
More Information needed
9router-dataCC-MAIN-2023-14-rawbert_pretrain_datasets
Dataset Card for "bert_pretrain_datasets"
This dataset is essentially a concatenation of the training set of the English Wikipedia (wikipedia.20220301.en.train) and the Book Corpus (bookcorpus.train).
This is exactly how I get this dataset:
from datasets import load_dataset, concatenate_datasets, load_from_disk
cache_dir = "/data/haob2/cache/"
# book corpus
bookcorpus = load_dataset("bookcorpus", split="train", cache_dir=cache_dir)
# english wikipedia
wiki =… See the full description on the dataset page: https://huggingface.co/datasets/JackBAI/bert_pretrain_datasets.processed_bert_dataset
Dataset Card for "processed_bert_dataset"
More Information needed
CC-MAIN-2022-33-rawprocessed_bert_dataset
Dataset Card for "processed_bert_dataset"
More Information needed
CC-MAIN-2021-10-rawCC-MAIN-2023-40-rawbert_pretraining_data
Dataset Card for "bert_pretraining_data"
More Information needed
