Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AbstractPhil /conceptual-captions-12m-webdataset-bertstext10M<n<100M1 likes6.6k downloads2mo agoHugging Face02styletts2-community /multilingual-pl-bertAttribution: Wikipedia.org text100K<n<1M19 likes5.5k downloads3y agoHugging Face03nomic-ai /bert-128-grouped Dataset Card for "bert-128-grouped" More Information needed 10M<n<100M0 likes4.4k downloads3y agoHugging Face04Bertievidgen /SimpleSafetyTeststexttext-generationn<1K12 likes3.1k downloads3y agoHugging Face05upup-ashton-wang /temp-bert-train-tokenizedtextn<1K0 likes2.3k downloads5mo agoHugging Face06artefactory /BERTJudge-Dataset BERTJudge-Dataset Dataset Description BERTJudge-Dataset is the training dataset used for developing BERTJudge models, as introduced in the paper BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation. It comprises question–candidate–reference triplets generated by 36 recent open-weight, instruction-tuned models across 7 established tasks, and synthetically annotated using nvidia/Llama-3_3-Nemotron-Super-49B-v1_5. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/BERTJudge-Dataset.text-classification2 likes2.2k downloads6mo agoHugging Face07Jackmin108 /bert-base-uncased-refined-web-segment0 Dataset Card for "bert-base-uncased-refined-web-segment0" More Information needed 100M<n<1B0 likes2.1k downloads3y agoHugging Face08bertram-gilfoyle /CC-MAIN-2023-500 likes2k downloads3y agoHugging Face09bertram-gilfoyle /CC-MAIN-2021-49-rawtext10M<n<100M0 likes2k downloads3y agoHugging Face10bertram-gilfoyle /CC-MAIN-2022-40-rawtext10M<n<100M0 likes2k downloads3y agoHugging Face11gsgoncalves /bert_pretrain Dataset Card for "bert_pretrain" More Information needed text10M<n<100M0 likes1.9k downloads3y agoHugging Face12bertram-gilfoyle /CC-MAIN-2022-27-rawtext10M<n<100M0 likes1.6k downloads3y agoHugging Face13bertram-gilfoyle /CC-MAIN-2022-21-rawtext10M<n<100M0 likes1.6k downloads3y agoHugging Face14bertram-gilfoyle /CC-MAIN-2022-05-rawtext10M<n<100M0 likes1.5k downloads3y agoHugging Face15LakoreAI /bert-mlm-experiments-en Unified English MLM Pre-training Corpus (80M Rows) This dataset is a massive, diverse, multi-domain English text corpus explicitly engineered for pre-training and domain-adaptation of BERT-style models via Masked Language Modeling (MLM). It aggregates over 80 million rows of text, completely stripped of auxiliary metadata, labels, and identifiers to expose purely raw text strings. Dataset Details Repository ID: 8Opt/bert-mlm-experiments-en Total Rows: 80,489,226… See the full description on the dataset page: https://huggingface.co/datasets/LakoreAI/bert-mlm-experiments-en.textfill-mask10M<n<100M1 likes1.5k downloads4mo agoHugging Face16bertin-project /mc4-es-sampled50 million documents in Spanish extracted from mC4 applying perplexity sampling via mc4-sampling: "https://huggingface.co/datasets/bertin-project/mc4-sampling". Please, refer to BERTIN Project. The original dataset is the Multlingual Colossal, Cleaned version of Common Crawl's web crawl corpus (mC4), based on the Common Crawl dataset: "https://commoncrawl.org", and processed by AllenAI.texttext-generation1M<n<10M2 likes1.3k downloads4y agoHugging Face17bertram-gilfoyle /CC-MAIN-2023-06-rawtext10M<n<100M0 likes1.3k downloads3y agoHugging Face18nthngdy /bert_dataset_202203 Dataset Card for "bert_dataset_202203" More Information needed texttext-generation100M<n<1B0 likes1.3k downloads4y agoHugging Face19mild-rgb /bert_cot_em Can you tell a model is about to misbehave by reading its reasoning? Short answer: no — but you can change what it does by writing its reasoning for it. This repo studies a large language model that has been deliberately made misaligned, and asks whether its chain-of-thought (the "thinking out loud" it does before answering) gives away that a harmful answer is coming. The setup in plain terms Researchers found that fine-tuning a model on bad medical advice makes… See the full description on the dataset page: https://huggingface.co/datasets/mild-rgb/bert_cot_em.0 likes1.3k downloads27d agoHugging Face20bertram-gilfoyle /CC-MAIN-2021-25-rawtext10M<n<100M0 likes1.1k downloads3y agoHugging Face21nomic-ai /nomic-bert-2048-pretraining-data Dataset Card for "bert-pretokenized-2048-wiki-2023" More Information needed 1M<n<10M1 likes953 downloads3y agoHugging Face22bertandd /9router-data1 likes935 downloads3d agoHugging Face23bertram-gilfoyle /CC-MAIN-2023-14-rawtext10M<n<100M0 likes756 downloads3y agoHugging Face24JackBAI /bert_pretrain_datasets Dataset Card for "bert_pretrain_datasets" This dataset is essentially a concatenation of the training set of the English Wikipedia (wikipedia.20220301.en.train) and the Book Corpus (bookcorpus.train). This is exactly how I get this dataset: from datasets import load_dataset, concatenate_datasets, load_from_disk cache_dir = "/data/haob2/cache/" # book corpus bookcorpus = load_dataset("bookcorpus", split="train", cache_dir=cache_dir) # english wikipedia wiki =… See the full description on the dataset page: https://huggingface.co/datasets/JackBAI/bert_pretrain_datasets.text10M<n<100M1 likes691 downloads3y agoHugging Face25Tural /processed_bert_dataset Dataset Card for "processed_bert_dataset" More Information needed 10M<n<100M0 likes641 downloads3y agoHugging Face26bertram-gilfoyle /CC-MAIN-2022-33-rawtext10M<n<100M0 likes639 downloads3y agoHugging Face27delmeng /processed_bert_dataset Dataset Card for "processed_bert_dataset" More Information needed 1M<n<10M0 likes603 downloads3y agoHugging Face28bertram-gilfoyle /CC-MAIN-2021-10-rawtext10M<n<100M0 likes601 downloads3y agoHugging Face29bertram-gilfoyle /CC-MAIN-2023-40-rawtext10M<n<100M0 likes572 downloads3y agoHugging Face30angie-chen55 /bert_pretraining_data Dataset Card for "bert_pretraining_data" More Information needed 10M<n<100M0 likes560 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.