datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
roberta_pretrain
Dataset Card for RoBERTa Pretrain
Dataset Summary
This is the concatenation of the datasets used to Pretrain RoBERTa.
The dataset is not shuffled and contains raw text. It is packaged for convenicence.
Essentially is the same as:
from datasets import load_dataset, concatenate_datasets
bookcorpus = load_dataset("bookcorpus", split="train")
openweb = load_dataset("openwebtext", split="train")
cc_news = load_dataset("cc_news", split="train")
cc_news =… See the full description on the dataset page: https://huggingface.co/datasets/gsgoncalves/roberta_pretrain.chess-roberta-basewikitext-103-raw-v1_sents_min_len10_max_len30_princeton-nlp_sup-simcse-roberta-largecc12m_princeton-nlp_sup-simcse-roberta-largeroberta-pii-synth
Synthetic PII Detection Dataset (RoBERTa-PII-Synth)
A large-scale, fully synthetic dataset for training token-classification models to detect Personally Identifiable Information (PII) in realistic text.
This dataset was built using an enhanced synthetic generation pipeline, designed to better capture the linguistic and formatting variability of real-world user text. All samples are fully artificial — no real people or identifiers appear anywhere.
📘 Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/xorushi/roberta-pii-synth.roberta-largasPII-cleaned-roberta-classes-merged-ignoreimdb_prefix20_forDPO_gpt2-large-imdb-FT_siebert_sentiment-roberta-large-english
Dataset Card for "imdb_prefix20_forDPO_gpt2-large-imdb-FT_siebert_sentiment-roberta-large-english"
1. Purpose of creating the dataset
For reproduction of DPO (direct preference optimization) thesis experiments(https://arxiv.org/abs/2305.18290)
2. How data is produced
To reproduce the paper's experimental results, we need (x, chosen, rejected) data.However, imdb data only contains good or bad reviews, so the data must be readjusted.
2.1 prepare imdb… See the full description on the dataset page: https://huggingface.co/datasets/insub/imdb_prefix20_forDPO_gpt2-large-imdb-FT_siebert_sentiment-roberta-large-english.jaredjoss__pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model-details
Dataset Card for Evaluation run of jaredjoss/pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model
Dataset automatically created during the evaluation run of model jaredjoss/pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jaredjoss__pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model-details.retrieval_verification_roberta
Dataset Card for "retrieval_verification_roberta"
More Information needed
ecqa_model_generate_robertaCelebA_RoBERTa_Sp
Corpus Summary
This corpus contains 250000 entries made up of a pair of sentences in Spanish and their respective similarity value in the range 0 to 1. This corpus was used in the training of the
sentence-transformer library to improve the efficiency of the RoBERTa-large-bne base model.
Each of the pairs of sentences are textual descriptions of the faces of the CelebA dataset, which were previously translated into Spanish. The process followed to generate it was:
First, a… See the full description on the dataset page: https://huggingface.co/datasets/oeg/CelebA_RoBERTa_Sp.mscoco_2014_captions_princeton-nlp_sup-simcse-roberta-largeag_news_roberta_keywordsroberta_datasetbbq_roberta_large_race_custom_loss_our_datasetretrieval_verification_bm25_roberta
Dataset Card for "retrieval_verification_bm25_roberta"
More Information needed
imdb_prefix3_forDPO_gpt2-large-imdb-FT_siebert_sentiment-roberta-large-english
Dataset Card for "imdb_prefix3_forDPO_gpt2-large-imdb-FT_siebert_sentiment-roberta-large-english"
More Information needed
test_data_roberta_baseparsed-dataset-xlm-robertahc3-wiki-cleaned-text-for-domain-classification-roberta-tokenized-max-len-512
Dataset Card for "hc3-wiki-cleaned-text-for-domain-classification-roberta-tokenized-max-len-512"
More Information needed
dolly-code-migrationsst_roberta_keywords_embeddingswsd_myriade_synth_data_multilabel_robertachess-roberta-pretraining-sansconfigs:
config_name: default
data_files:
split: train
path: train/*.csv
split: eval
path: eval/*.csv
rapidapi-example-responses-tokenized-xlm-roberta
Dataset Card for "rapidapi-example-responses-tokenized-xlm-roberta"
More Information needed
rating_regression_no0_roberta_train.hfflickr30k_princeton-nlp_sup-simcse-roberta-largeSpirit_RoBERTa_FT
Dataset Card for "Spirit_RoBERTa_FT"
More Information needed
SA_RoBERTa_based
