datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
roberta-pt-checkpointsroberta_pretrain
Dataset Card for RoBERTa Pretrain
Dataset Summary
This is the concatenation of the datasets used to Pretrain RoBERTa.
The dataset is not shuffled and contains raw text. It is packaged for convenicence.
Essentially is the same as:
from datasets import load_dataset, concatenate_datasets
bookcorpus = load_dataset("bookcorpus", split="train")
openweb = load_dataset("openwebtext", split="train")
cc_news = load_dataset("cc_news", split="train")
cc_news =… See the full description on the dataset page: https://huggingface.co/datasets/gsgoncalves/roberta_pretrain.roberta-dataroberta-pii-synth
Synthetic PII Detection Dataset (RoBERTa-PII-Synth)
A large-scale, fully synthetic dataset for training token-classification models to detect Personally Identifiable Information (PII) in realistic text.
This dataset was built using an enhanced synthetic generation pipeline, designed to better capture the linguistic and formatting variability of real-world user text. All samples are fully artificial — no real people or identifiers appear anywhere.
📘 Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/tursunait/roberta-pii-synth.recipe_RL_data_ONLY_CORRECT_roberta-baserecipe_RL_data_roberta-base
Dataset Description
Structure
Consists of 5 fields
Each row corresponds to a policy - sequence of actions, given an initial <START> state, and corresponding rewards at each step.
Fields
steps, step_attn_masks, rewards, actions, dones
Field descriptions
steps (List of lists of Ints) - tokenized step tokens of all the steps in the policy sequence (here we use the roberta-base tokenizer, as roberta-base would be used to encode each step of a recipe)… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousSub/recipe_RL_data_roberta-base.wikitext-103-raw-v1_sents_min_len10_max_len30_princeton-nlp_sup-simcse-roberta-largecc12m_princeton-nlp_sup-simcse-roberta-largerecipe_RL_data_roberta-base_BASE_REWARD_32_STEPDIFF_rewardchess-roberta-baseroberta-largasroberta-tokenized-dataroberta-pii-synth
Synthetic PII Detection Dataset (RoBERTa-PII-Synth)
A large-scale, fully synthetic dataset for training token-classification models to detect Personally Identifiable Information (PII) in realistic text.
This dataset was built using an enhanced synthetic generation pipeline, designed to better capture the linguistic and formatting variability of real-world user text. All samples are fully artificial — no real people or identifiers appear anywhere.
📘 Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/xorushi/roberta-pii-synth.PII-cleaned-roberta-classes-merged-ignorewikitext-tags-robertavitmae-roberta-processed
Dataset Card for "vitmae-roberta-processed"
More Information needed
cond_ft_none_on_reddit__prcnt_100__test_run_False__xlm-roberta-baseimdb_prefix20_forDPO_gpt2-large-imdb-FT_siebert_sentiment-roberta-large-english
Dataset Card for "imdb_prefix20_forDPO_gpt2-large-imdb-FT_siebert_sentiment-roberta-large-english"
1. Purpose of creating the dataset
For reproduction of DPO (direct preference optimization) thesis experiments(https://arxiv.org/abs/2305.18290)
2. How data is produced
To reproduce the paper's experimental results, we need (x, chosen, rejected) data.However, imdb data only contains good or bad reviews, so the data must be readjusted.
2.1 prepare imdb… See the full description on the dataset page: https://huggingface.co/datasets/insub/imdb_prefix20_forDPO_gpt2-large-imdb-FT_siebert_sentiment-roberta-large-english.cond_ft_none_on_reddit__prcnt_100__test_run_False__roberta-basecond_ft_subreddit_on_reddit__prcnt_100__test_run_False__roberta-baseretrieval_verification_roberta
Dataset Card for "retrieval_verification_roberta"
More Information needed
COVID_Vaccine_Tweet_sentiment_analysis_roberta
Dataset Card for "COVID_Vaccine_Tweet_sentiment_analysis_roberta"
More Information needed
jaredjoss__pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model-details
Dataset Card for Evaluation run of jaredjoss/pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model
Dataset automatically created during the evaluation run of model jaredjoss/pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jaredjoss__pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model-details.NEWS5M-simcse-roberta-large-embeddings-pca-256A dataset that contains all data in 'ffgcc/NEWS5M' which the corresponding text embedding produced by 'princeton-nlp/unsup-simcse-roberta-large'. The features are transformed to a size of 256 by PCA.
The usage:
news5M_kd_pca_dataset_unsup = torch.load('./NEWS5M-simcse-roberta-large-embeddings-pca-256/news5M_kd_pca_dataset_unsup.pt')
test_data_roberta_basecovid-tweet-sentiment-analyzer-roberta-latest-dataecqa_model_generate_robertacond_ft_subreddit_on_reddit__prcnt_na__test_run_True__roberta-baseCelebA_RoBERTa_Sp
Corpus Summary
This corpus contains 250000 entries made up of a pair of sentences in Spanish and their respective similarity value in the range 0 to 1. This corpus was used in the training of the
sentence-transformer library to improve the efficiency of the RoBERTa-large-bne base model.
Each of the pairs of sentences are textual descriptions of the faces of the CelebA dataset, which were previously translated into Spanish. The process followed to generate it was:
First, a… See the full description on the dataset page: https://huggingface.co/datasets/oeg/CelebA_RoBERTa_Sp.tokenized-chitanka-roberta
