Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01eduagarcia-temp /roberta-pt-checkpoints0 likes1.6k downloads3y agoHugging Face02gsgoncalves /roberta_pretrain Dataset Card for RoBERTa Pretrain Dataset Summary This is the concatenation of the datasets used to Pretrain RoBERTa. The dataset is not shuffled and contains raw text. It is packaged for convenicence. Essentially is the same as: from datasets import load_dataset, concatenate_datasets bookcorpus = load_dataset("bookcorpus", split="train") openweb = load_dataset("openwebtext", split="train") cc_news = load_dataset("cc_news", split="train") cc_news =… See the full description on the dataset page: https://huggingface.co/datasets/gsgoncalves/roberta_pretrain.textfill-mask10M<n<100M5 likes398 downloads3y agoHugging Face03elricwan /roberta-data10M<n<100M0 likes246 downloads5y agoHugging Face04tursunait /roberta-pii-synth Synthetic PII Detection Dataset (RoBERTa-PII-Synth) A large-scale, fully synthetic dataset for training token-classification models to detect Personally Identifiable Information (PII) in realistic text. This dataset was built using an enhanced synthetic generation pipeline, designed to better capture the linguistic and formatting variability of real-world user text. All samples are fully artificial — no real people or identifiers appear anywhere. 📘 Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/tursunait/roberta-pii-synth.token-classification100K<n<1M0 likes209 downloads10mo agoHugging Face05AnonymousSub /recipe_RL_data_ONLY_CORRECT_roberta-base1M<n<10M0 likes163 downloads4y agoHugging Face06AnonymousSub /recipe_RL_data_roberta-base Dataset Description Structure Consists of 5 fields Each row corresponds to a policy - sequence of actions, given an initial <START> state, and corresponding rewards at each step. Fields steps, step_attn_masks, rewards, actions, dones Field descriptions steps (List of lists of Ints) - tokenized step tokens of all the steps in the policy sequence (here we use the roberta-base tokenizer, as roberta-base would be used to encode each step of a recipe)… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousSub/recipe_RL_data_roberta-base.1M<n<10M0 likes146 downloads4y agoHugging Face07closji /wikitext-103-raw-v1_sents_min_len10_max_len30_princeton-nlp_sup-simcse-roberta-largetext1M<n<10M0 likes134 downloads4y agoHugging Face08closji /cc12m_princeton-nlp_sup-simcse-roberta-largeimage10M<n<100M0 likes131 downloads4y agoHugging Face09AnonymousSub /recipe_RL_data_roberta-base_BASE_REWARD_32_STEPDIFF_reward1M<n<10M0 likes113 downloads4y agoHugging Face10TannerGladson /chess-roberta-basetabular100M<n<1B0 likes113 downloads2y agoHugging Face11Aadithyak /roberta-largastabular100K<n<1M0 likes64 downloads2y agoHugging Face12joegolk /roberta-tokenized-data1M<n<10M0 likes59 downloads9mo agoHugging Face13xorushi /roberta-pii-synth Synthetic PII Detection Dataset (RoBERTa-PII-Synth) A large-scale, fully synthetic dataset for training token-classification models to detect Personally Identifiable Information (PII) in realistic text. This dataset was built using an enhanced synthetic generation pipeline, designed to better capture the linguistic and formatting variability of real-world user text. All samples are fully artificial — no real people or identifiers appear anywhere. 📘 Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/xorushi/roberta-pii-synth.texttoken-classification100K<n<1M0 likes58 downloads1mo agoHugging Face14Warawreh /PII-cleaned-roberta-classes-merged-ignoretext10K<n<100K0 likes57 downloads1y agoHugging Face15hriaz /wikitext-tags-roberta1M<n<10M0 likes54 downloads1y agoHugging Face16peeper /vitmae-roberta-processed Dataset Card for "vitmae-roberta-processed" More Information needed 1K<n<10K0 likes50 downloads4y agoHugging Face17emilylearning /cond_ft_none_on_reddit__prcnt_100__test_run_False__xlm-roberta-base1M<n<10M0 likes49 downloads4y agoHugging Face18insub /imdb_prefix20_forDPO_gpt2-large-imdb-FT_siebert_sentiment-roberta-large-english Dataset Card for "imdb_prefix20_forDPO_gpt2-large-imdb-FT_siebert_sentiment-roberta-large-english" 1. Purpose of creating the dataset For reproduction of DPO (direct preference optimization) thesis experiments(https://arxiv.org/abs/2305.18290) 2. How data is produced To reproduce the paper's experimental results, we need (x, chosen, rejected) data.However, imdb data only contains good or bad reviews, so the data must be readjusted. 2.1 prepare imdb… See the full description on the dataset page: https://huggingface.co/datasets/insub/imdb_prefix20_forDPO_gpt2-large-imdb-FT_siebert_sentiment-roberta-large-english.text10K<n<100K2 likes45 downloads3y agoHugging Face19emilylearning /cond_ft_none_on_reddit__prcnt_100__test_run_False__roberta-base1M<n<10M0 likes41 downloads4y agoHugging Face20emilylearning /cond_ft_subreddit_on_reddit__prcnt_100__test_run_False__roberta-base1M<n<10M0 likes40 downloads4y agoHugging Face21nikchar /retrieval_verification_roberta Dataset Card for "retrieval_verification_roberta" More Information needed tabular10K<n<100K0 likes39 downloads3y agoHugging Face22bambadij /COVID_Vaccine_Tweet_sentiment_analysis_roberta Dataset Card for "COVID_Vaccine_Tweet_sentiment_analysis_roberta" More Information needed 1K<n<10K1 likes38 downloads3y agoHugging Face23open-llm-leaderboard /jaredjoss__pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model-detailsgated Dataset Card for Evaluation run of jaredjoss/pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model Dataset automatically created during the evaluation run of model jaredjoss/pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jaredjoss__pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model-details.tabular10K<n<100K0 likes38 downloads2y agoHugging Face24zen-E /NEWS5M-simcse-roberta-large-embeddings-pca-256A dataset that contains all data in 'ffgcc/NEWS5M' which the corresponding text embedding produced by 'princeton-nlp/unsup-simcse-roberta-large'. The features are transformed to a size of 256 by PCA. The usage: news5M_kd_pca_dataset_unsup = torch.load('./NEWS5M-simcse-roberta-large-embeddings-pca-256/news5M_kd_pca_dataset_unsup.pt') sentence-similarity1M<n<10M0 likes34 downloads3y agoHugging Face25Shweta-singh /test_data_roberta_basetabular10K<n<100K0 likes32 downloads2y agoHugging Face26fantasticrambo /covid-tweet-sentiment-analyzer-roberta-latest-data1K<n<10K1 likes31 downloads3y agoHugging Face27spoiled /ecqa_model_generate_robertatext10K<n<100K0 likes30 downloads4y agoHugging Face28emilylearning /cond_ft_subreddit_on_reddit__prcnt_na__test_run_True__roberta-base0 likes30 downloads4y agoHugging Face29oeg /CelebA_RoBERTa_Sp Corpus Summary This corpus contains 250000 entries made up of a pair of sentences in Spanish and their respective similarity value in the range 0 to 1. This corpus was used in the training of the sentence-transformer library to improve the efficiency of the RoBERTa-large-bne base model. Each of the pairs of sentences are textual descriptions of the faces of the CelebA dataset, which were previously translated into Spanish. The process followed to generate it was: First, a… See the full description on the dataset page: https://huggingface.co/datasets/oeg/CelebA_RoBERTa_Sp.texttable-question-answering100K<n<1M1 likes29 downloads3y agoHugging Face30mor40 /tokenized-chitanka-roberta100K<n<1M0 likes29 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.