Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gsgoncalves /roberta_pretrain Dataset Card for RoBERTa Pretrain Dataset Summary This is the concatenation of the datasets used to Pretrain RoBERTa. The dataset is not shuffled and contains raw text. It is packaged for convenicence. Essentially is the same as: from datasets import load_dataset, concatenate_datasets bookcorpus = load_dataset("bookcorpus", split="train") openweb = load_dataset("openwebtext", split="train") cc_news = load_dataset("cc_news", split="train") cc_news =… See the full description on the dataset page: https://huggingface.co/datasets/gsgoncalves/roberta_pretrain.textfill-mask10M<n<100M5 likes314 downloads3y agoHugging Face02TannerGladson /chess-roberta-basetabular100M<n<1B0 likes113 downloads2y agoHugging Face03closji /wikitext-103-raw-v1_sents_min_len10_max_len30_princeton-nlp_sup-simcse-roberta-largetext1M<n<10M0 likes102 downloads4y agoHugging Face04closji /cc12m_princeton-nlp_sup-simcse-roberta-largeimage10M<n<100M0 likes94 downloads4y agoHugging Face05xorushi /roberta-pii-synth Synthetic PII Detection Dataset (RoBERTa-PII-Synth) A large-scale, fully synthetic dataset for training token-classification models to detect Personally Identifiable Information (PII) in realistic text. This dataset was built using an enhanced synthetic generation pipeline, designed to better capture the linguistic and formatting variability of real-world user text. All samples are fully artificial — no real people or identifiers appear anywhere. 📘 Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/xorushi/roberta-pii-synth.texttoken-classification100K<n<1M0 likes64 downloads2mo agoHugging Face06Aadithyak /roberta-largastabular100K<n<1M0 likes62 downloads2y agoHugging Face07Warawreh /PII-cleaned-roberta-classes-merged-ignoretext10K<n<100K0 likes57 downloads1y agoHugging Face08insub /imdb_prefix20_forDPO_gpt2-large-imdb-FT_siebert_sentiment-roberta-large-english Dataset Card for "imdb_prefix20_forDPO_gpt2-large-imdb-FT_siebert_sentiment-roberta-large-english" 1. Purpose of creating the dataset For reproduction of DPO (direct preference optimization) thesis experiments(https://arxiv.org/abs/2305.18290) 2. How data is produced To reproduce the paper's experimental results, we need (x, chosen, rejected) data.However, imdb data only contains good or bad reviews, so the data must be readjusted. 2.1 prepare imdb… See the full description on the dataset page: https://huggingface.co/datasets/insub/imdb_prefix20_forDPO_gpt2-large-imdb-FT_siebert_sentiment-roberta-large-english.text10K<n<100K2 likes48 downloads3y agoHugging Face09open-llm-leaderboard /jaredjoss__pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model-detailsgated Dataset Card for Evaluation run of jaredjoss/pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model Dataset automatically created during the evaluation run of model jaredjoss/pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jaredjoss__pythia-410m-roberta-lr_8e7-kl_01-steps_12000-rlhf-model-details.tabular10K<n<100K0 likes39 downloads2y agoHugging Face10nikchar /retrieval_verification_roberta Dataset Card for "retrieval_verification_roberta" More Information needed tabular10K<n<100K0 likes32 downloads3y agoHugging Face11spoiled /ecqa_model_generate_robertatext10K<n<100K0 likes29 downloads4y agoHugging Face12oeg /CelebA_RoBERTa_Sp Corpus Summary This corpus contains 250000 entries made up of a pair of sentences in Spanish and their respective similarity value in the range 0 to 1. This corpus was used in the training of the sentence-transformer library to improve the efficiency of the RoBERTa-large-bne base model. Each of the pairs of sentences are textual descriptions of the faces of the CelebA dataset, which were previously translated into Spanish. The process followed to generate it was: First, a… See the full description on the dataset page: https://huggingface.co/datasets/oeg/CelebA_RoBERTa_Sp.texttable-question-answering100K<n<1M1 likes29 downloads3y agoHugging Face13closji /mscoco_2014_captions_princeton-nlp_sup-simcse-roberta-largetext100K<n<1M0 likes22 downloads4y agoHugging Face14sunhaozhepy /ag_news_roberta_keywordstext100K<n<1M1 likes22 downloads3y agoHugging Face15ejdis /roberta_datasettext100M<n<1B0 likes22 downloads2y agoHugging Face16artianand /bbq_roberta_large_race_custom_loss_our_datasettabular10K<n<100K0 likes22 downloads1y agoHugging Face17nikchar /retrieval_verification_bm25_roberta Dataset Card for "retrieval_verification_bm25_roberta" More Information needed tabular10K<n<100K0 likes21 downloads3y agoHugging Face18insub /imdb_prefix3_forDPO_gpt2-large-imdb-FT_siebert_sentiment-roberta-large-english Dataset Card for "imdb_prefix3_forDPO_gpt2-large-imdb-FT_siebert_sentiment-roberta-large-english" More Information needed text10K<n<100K1 likes21 downloads3y agoHugging Face19Shweta-singh /test_data_roberta_basetabular10K<n<100K0 likes21 downloads2y agoHugging Face20merve /parsed-dataset-xlm-robertatextn<1K0 likes20 downloads4y agoHugging Face21rajendrabaskota /hc3-wiki-cleaned-text-for-domain-classification-roberta-tokenized-max-len-512 Dataset Card for "hc3-wiki-cleaned-text-for-domain-classification-roberta-tokenized-max-len-512" More Information needed tabular100K<n<1M0 likes20 downloads3y agoHugging Face22robert-altmiller /dolly-code-migrationtextn<1K1 likes19 downloads3y agoHugging Face23sunhaozhepy /sst_roberta_keywords_embeddingstext10K<n<100K0 likes19 downloads3y agoHugging Face24gguichard /wsd_myriade_synth_data_multilabel_robertatext10K<n<100K0 likes18 downloads2y agoHugging Face25TannerGladson /chess-roberta-pretraining-sansconfigs: config_name: default data_files: split: train path: train/*.csv split: eval path: eval/*.csv tabular100M<n<1B0 likes18 downloads2y agoHugging Face26davidfant /rapidapi-example-responses-tokenized-xlm-roberta Dataset Card for "rapidapi-example-responses-tokenized-xlm-roberta" More Information needed text10K<n<100K0 likes17 downloads3y agoHugging Face27pppereira3 /rating_regression_no0_roberta_train.hftext1K<n<10K0 likes17 downloads2y agoHugging Face28closji /flickr30k_princeton-nlp_sup-simcse-roberta-largetext100K<n<1M0 likes16 downloads4y agoHugging Face29EgilKarlsen /Spirit_RoBERTa_FT Dataset Card for "Spirit_RoBERTa_FT" More Information needed tabular10K<n<100K0 likes16 downloads3y agoHugging Face30enansari /SA_RoBERTa_basedtext10K<n<100K0 likes15 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.