Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01princeton-nlp /datasets-for-simcsetext1M<n<10M8 likes1k downloads5y agoHugging Face02closji /bookcorpus_filtered_len_17_simcsetext10M<n<100M0 likes214 downloads4y agoHugging Face03sentence-transformers /nli-for-simcse Dataset Card for NLI for SimCSE This is a reformatting of the NLI for SimCSE Dataset used to train the BGE-M3 model. See the full BGE-M3 dataset in Shitao/bge-m3-data. Despite being labeled as Natural Language Inference (NLI), this dataset can be used for training/finetuning an embedding model for semantic textual similarity. Dataset Subsets triplet subset Columns: "anchor", "positive", "negative" Column types: str, str, str Examples:{ 'anchor': 'One… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/nli-for-simcse.textfeature-extraction1M<n<10M3 likes117 downloads2y agoHugging Face04closji /wikitext-103-raw-v1_sents_min_len10_max_len30_princeton-nlp_sup-simcse-roberta-largetext1M<n<10M0 likes102 downloads4y agoHugging Face05closji /cc12m_princeton-nlp_sup-simcse-roberta-largeimage10M<n<100M0 likes94 downloads4y agoHugging Face06llm-book /jawiki-paragraphs-unsup-simcse-bert-base-japanese-v3 Dataset Card for "jawiki-paragraphs-unsup-simcse-bert-base-japanese-v3" More Information needed tabular1M<n<10M1 likes87 downloads3y agoHugging Face07sentence-transformers /wiki1m-for-simcse Dataset Card for Wiki1m for SimCSE This is a reupload of the wiki1m_for_simcse.txt file from princeton-nlp/datasets-for-simcse, which can no longer be downloaded with recent datasets versions. Columns: "text" Column types: str Examples:{'text': 'YMCA in South Australia'} Collection strategy: Downloading the princeton-nlp/datasets-for-simcse dataset with datasets==2.21.0 and reuploading it to make the format compatible with datasets. Deduplicated: No textfeature-extraction1M<n<10M1 likes77 downloads9mo agoHugging Face08dkoterwa /kor_nli_simcse Korean Natural Language Inference (KorNLI) for SimCSE Dataset For a better dataset description, please visit this GitHub repository prepared by the authors of the article: LINK This dataset was prepared by converting KorNLI dataset. I took every unique premise of the dataset and searched for its entailment (positive example) and contradiction (negative example). These changes have been made in order to apply SimCSE method. I additionaly share the code, which I used to convert the… See the full description on the dataset page: https://huggingface.co/datasets/dkoterwa/kor_nli_simcse.text100K<n<1M1 likes62 downloads3y agoHugging Face09closji /flickr30k_CLIP_ViT-B-32_subset_pairs_SimCSE_similarity_copytabular10M<n<100M0 likes50 downloads4y agoHugging Face10andersonbcdefg /simcse_nlitext100K<n<1M0 likes42 downloads3y agoHugging Face11closji /bookcorpus_filtered_len_17_simcse_retrieval_top32__source_tranch_14__target_tranch_11__from_120text100K<n<1M0 likes35 downloads4y agoHugging Face12closji /bookcorpus_filtered_len_17_simcse_retrieval_top32__source_tranch_17__target_tranch_4__from_120text100K<n<1M0 likes34 downloads4y agoHugging Face13anti-ai /ViNLI-SimCSE-supervisedtextsentence-similarity100K<n<1M1 likes32 downloads3y agoHugging Face14emrecan /nli_tr_for_simcse NLI-TR for Supervised SimCSE This dataset is a modified version of NLI-TR dataset. Its intended use is to train Supervised SimCSE models for sentence-embeddings. Steps followed to produce this dataset are listed below: Merge train split of snli_tr and multinli_tr subsets. Find every premise that has an entailment hypothesis and a contradiction hypothesis. Write found triplets into sent0 (premise), sent1 (entailment hypothesis), hard_neg (contradiction hypothesis) format. See this… See the full description on the dataset page: https://huggingface.co/datasets/emrecan/nli_tr_for_simcse.texttext-classification100K<n<1M1 likes31 downloads4y agoHugging Face15closji /flickr30k_captions_simCSEtext100K<n<1M0 likes30 downloads4y agoHugging Face16closji /bookcorpus_filtered_len_17_simcse_retrieval_top32__source_tranch_16__target_tranch_11__from_120text100K<n<1M0 likes30 downloads4y agoHugging Face17closji /bookcorpus_filtered_len_17_simcse_retrieval_top32__source_tranch_14__target_tranch_17__from_120text100K<n<1M0 likes28 downloads4y agoHugging Face18closji /flickr30k_CLIP_ViT-B-32_subset_pairs_SimCSE_similaritytabular10M<n<100M1 likes27 downloads4y agoHugging Face19closji /bookcorpus_filtered_len_17_simcse_retrieval_top32__source_tranch_13__target_tranch_10__from_120text100K<n<1M0 likes27 downloads4y agoHugging Face20closji /bookcorpus_filtered_len_17_simcse_retrieval_top32__source_tranch_13__target_tranch_16__from_120text100K<n<1M0 likes27 downloads4y agoHugging Face21zen-E /NEWS5M-simcse-roberta-large-embeddings-pca-256A dataset that contains all data in 'ffgcc/NEWS5M' which the corresponding text embedding produced by 'princeton-nlp/unsup-simcse-roberta-large'. The features are transformed to a size of 256 by PCA. The usage: news5M_kd_pca_dataset_unsup = torch.load('./NEWS5M-simcse-roberta-large-embeddings-pca-256/news5M_kd_pca_dataset_unsup.pt') sentence-similarity1M<n<10M0 likes27 downloads3y agoHugging Face22anti-ai /ViNLI-SimCSE-supervised_v2textsentence-similarity100K<n<1M0 likes26 downloads2y agoHugging Face23closji /bookcorpus_filtered_len_17_simcse_retrieval_top32__source_tranch_13__target_tranch_19__from_120text100K<n<1M0 likes25 downloads4y agoHugging Face24closji /bookcorpus_filtered_len_17_simcse_retrieval_top32__source_tranch_17__target_tranch_18__from_120text100K<n<1M0 likes25 downloads4y agoHugging Face25closji /bookcorpus_filtered_len_17_simcse_retrieval_top32__source_tranch_13__target_tranch_9__from_120text100K<n<1M0 likes24 downloads4y agoHugging Face26closji /bookcorpus_filtered_len_17_simcse_retrieval_top32__source_tranch_19__target_tranch_28__from_120text100K<n<1M0 likes23 downloads4y agoHugging Face27closji /bookcorpus_filtered_len_17_simcse_retrieval_top32__source_tranch_18__target_tranch_9__from_120text100K<n<1M0 likes23 downloads4y agoHugging Face28closji /mscoco_2014_captions_princeton-nlp_sup-simcse-roberta-largetext100K<n<1M0 likes22 downloads4y agoHugging Face29closji /bookcorpus_filtered_len_17_simcse_retrieval_top32__source_tranch_19__target_tranch_27__from_120text100K<n<1M0 likes22 downloads4y agoHugging Face30phnyxlab /klue-nli-simcse KLUENLI for SimCSE Dataset For a better dataset description, please visit: LINK This dataset was prepared by converting KLUENLI dataset to use it for contrastive training (SimCSE). The code used to prepare the data is given below: import pandas as pd from datasets import load_dataset, concatenate_datasets, Dataset from torch.utils.data import random_split class PrepTriplets: @staticmethod def make_dataset(): train_dataset = load_dataset("klue", "nli"… See the full description on the dataset page: https://huggingface.co/datasets/phnyxlab/klue-nli-simcse.text1K<n<10K1 likes22 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.