Team Ai
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01flax-sentence-embeddings /stackexchange_title_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M9 likes1.2k downloads4y agoHugging Face02flax-sentence-embeddings /stackexchange_titlebody_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M11 likes1.2k downloads4y agoHugging Face03flax-sentence-embeddings /stackexchange_title_body_jsonljsonl.gz format from https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_xml Each line contains a dict in the format: {"text": ["title", "body"], "tags": ["tag1", "tag2"]} The following parameters have been used for filtering: min_title_len = 20 min_body_len = 20 max_body_len = 4096 min_score = 0 If a stackexchange contained less than 10k questions (after filtering), it is written to the small_stackexchanges.jsonl.gz file. This is a dump of the files from… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_title_body_jsonl.text1M<n<10M2 likes1k downloads5y agoHugging Face04flax-sentence-embeddings /stackexchange_titlebody_best_and_down_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering100K<n<1M13 likes783 downloads4y agoHugging Face05flax-sentence-embeddings /stackexchange_math_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M20 likes531 downloads4y agoHugging Face06biomedical-translator /pubmed2024_sentence_embeddingstext100M<n<1B0 likes400 downloads2y agoHugging Face07embedding-data /sentence-compression Dataset Card for "sentence-compression" Dataset Summary Dataset with pairs of equivalent sentences. The dataset is provided "AS IS" without any warranty, express or implied. Google disclaims all liability for any damages, direct or indirect, resulting from using the dataset. Disclaimer: The team releasing sentence-compression did not upload the dataset to the Hub and did not write a dataset card. These steps were done by the Hugging Face team. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/embedding-data/sentence-compression.textsentence-similarity100K<n<1M22 likes231 downloads4y agoHugging Face08rjac /all-the-news-2-1-Component-one-sentence-embeddingSimilar dataset to rjac/all-the-news-2-1-Component-one with Embedding generated by Sentence Transformer - model : "all-MiniLM-L6-v2" per small paragraph of an Article. tabular1M<n<10M0 likes108 downloads4y agoHugging Face09hejazizo /mimic-iii-Qwen3-Embedding-0.6B-sentence_transformertext10K<n<100K0 likes43 downloads1y agoHugging Face10flax-sentence-embeddings /Gender_Bias_Evaluation_SetThis dataset has been created as part of the Flax/JAX community week for testing the flax-sentence-embeddings Sentence Similarity models for Gender Bias but can be used for other use-cases as well related to evaluating Gender Bias. The Following Dataset has been created for Evaluating Gender Bias for different models, based on various stereotypical occupations. The Structure of the dataset is of the following type: Base Sentence Occupation Steretypical_Gender Male Sentence Female… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/Gender_Bias_Evaluation_Set.text1K<n<10K4 likes42 downloads2mo agoHugging Face11jahb57 /gpt2_sentence_embeddingstextn<1K0 likes24 downloads3y agoHugging Face12vlhandfo /nordic-sentence-embedding-hard-negatives-cleaned-v2 Dataset Card for nordic-sentence-embedding-hard-negatives-cleaned Dataset Summary This dataset is a cleaned and filtered triplet dataset for training sentence embedding models with hard negatives. Repository: vlhandfo/nordic-sentence-embedding-hard-negatives-cleaned-v2 Languages: Danish, Norwegian, Swedish Final schema: anchor, positive, negative, language, task, id, source Splits: train: 334,801 validation: 8,811 test: 8,811 Total rows: 352,423… See the full description on the dataset page: https://huggingface.co/datasets/vlhandfo/nordic-sentence-embedding-hard-negatives-cleaned-v2.tabulartext-classification100K<n<1M1 likes24 downloads3mo agoHugging Face13beanjar /sp500-business-description-sentence-bert-embeddingsEmbeddings derived from business descriptions of S&P500 companies using sentence-BERT, SentenceTransformer('all-MiniLM-L6-v2') to be exact. For more info on evaluation of sentence transformers (specifcailly the huge GPT-3 versus smaller models see: https://twitter.com/Nils_Reimers/status/1487014195568775173) tabularn<1K0 likes23 downloads4y agoHugging Face14AtefeHassani /mimic-iii-Qwen3-Embedding-0.6B-sentence_transformertext10K<n<100K0 likes23 downloads9mo agoHugging Face15flax-sentence-embeddings /paws-jsonl Introduction This dataset is a jsonl format for PAWS dataset from: https://github.com/google-research-datasets/paws. It only contains the PAWS-Wiki Labeled (Final) and PAWS-Wiki Labeled (Swap-only) training sections of the original PAWS dataset. Duplicates data are removed. Each line contains a dict in the following format: {"guid": <id>, "texts": [anchor, positive]} or {"guid": <id>, "texts": [anchor, positive, negative]} positives_negatives.jsonl.gz: 24,723… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/paws-jsonl.text10K<n<100K1 likes22 downloads5y agoHugging Face16buseskorkmaz /wants_to_hired_gendered_sentence_embeddingstext10K<n<100K0 likes22 downloads2y agoHugging Face17Tm24sense /sentence-embeddingstext100K<n<1M1 likes21 downloads2y agoHugging Face18biomedical-translator /pubmed_sentence_embeddingstext10K<n<100K0 likes19 downloads2y agoHugging Face19vlhandfo /nordic-sentence-embedding-hard-negatives-cleaned Dataset Card for nordic-sentence-embedding-hard-negatives-cleaned Dataset Summary This dataset is a cleaned and filtered triplet dataset for training sentence embedding models with hard negatives. Repository: vlhandfo/nordic-sentence-embedding-hard-negatives-cleaned Languages: Danish, Norwegian, Swedish Final schema: anchor, positive, negative, language, task, id, source Splits: train: 281,938 validation: 35,242 test: 35,243 Total rows: 352,423 Source Data… See the full description on the dataset page: https://huggingface.co/datasets/vlhandfo/nordic-sentence-embedding-hard-negatives-cleaned.text100K<n<1M0 likes14 downloads7mo agoHugging Face20atlasia /Morocco-Darija-Sentence-Embedding-Benchmarkgated Moroccan Darija Sentence Embedding Benchmark This dataset is human-annotated benchmark for evaluating sentence embeddings models in Moroccan Darija (الدارجة المغربية) also known as ary. It was currated by Abdeljalil EL Majjodi, Abdelaziz Bounhar and Amine Hani. The dataset consists of sentence pairs with similarity scores assigned by the human annotators listed above, specifically designed to assess the performance of sentence embedding models on Moroccan Darija text. In particular… See the full description on the dataset page: https://huggingface.co/datasets/atlasia/Morocco-Darija-Sentence-Embedding-Benchmark.textn<1K0 likes12 downloads2y agoHugging Face21jahb57 /bert-base-uncased_sentence_embeddingstextn<1K0 likes9 downloads3y agoHugging Face22Philseok /sentence-embeddingtext1K<n<10K0 likes8 downloads7mo agoHugging Face23the-french-artist /wikipedia_20220301.simple_sentence_embeddingstabular100K<n<1M0 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.