Team Ai
28 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01sentence-transformers /embedding-training-data Training Data for Text Embedding Models [!NOTE] This repository contains raw datasets, all of which have also been formatted for easy training in the Embedding Model Datasets collection. We recommend looking there first. This repository contains training files to train text embedding models, e.g. using sentence-transformers. Data Format All files are in a jsonl.gz format: Each line contains a JSON-object that represent one training example. The JSON objects can… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/embedding-training-data.feature-extraction144 likes2.4k downloads2mo agoHugging Face02flax-sentence-embeddings /stackexchange_xmlThis is a dump of the files from https://archive.org/details/stackexchange downloaded via torrent on 2021-07-01. Publication date 2021-06-07 Usage Attribution-ShareAlike 4.0 International Creative Commons License by sa Topics Stack Exchange Data Dump Contributor Stack Exchange Community Please see the license information at: https://archive.org/details/stackexchange The dataset has been split into following for cleaner formatting.… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_xml.1 likes1.4k downloads5y agoHugging Face03flax-sentence-embeddings /stackexchange_title_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M9 likes1.2k downloads4y agoHugging Face04flax-sentence-embeddings /stackexchange_titlebody_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M11 likes1.2k downloads4y agoHugging Face05flax-sentence-embeddings /stackexchange_title_body_jsonljsonl.gz format from https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_xml Each line contains a dict in the format: {"text": ["title", "body"], "tags": ["tag1", "tag2"]} The following parameters have been used for filtering: min_title_len = 20 min_body_len = 20 max_body_len = 4096 min_score = 0 If a stackexchange contained less than 10k questions (after filtering), it is written to the small_stackexchanges.jsonl.gz file. This is a dump of the files from… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_title_body_jsonl.text1M<n<10M2 likes1k downloads5y agoHugging Face06flax-sentence-embeddings /stackexchange_titlebody_best_and_down_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering100K<n<1M13 likes783 downloads4y agoHugging Face07flax-sentence-embeddings /stackexchange_math_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.textquestion-answering1M<n<10M20 likes531 downloads4y agoHugging Face08biomedical-translator /pubmed2024_sentence_embeddingstext100M<n<1B0 likes400 downloads2y agoHugging Face09embedding-data /sentence-compression Dataset Card for "sentence-compression" Dataset Summary Dataset with pairs of equivalent sentences. The dataset is provided "AS IS" without any warranty, express or implied. Google disclaims all liability for any damages, direct or indirect, resulting from using the dataset. Disclaimer: The team releasing sentence-compression did not upload the dataset to the Hub and did not write a dataset card. These steps were done by the Hugging Face team. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/embedding-data/sentence-compression.textsentence-similarity100K<n<1M22 likes231 downloads4y agoHugging Face10rjac /all-the-news-2-1-Component-one-sentence-embeddingSimilar dataset to rjac/all-the-news-2-1-Component-one with Embedding generated by Sentence Transformer - model : "all-MiniLM-L6-v2" per small paragraph of an Article. tabular1M<n<10M0 likes108 downloads4y agoHugging Face11weirdjet /scottish-councils-sentence-embeddings Scottish Council Embeddings Site content from all* Scottish council sites scraped and embedded using Sentence Transformers and all-mpnet-base-v2 model. * Some councils were unable to be scraped effectively, resulting in little or no embeddings: Aberdeenshire Council Aberdeen City Council Angus Council Glasgow City Council sentence-similarity0 likes53 downloads3y agoHugging Face12hejazizo /mimic-iii-Qwen3-Embedding-0.6B-sentence_transformertext10K<n<100K0 likes43 downloads1y agoHugging Face13flax-sentence-embeddings /Gender_Bias_Evaluation_SetThis dataset has been created as part of the Flax/JAX community week for testing the flax-sentence-embeddings Sentence Similarity models for Gender Bias but can be used for other use-cases as well related to evaluating Gender Bias. The Following Dataset has been created for Evaluating Gender Bias for different models, based on various stereotypical occupations. The Structure of the dataset is of the following type: Base Sentence Occupation Steretypical_Gender Male Sentence Female… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/Gender_Bias_Evaluation_Set.text1K<n<10K4 likes42 downloads2mo agoHugging Face14jahb57 /gpt2_sentence_embeddingstextn<1K0 likes24 downloads3y agoHugging Face15vlhandfo /nordic-sentence-embedding-hard-negatives-cleaned-v2 Dataset Card for nordic-sentence-embedding-hard-negatives-cleaned Dataset Summary This dataset is a cleaned and filtered triplet dataset for training sentence embedding models with hard negatives. Repository: vlhandfo/nordic-sentence-embedding-hard-negatives-cleaned-v2 Languages: Danish, Norwegian, Swedish Final schema: anchor, positive, negative, language, task, id, source Splits: train: 334,801 validation: 8,811 test: 8,811 Total rows: 352,423… See the full description on the dataset page: https://huggingface.co/datasets/vlhandfo/nordic-sentence-embedding-hard-negatives-cleaned-v2.tabulartext-classification100K<n<1M1 likes24 downloads3mo agoHugging Face16beanjar /sp500-business-description-sentence-bert-embeddingsEmbeddings derived from business descriptions of S&P500 companies using sentence-BERT, SentenceTransformer('all-MiniLM-L6-v2') to be exact. For more info on evaluation of sentence transformers (specifcailly the huge GPT-3 versus smaller models see: https://twitter.com/Nils_Reimers/status/1487014195568775173) tabularn<1K0 likes23 downloads4y agoHugging Face17AtefeHassani /mimic-iii-Qwen3-Embedding-0.6B-sentence_transformertext10K<n<100K0 likes23 downloads9mo agoHugging Face18flax-sentence-embeddings /paws-jsonl Introduction This dataset is a jsonl format for PAWS dataset from: https://github.com/google-research-datasets/paws. It only contains the PAWS-Wiki Labeled (Final) and PAWS-Wiki Labeled (Swap-only) training sections of the original PAWS dataset. Duplicates data are removed. Each line contains a dict in the following format: {"guid": <id>, "texts": [anchor, positive]} or {"guid": <id>, "texts": [anchor, positive, negative]} positives_negatives.jsonl.gz: 24,723… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/paws-jsonl.text10K<n<100K1 likes22 downloads5y agoHugging Face19buseskorkmaz /wants_to_hired_gendered_sentence_embeddingstext10K<n<100K0 likes22 downloads2y agoHugging Face20Tm24sense /sentence-embeddingstext100K<n<1M1 likes21 downloads2y agoHugging Face21biomedical-translator /pubmed_sentence_embeddingstext10K<n<100K0 likes19 downloads2y agoHugging Face22vlhandfo /nordic-sentence-embedding-hard-negatives-cleaned Dataset Card for nordic-sentence-embedding-hard-negatives-cleaned Dataset Summary This dataset is a cleaned and filtered triplet dataset for training sentence embedding models with hard negatives. Repository: vlhandfo/nordic-sentence-embedding-hard-negatives-cleaned Languages: Danish, Norwegian, Swedish Final schema: anchor, positive, negative, language, task, id, source Splits: train: 281,938 validation: 35,242 test: 35,243 Total rows: 352,423 Source Data… See the full description on the dataset page: https://huggingface.co/datasets/vlhandfo/nordic-sentence-embedding-hard-negatives-cleaned.text100K<n<1M0 likes14 downloads7mo agoHugging Face23atlasia /Morocco-Darija-Sentence-Embedding-Benchmarkgated Moroccan Darija Sentence Embedding Benchmark This dataset is human-annotated benchmark for evaluating sentence embeddings models in Moroccan Darija (الدارجة المغربية) also known as ary. It was currated by Abdeljalil EL Majjodi, Abdelaziz Bounhar and Amine Hani. The dataset consists of sentence pairs with similarity scores assigned by the human annotators listed above, specifically designed to assess the performance of sentence embedding models on Moroccan Darija text. In particular… See the full description on the dataset page: https://huggingface.co/datasets/atlasia/Morocco-Darija-Sentence-Embedding-Benchmark.textn<1K0 likes12 downloads2y agoHugging Face24grawcse /Sinhala_Facebook_posts_sentence_embeddings0 likes11 downloads4y agoHugging Face25jahb57 /bert-base-uncased_sentence_embeddingstextn<1K0 likes9 downloads3y agoHugging Face26Philseok /sentence-embeddingtext1K<n<10K0 likes8 downloads7mo agoHugging Face27dbena /sentence-embedding-datasetsentence-similarity0 likes5 downloads2y agoHugging Face28the-french-artist /wikipedia_20220301.simple_sentence_embeddingstabular100K<n<1M0 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.