datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
stackexchange_title_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.stackexchange_titlebody_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.stackexchange_title_body_jsonljsonl.gz format from https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_xml
Each line contains a dict in the format: {"text": ["title", "body"], "tags": ["tag1", "tag2"]}
The following parameters have been used for filtering: min_title_len = 20 min_body_len = 20 max_body_len = 4096 min_score = 0
If a stackexchange contained less than 10k questions (after filtering), it is written to the small_stackexchanges.jsonl.gz file.
This is a dump of the files from… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_title_body_jsonl.stackexchange_titlebody_best_and_down_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.stackexchange_math_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.pubmed2024_sentence_embeddingssentence-compression
Dataset Card for "sentence-compression"
Dataset Summary
Dataset with pairs of equivalent sentences.
The dataset is provided "AS IS" without any warranty, express or implied.
Google disclaims all liability for any damages, direct or indirect, resulting from using the dataset.
Disclaimer: The team releasing sentence-compression did not upload the dataset to the Hub and did not write a dataset card. These steps were done by the Hugging Face team.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/embedding-data/sentence-compression.all-the-news-2-1-Component-one-sentence-embeddingSimilar dataset to rjac/all-the-news-2-1-Component-one with Embedding generated by Sentence Transformer - model : "all-MiniLM-L6-v2" per small paragraph of an Article.
mimic-iii-Qwen3-Embedding-0.6B-sentence_transformerGender_Bias_Evaluation_SetThis dataset has been created as part of the Flax/JAX community week for testing the flax-sentence-embeddings Sentence Similarity models for Gender Bias but can be used for other use-cases as well related to evaluating Gender Bias.
The Following Dataset has been created for Evaluating Gender Bias for different models, based on various stereotypical occupations.
The Structure of the dataset is of the following type:
Base Sentence
Occupation
Steretypical_Gender
Male Sentence
Female… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/Gender_Bias_Evaluation_Set.gpt2_sentence_embeddingsnordic-sentence-embedding-hard-negatives-cleaned-v2
Dataset Card for nordic-sentence-embedding-hard-negatives-cleaned
Dataset Summary
This dataset is a cleaned and filtered triplet dataset for training sentence embedding models with hard negatives.
Repository: vlhandfo/nordic-sentence-embedding-hard-negatives-cleaned-v2
Languages: Danish, Norwegian, Swedish
Final schema: anchor, positive, negative, language, task, id, source
Splits:
train: 334,801
validation: 8,811
test: 8,811
Total rows: 352,423… See the full description on the dataset page: https://huggingface.co/datasets/vlhandfo/nordic-sentence-embedding-hard-negatives-cleaned-v2.sp500-business-description-sentence-bert-embeddingsEmbeddings derived from business descriptions of S&P500 companies using sentence-BERT, SentenceTransformer('all-MiniLM-L6-v2') to be exact. For more info on evaluation of sentence transformers (specifcailly the huge GPT-3 versus smaller models see: https://twitter.com/Nils_Reimers/status/1487014195568775173)
mimic-iii-Qwen3-Embedding-0.6B-sentence_transformerpaws-jsonl
Introduction
This dataset is a jsonl format for PAWS dataset from: https://github.com/google-research-datasets/paws. It only contains the PAWS-Wiki Labeled (Final) and
PAWS-Wiki Labeled (Swap-only) training sections of the original PAWS dataset. Duplicates data are removed.
Each line contains a dict in the following format:
{"guid": <id>, "texts": [anchor, positive]} or
{"guid": <id>, "texts": [anchor, positive, negative]}
positives_negatives.jsonl.gz: 24,723… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/paws-jsonl.wants_to_hired_gendered_sentence_embeddingssentence-embeddingspubmed_sentence_embeddingsnordic-sentence-embedding-hard-negatives-cleaned
Dataset Card for nordic-sentence-embedding-hard-negatives-cleaned
Dataset Summary
This dataset is a cleaned and filtered triplet dataset for training sentence embedding models with hard negatives.
Repository: vlhandfo/nordic-sentence-embedding-hard-negatives-cleaned
Languages: Danish, Norwegian, Swedish
Final schema: anchor, positive, negative, language, task, id, source
Splits:
train: 281,938
validation: 35,242
test: 35,243
Total rows: 352,423
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/vlhandfo/nordic-sentence-embedding-hard-negatives-cleaned.Morocco-Darija-Sentence-Embedding-Benchmark
Moroccan Darija Sentence Embedding Benchmark
This dataset is human-annotated benchmark for evaluating sentence embeddings models in Moroccan Darija (الدارجة المغربية) also known as ary.
It was currated by Abdeljalil EL Majjodi, Abdelaziz Bounhar and Amine Hani.
The dataset consists of sentence pairs with similarity scores assigned by the human annotators listed above, specifically designed to assess the performance of sentence embedding models on Moroccan Darija text. In particular… See the full description on the dataset page: https://huggingface.co/datasets/atlasia/Morocco-Darija-Sentence-Embedding-Benchmark.bert-base-uncased_sentence_embeddingssentence-embeddingwikipedia_20220301.simple_sentence_embeddings
