datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
embedding-training-data
Training Data for Text Embedding Models
[!NOTE]
This repository contains raw datasets, all of which have also been formatted for easy training in the Embedding Model Datasets collection. We recommend looking there first.
This repository contains training files to train text embedding models, e.g. using sentence-transformers.
Data Format
All files are in a jsonl.gz format: Each line contains a JSON-object that represent one training example.
The JSON objects can… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/embedding-training-data.stackexchange_xmlThis is a dump of the files from
https://archive.org/details/stackexchange
downloaded via torrent on 2021-07-01.
Publication date 2021-06-07 Usage Attribution-ShareAlike 4.0 International Creative Commons License by sa Topics Stack Exchange Data Dump Contributor Stack Exchange Community
Please see the license information at:
https://archive.org/details/stackexchange
The dataset has been split into following for cleaner formatting.… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_xml.stackexchange_title_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.stackexchange_titlebody_best_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.stackexchange_title_body_jsonljsonl.gz format from https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_xml
Each line contains a dict in the format: {"text": ["title", "body"], "tags": ["tag1", "tag2"]}
The following parameters have been used for filtering: min_title_len = 20 min_body_len = 20 max_body_len = 4096 min_score = 0
If a stackexchange contained less than 10k questions (after filtering), it is written to the small_stackexchanges.jsonl.gz file.
This is a dump of the files from… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/stackexchange_title_body_jsonl.stackexchange_titlebody_best_and_down_voted_answer_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.stackexchange_math_jsonlThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.pubmed2024_sentence_embeddingssentence-compression
Dataset Card for "sentence-compression"
Dataset Summary
Dataset with pairs of equivalent sentences.
The dataset is provided "AS IS" without any warranty, express or implied.
Google disclaims all liability for any damages, direct or indirect, resulting from using the dataset.
Disclaimer: The team releasing sentence-compression did not upload the dataset to the Hub and did not write a dataset card. These steps were done by the Hugging Face team.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/embedding-data/sentence-compression.all-the-news-2-1-Component-one-sentence-embeddingSimilar dataset to rjac/all-the-news-2-1-Component-one with Embedding generated by Sentence Transformer - model : "all-MiniLM-L6-v2" per small paragraph of an Article.
scottish-councils-sentence-embeddings
Scottish Council Embeddings
Site content from all* Scottish council sites scraped and embedded using Sentence Transformers and all-mpnet-base-v2 model.
* Some councils were unable to be scraped effectively, resulting in little or no embeddings:
Aberdeenshire Council
Aberdeen City Council
Angus Council
Glasgow City Council
mimic-iii-Qwen3-Embedding-0.6B-sentence_transformerGender_Bias_Evaluation_SetThis dataset has been created as part of the Flax/JAX community week for testing the flax-sentence-embeddings Sentence Similarity models for Gender Bias but can be used for other use-cases as well related to evaluating Gender Bias.
The Following Dataset has been created for Evaluating Gender Bias for different models, based on various stereotypical occupations.
The Structure of the dataset is of the following type:
Base Sentence
Occupation
Steretypical_Gender
Male Sentence
Female… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/Gender_Bias_Evaluation_Set.gpt2_sentence_embeddingsnordic-sentence-embedding-hard-negatives-cleaned-v2
Dataset Card for nordic-sentence-embedding-hard-negatives-cleaned
Dataset Summary
This dataset is a cleaned and filtered triplet dataset for training sentence embedding models with hard negatives.
Repository: vlhandfo/nordic-sentence-embedding-hard-negatives-cleaned-v2
Languages: Danish, Norwegian, Swedish
Final schema: anchor, positive, negative, language, task, id, source
Splits:
train: 334,801
validation: 8,811
test: 8,811
Total rows: 352,423… See the full description on the dataset page: https://huggingface.co/datasets/vlhandfo/nordic-sentence-embedding-hard-negatives-cleaned-v2.sp500-business-description-sentence-bert-embeddingsEmbeddings derived from business descriptions of S&P500 companies using sentence-BERT, SentenceTransformer('all-MiniLM-L6-v2') to be exact. For more info on evaluation of sentence transformers (specifcailly the huge GPT-3 versus smaller models see: https://twitter.com/Nils_Reimers/status/1487014195568775173)
mimic-iii-Qwen3-Embedding-0.6B-sentence_transformerpaws-jsonl
Introduction
This dataset is a jsonl format for PAWS dataset from: https://github.com/google-research-datasets/paws. It only contains the PAWS-Wiki Labeled (Final) and
PAWS-Wiki Labeled (Swap-only) training sections of the original PAWS dataset. Duplicates data are removed.
Each line contains a dict in the following format:
{"guid": <id>, "texts": [anchor, positive]} or
{"guid": <id>, "texts": [anchor, positive, negative]}
positives_negatives.jsonl.gz: 24,723… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/paws-jsonl.wants_to_hired_gendered_sentence_embeddingssentence-embeddingspubmed_sentence_embeddingsnordic-sentence-embedding-hard-negatives-cleaned
Dataset Card for nordic-sentence-embedding-hard-negatives-cleaned
Dataset Summary
This dataset is a cleaned and filtered triplet dataset for training sentence embedding models with hard negatives.
Repository: vlhandfo/nordic-sentence-embedding-hard-negatives-cleaned
Languages: Danish, Norwegian, Swedish
Final schema: anchor, positive, negative, language, task, id, source
Splits:
train: 281,938
validation: 35,242
test: 35,243
Total rows: 352,423
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/vlhandfo/nordic-sentence-embedding-hard-negatives-cleaned.Morocco-Darija-Sentence-Embedding-Benchmark
Moroccan Darija Sentence Embedding Benchmark
This dataset is human-annotated benchmark for evaluating sentence embeddings models in Moroccan Darija (الدارجة المغربية) also known as ary.
It was currated by Abdeljalil EL Majjodi, Abdelaziz Bounhar and Amine Hani.
The dataset consists of sentence pairs with similarity scores assigned by the human annotators listed above, specifically designed to assess the performance of sentence embedding models on Moroccan Darija text. In particular… See the full description on the dataset page: https://huggingface.co/datasets/atlasia/Morocco-Darija-Sentence-Embedding-Benchmark.Sinhala_Facebook_posts_sentence_embeddingsbert-base-uncased_sentence_embeddingssentence-embeddingsentence-embedding-datasetwikipedia_20220301.simple_sentence_embeddings
