Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01maloyan /wikipedia-22-12-en-embeddings-all-MiniLM-L6-v2 Dataset Card for "wikipedia-22-12-en-embeddings-all-MiniLM-L6-v2" More Information needed tabular10M<n<100M4 likes976 downloads3y agoHugging Face02sentence-transformers /msmarco-msmarco-MiniLM-L6-v3 MS MARCO with hard negatives from msmarco-MiniLM-L6-v3 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models: msmarco-bm25… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-MiniLM-L6-v3.tabularfeature-extraction10M<n<100M3 likes646 downloads2y agoHugging Face03lsb /enwiki20230101-pageid-minilml6v2embeddings Dataset Card for "enwiki20230101-pageid-minilml6v2embeddings" More Information needed text10M<n<100M0 likes562 downloads4y agoHugging Face04ahsanayub /malicious-prompts-minilm-embeddingstabular100K<n<1M0 likes551 downloads2y agoHugging Face05lsb /enwiki20230101-pageid-minilml6v2embeddingsjson Dataset Card for "enwiki20230101-pageid-minilml6v2embeddingsjson" More Information needed text10M<n<100M0 likes431 downloads4y agoHugging Face06lsb /openwebtext-all-minilm-l6-v2-embedding Dataset Card for "openwebtext-all-minilm-l6-v2-embedding" More Information needed text1M<n<10M0 likes334 downloads4y agoHugging Face07bdanko /wixqa-gpt-oss-120b-all-MiniLM-L6-v2-pgvector-evalstext10K<n<100K0 likes302 downloads7mo agoHugging Face08lsb /enwiki20230101-minilml6v2-avgembeddings Dataset Card for "enwiki20230101-minilml6v2-avgembeddings" More Information needed text1M<n<10M0 likes292 downloads4y agoHugging Face09deadbits /vigil-jailbreak-all-MiniLM-L6-v2 Vigil: LLM Jailbreak all-MiniLM-L6-v2 Repo: github.com/deadbits/vigil-llm Vigil is a Python framework and REST API for assessing Large Language Model (LLM) prompts against a set of scanners to detect prompt injections, jailbreaks, and other potentially risky inputs. This repository contains all-MiniLM-L6-v2 embeddings for all "jailbreak" prompts used by Vigil. You can use the parquet2vdb.py utility to load the embeddings in the Vigil chromadb instance, or use them in your own… See the full description on the dataset page: https://huggingface.co/datasets/deadbits/vigil-jailbreak-all-MiniLM-L6-v2.textn<1K2 likes157 downloads3y agoHugging Face10lsb /enwiki20230101-bysize-minilml6v2-avgembeddings Dataset Card for "enwiki20230101-bysize-minilml6v2-avgembeddings" More Information needed text1M<n<10M0 likes92 downloads4y agoHugging Face11Yalexk /newsroom-embeddings-minilmtext1M<n<10M0 likes91 downloads6mo agoHugging Face12lsb /enwiki20250301_paraphrase_multilingual_minilm_l12_v2text10M<n<100M0 likes66 downloads1y agoHugging Face13deadbits /vigil-instruction-bypass-all-MiniLM-L6-v2 Vigil: LLM Instruction Bypass all-MiniLM-L6-v2 Repo: github.com/deadbits/vigil-llm Vigil is a Python framework and REST API for assessing Large Language Model (LLM) prompts against a set of scanners to detect prompt injections, jailbreaks, and other potentially risky inputs. This repository contains all-MiniLM-L6-v2 embeddings for all Instruction Bypass style prompts ("Ignore instructions ...") used by Vigil. You can use the parquet2vdb.py utility to load the embeddings in the… See the full description on the dataset page: https://huggingface.co/datasets/deadbits/vigil-instruction-bypass-all-MiniLM-L6-v2.text1K<n<10K0 likes58 downloads3y agoHugging Face14tommymarto /STEM-wikipedia-22-12-en-embeddings-all-MiniLM-L6-v2tabular1M<n<10M1 likes40 downloads2y agoHugging Face15cat-claws /hotpotqa_clustered_minibatchkmeans_paraphrase-MiniLM-L3-v2_20text10K<n<100K0 likes39 downloads1y agoHugging Face16lsb /fineweb-multilingual-minilm-l12-v2-shard-27468-20000tabular1M<n<10M1 likes38 downloads1y agoHugging Face17cat-claws /hotpotqa_clustered_agglomerative_all-MiniLM-L6-v2_2text10K<n<100K0 likes38 downloads1y agoHugging Face18cat-claws /hotpotqa_clustered_spectral_multi-qa-MiniLM-L6-cos-v1_50text10K<n<100K0 likes38 downloads1y agoHugging Face19cat-claws /hotpotqa_clustered_spectral_all-MiniLM-L6-v2_5text10K<n<100K0 likes37 downloads1y agoHugging Face20cat-claws /hotpotqa_clustered_spectral_paraphrase-MiniLM-L3-v2_50text10K<n<100K0 likes37 downloads1y agoHugging Face21lsb /simplewiki2023-all-minilm-l6-v2-embedding Dataset Card for "simplewiki2023-all-minilm-l6-v2-embedding" More Information needed text100K<n<1M0 likes36 downloads4y agoHugging Face22ArkhAngelLifeJiggy /vigil-instruction-bypass-all-MiniLM-L6-v2 Vigil: LLM Instruction Bypass all-MiniLM-L6-v2 Repo: github.com/deadbits/vigil-llm Vigil is a Python framework and REST API for assessing Large Language Model (LLM) prompts against a set of scanners to detect prompt injections, jailbreaks, and other potentially risky inputs. This repository contains all-MiniLM-L6-v2 embeddings for all Instruction Bypass style prompts ("Ignore instructions ...") used by Vigil. You can use the parquet2vdb.py utility to load the embeddings in the… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/vigil-instruction-bypass-all-MiniLM-L6-v2.text1K<n<10K0 likes34 downloads10d agoHugging Face23cat-claws /hotpotqa_clustered_dbscan_paraphrase-MiniLM-L3-v2_autotext10K<n<100K0 likes33 downloads1y agoHugging Face24cat-claws /hotpotqa_clustered_dbscan_multi-qa-MiniLM-L6-cos-v1_autotext10K<n<100K0 likes32 downloads1y agoHugging Face25karmiq /wikipedia-embeddings-cs-minilmThis dataset contains the Czech subset of the wikimedia/wikipedia dataset. Each page is divided into paragraphs, stored as a list in the chunks column. For every paragraph, embeddings are created using the sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 model. Usage Load the dataset: from datasets import load_dataset ds = load_dataset("karmiq/wikipedia-embeddings-cs-e5-base", split="train") ds[1] { 'id': '1', 'url': 'https://cs.wikipedia.org/wiki/Astronomie'… See the full description on the dataset page: https://huggingface.co/datasets/karmiq/wikipedia-embeddings-cs-minilm.texttext-generation100K<n<1M0 likes31 downloads3y agoHugging Face26cat-claws /hotpotqa_clustered_dbscan_all-MiniLM-L6-v2_autotext10K<n<100K0 likes31 downloads1y agoHugging Face27cat-claws /hotpotqa_clustered_spectral_all-MiniLM-L6-v2_2text10K<n<100K0 likes31 downloads1y agoHugging Face28cat-claws /hotpotqa_clustered_spectral_paraphrase-MiniLM-L3-v2_2text10K<n<100K0 likes31 downloads1y agoHugging Face29cat-claws /hotpotqa_clustered_minibatchkmeans_multi-qa-MiniLM-L6-cos-v1_2text10K<n<100K0 likes30 downloads1y agoHugging Face30cat-claws /hotpotqa_clustered_kmeans_all-MiniLM-L6-v2_20text100K<n<1M0 likes28 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.