Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01maloyan /wikipedia-22-12-en-embeddings-all-MiniLM-L6-v2 Dataset Card for "wikipedia-22-12-en-embeddings-all-MiniLM-L6-v2" More Information needed tabular10M<n<100M4 likes1k downloads3y agoHugging Face02sentence-transformers /msmarco-msmarco-MiniLM-L6-v3 MS MARCO with hard negatives from msmarco-MiniLM-L6-v3 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models. Related Datasets These are the datasets generated using the 13 different models: msmarco-bm25… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-MiniLM-L6-v3.tabularfeature-extraction10M<n<100M3 likes667 downloads2y agoHugging Face03ahsanayub /malicious-prompts-minilm-embeddingstabular100K<n<1M0 likes550 downloads2y agoHugging Face04lsb /enwiki20230101-pageid-minilml6v2embeddings Dataset Card for "enwiki20230101-pageid-minilml6v2embeddings" More Information needed text10M<n<100M0 likes517 downloads4y agoHugging Face05lsb /enwiki20230101-pageid-minilml6v2embeddingsjson Dataset Card for "enwiki20230101-pageid-minilml6v2embeddingsjson" More Information needed text10M<n<100M0 likes449 downloads4y agoHugging Face06lsb /openwebtext-all-minilm-l6-v2-embedding Dataset Card for "openwebtext-all-minilm-l6-v2-embedding" More Information needed text1M<n<10M0 likes311 downloads4y agoHugging Face07lsb /enwiki20230101-minilml6v2-avgembeddings Dataset Card for "enwiki20230101-minilml6v2-avgembeddings" More Information needed text1M<n<10M0 likes305 downloads4y agoHugging Face08bdanko /wixqa-gpt-oss-120b-all-MiniLM-L6-v2-pgvector-evalstext10K<n<100K0 likes285 downloads7mo agoHugging Face09deadbits /vigil-jailbreak-all-MiniLM-L6-v2 Vigil: LLM Jailbreak all-MiniLM-L6-v2 Repo: github.com/deadbits/vigil-llm Vigil is a Python framework and REST API for assessing Large Language Model (LLM) prompts against a set of scanners to detect prompt injections, jailbreaks, and other potentially risky inputs. This repository contains all-MiniLM-L6-v2 embeddings for all "jailbreak" prompts used by Vigil. You can use the parquet2vdb.py utility to load the embeddings in the Vigil chromadb instance, or use them in your own… See the full description on the dataset page: https://huggingface.co/datasets/deadbits/vigil-jailbreak-all-MiniLM-L6-v2.textn<1K2 likes154 downloads3y agoHugging Face10lsb /enwiki20230101-bysize-minilml6v2-avgembeddings Dataset Card for "enwiki20230101-bysize-minilml6v2-avgembeddings" More Information needed text1M<n<10M0 likes95 downloads4y agoHugging Face11Yalexk /newsroom-embeddings-minilmtext1M<n<10M0 likes91 downloads6mo agoHugging Face12sentence-transformers /msmarco-scores-ms-marco-MiniLM-L6-v2 MS MARCO query-passage scores using cross-encoder/ms-marco-MiniLM-L6-v2 MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine. This dataset contains 160 million CrossEncoder scores on the MS MARCO dataset, using the cross-encoder/ms-marco-MiniLM-L6-v2 model. The scores are unprocessed logits, i.e. they don't range between 0...1, and they can be used for finetuning search models using distillation. See… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-scores-ms-marco-MiniLM-L6-v2.tabularfeature-extraction100M<n<1B3 likes88 downloads1y agoHugging Face13deadbits /vigil-instruction-bypass-all-MiniLM-L6-v2 Vigil: LLM Instruction Bypass all-MiniLM-L6-v2 Repo: github.com/deadbits/vigil-llm Vigil is a Python framework and REST API for assessing Large Language Model (LLM) prompts against a set of scanners to detect prompt injections, jailbreaks, and other potentially risky inputs. This repository contains all-MiniLM-L6-v2 embeddings for all Instruction Bypass style prompts ("Ignore instructions ...") used by Vigil. You can use the parquet2vdb.py utility to load the embeddings in the… See the full description on the dataset page: https://huggingface.co/datasets/deadbits/vigil-instruction-bypass-all-MiniLM-L6-v2.text1K<n<10K0 likes57 downloads3y agoHugging Face14tommymarto /STEM-wikipedia-22-12-en-embeddings-all-MiniLM-L6-v2tabular1M<n<10M1 likes44 downloads2y agoHugging Face15lsb /fineweb-multilingual-minilm-l12-v2-shard-27468-20000tabular1M<n<10M1 likes38 downloads1y agoHugging Face16cat-claws /hotpotqa_clustered_agglomerative_all-MiniLM-L6-v2_2text10K<n<100K0 likes38 downloads1y agoHugging Face17cat-claws /hotpotqa_clustered_minibatchkmeans_paraphrase-MiniLM-L3-v2_20text10K<n<100K0 likes38 downloads1y agoHugging Face18cat-claws /hotpotqa_clustered_spectral_multi-qa-MiniLM-L6-cos-v1_50text10K<n<100K0 likes38 downloads1y agoHugging Face19alperk3003 /frca-adversarialqa-dbert-minilm-raw-results-20260924 AdversarialQA DeBERTa+MiniLM exploratory raw results This repository contains the full adapter-bearing raw archive for the pre-named AdversarialQA exploratory cell reported as certificate-backed in the associated prompt-visible feedback study. It is an auxiliary result, not the paper's primary endpoint. Contents adversarialqa_dbert_minilm_cuda130_raw_results_20260924.zip ? full raw state, including adapters and the execution receipt (1,289,053,860 bytes).… See the full description on the dataset page: https://huggingface.co/datasets/alperk3003/frca-adversarialqa-dbert-minilm-raw-results-20260924.0 likes38 downloads12d agoHugging Face20cat-claws /hotpotqa_clustered_spectral_all-MiniLM-L6-v2_5text10K<n<100K0 likes37 downloads1y agoHugging Face21cat-claws /hotpotqa_clustered_spectral_paraphrase-MiniLM-L3-v2_50text10K<n<100K0 likes37 downloads1y agoHugging Face22lsb /simplewiki2023-all-minilm-l6-v2-embedding Dataset Card for "simplewiki2023-all-minilm-l6-v2-embedding" More Information needed text100K<n<1M0 likes36 downloads4y agoHugging Face23ArkhAngelLifeJiggy /vigil-instruction-bypass-all-MiniLM-L6-v2 Vigil: LLM Instruction Bypass all-MiniLM-L6-v2 Repo: github.com/deadbits/vigil-llm Vigil is a Python framework and REST API for assessing Large Language Model (LLM) prompts against a set of scanners to detect prompt injections, jailbreaks, and other potentially risky inputs. This repository contains all-MiniLM-L6-v2 embeddings for all Instruction Bypass style prompts ("Ignore instructions ...") used by Vigil. You can use the parquet2vdb.py utility to load the embeddings in the… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/vigil-instruction-bypass-all-MiniLM-L6-v2.text1K<n<10K0 likes34 downloads9d agoHugging Face24cat-claws /hotpotqa_clustered_dbscan_multi-qa-MiniLM-L6-cos-v1_autotext10K<n<100K0 likes32 downloads1y agoHugging Face25cat-claws /hotpotqa_clustered_dbscan_all-MiniLM-L6-v2_autotext10K<n<100K0 likes31 downloads1y agoHugging Face26cat-claws /hotpotqa_clustered_dbscan_paraphrase-MiniLM-L3-v2_autotext10K<n<100K0 likes31 downloads1y agoHugging Face27cat-claws /hotpotqa_clustered_spectral_all-MiniLM-L6-v2_2text10K<n<100K0 likes31 downloads1y agoHugging Face28cat-claws /hotpotqa_clustered_spectral_paraphrase-MiniLM-L3-v2_2text10K<n<100K0 likes31 downloads1y agoHugging Face29jamescalam /oscar-en-minilm-2m Oscar EN 2M Embeddings This dataset contains 2M sentences extracted from the English subset of the OSCAR dataset, and encoded into sentence embeddings using the sentence-transformers/all-MiniLM-L6-v2 model. sentence-similarity1M<n<10M1 likes30 downloads4y agoHugging Face30karmiq /wikipedia-embeddings-cs-minilmThis dataset contains the Czech subset of the wikimedia/wikipedia dataset. Each page is divided into paragraphs, stored as a list in the chunks column. For every paragraph, embeddings are created using the sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 model. Usage Load the dataset: from datasets import load_dataset ds = load_dataset("karmiq/wikipedia-embeddings-cs-e5-base", split="train") ds[1] { 'id': '1', 'url': 'https://cs.wikipedia.org/wiki/Astronomie'… See the full description on the dataset page: https://huggingface.co/datasets/karmiq/wikipedia-embeddings-cs-minilm.texttext-generation100K<n<1M0 likes30 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.