datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
seamless-align-enA-esA.speaker-embedding.w2vbert-600mseamless-align-enA-esA.speaker-embedding.xlsr-2bmultilingual-embeddings-pre-training-curated
📚 Collection | 📝 Multilingual Blog | 📝 English Blog
Contrastive Multilingual Pre-Training
2.16B query–document pairs across eight languages, plus cross-lingual pairs
mDenseOn |
mLateOn |
DenseOn |
LateOn |
PyLate |
FastPlaid
🎯 TL;DR: The multilingual contrastive pre-training corpus used to train mDenseOn and mLateOn. It extends our curated English data recipe (embeddings-pre-training-curated) to French, German, Italian, Spanish, Portuguese, Swedish, Norwegian, and Arabic… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/multilingual-embeddings-pre-training-curated.daily-papers-embeddingsseamless-align-enA-jaA.speaker-embedding.w2vbert-600mAthar-Embeddingsseamless-align-enA-viA.speaker-embedding.xlsr-2bpaired-llama-3.2-1b-embeddings-lmsys-chat-1m
Paired Llama 3.2 1B Token Embeddings (LMSYS-Chat-1M)
This dataset contains paired activations corresponding to single token locations extracted from Meta's Llama 3.2 1B Instruct on conversations from LMSYS-Chat-1M.
Embeddings are provided for layers 5 through 14, which capture the most interesting intermediate representations.
This dataset was built to study things like:
Learning different basis for activations at a given layer
Studying if there are cases where position encodes… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/paired-llama-3.2-1b-embeddings-lmsys-chat-1m.seamless-align-enA-zhA.speaker-embedding.w2vbert-600mlatent-sr-embeddings
Latent-SR Embeddings: Precomputed VAE Latents for Medical Image Super-Resolution
Precomputed VAE latent embeddings from the paper:
"Domain-Specific Latent Representations Improve the Fidelity of Diffusion-Based Medical Image Super-Resolution"Sebastian Cajas, Ashaba Judith, Rahul Gorijavolu, Sahil Kapadia, Hillary Clinton Kasimbazi, Leo Kinyera, Emmanuel Paul Kwesiga, Sri Sri Jaithra Varma Manthena, Luis Filipe Nakayama, Ninsiima Doreen, Leo Anthony Celi.arXiv:2604.12152 (2026)… See the full description on the dataset page: https://huggingface.co/datasets/sebasmos/latent-sr-embeddings.seamless-align-enA-frA.speaker-embedding.xlsr-2bseamless-align-enA-zhA.speaker-embedding.hubert-xlembeddings-fine-tuning
Overview
This dataset is composed of high quality data sources with mined hard negatives. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using this dataset or its curated version.
This dataset has originally been created to follow the nv-retrieve setup, that mines the closest negatives to the query in a dataset and filter false negatives if their bi-encoder similarity is higher than a… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning.seamless-align-enA-zhA.speaker-embedding.xlsr-2bCaselaw_Access_Project_embeddingsThis is an embeddings dataset for the Caselaw Access Project, created by a user named Endomorphosis.
Each caselaw entry is hashed with IPFS / multiformats, so retrieval of the document can be made over the IPFS / filecoin network
The ipfs content id "cid" is the primary key that links the dataset to the embeddings, should you want to retrieve from the dataset instead.
The dataset has been had embeddings generated with three models: thenlper/gte-small, Alibaba-NLP/gte-large-en-v1.5, and… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/Caselaw_Access_Project_embeddings.seamless-align-deA-enA.speaker-embedding.xlsr-2bseamless-align-enA-hiA.speaker-embedding.hubert-xlseamless-align-enA-jaA.speaker-embedding.xlsr-2bseamless-align-enA-frA.speaker-embedding.w2vbert-600mCaselaw_Access_Project_embeddingsOriginal Repository:
https://huggingface.co/datasets/justicedao/Caselaw_Access_Project_embeddings/
This is an embeddings dataset for the Caselaw Access Project, created by a user named Endomorphosis.
Each caselaw entry is hashed with IPFS / multiformats, so retrieval of the document can be made over the IPFS / filecoin network
The ipfs content id "cid" is the primary key that links the dataset to the embeddings, should you want to retrieve from the dataset instead.
The dataset has been had… See the full description on the dataset page: https://huggingface.co/datasets/laion/Caselaw_Access_Project_embeddings.seamless-align-enA-jaA.speaker-embedding.hubert-xlembeddingssubset_arxiv_papers_with_embeddingsThis dataset is a curated subset of the original arXiv dataset, each entry enriched with a 256-dimensional embedding vector. The embeddings are generated using OpenAI's "text-embedding-3-small" model. For each data point, the embedding is created by concatenating the text of the title, author(s), and abstract into a single string, which is then processed by the embedding model. This approach captures the semantic essence of each document, facilitating tasks such as similarity search… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/subset_arxiv_papers_with_embeddings.seamless-align-enA-hiA.speaker-embedding.w2vbert-600mopenalex-multilingual-embeddings
OpenAlex Multilingual Embeddings
This dataset contains multilingual text embeddings of all records in OpenAlex with a title or an abstract from the snapshot of 2023-10-20.
The dataset was created for the FORAS project to investigate the efficacy of
different methods of searching in databases of academic publications. All scripts will be available in a GitHub repository.
The project is supported by a grant from the Dutch Research Council (grant no. 406.22.GO.048)… See the full description on the dataset page: https://huggingface.co/datasets/GlobalCampus/openalex-multilingual-embeddings.emilia-yodas-en-speaker-embeddings
Emilia-YODAS English Qwen3-TTS Speaker Embeddings
This dataset contains precomputed speaker embeddings for the English subset of
Emilia-YODAS. Each row maps an Emilia-YODAS sample ID to one speaker embedding
extracted from the corresponding audio.
Dataset Details
Source dataset: amphion/Emilia-Dataset
Source subset: Emilia-YODAS English
Embedding model: Qwen/Qwen3-TTS-12Hz-1.7B-Base
Embedding shape: (2048,)
Embedding dtype: float16
Rows: 4,516,833
Split: train
Additional… See the full description on the dataset page: https://huggingface.co/datasets/duplexio/emilia-yodas-en-speaker-embeddings.seamless-align-enA-viA.speaker-embedding.w2vbert-600mseamless-align-enA-hiA.speaker-embedding.xlsr-2bseamless-align-enA-koA.speaker-embedding.w2vbert-600mKaLM-embedding-pretrain-dataThe finetuning dataset is is available at this link:KaLM-Embedding/KaLM-embedding-finetuning-data.
Citation
If you find these datasets useful, please consider giving a star and citation.
@misc{zhao2025kalmembeddingv2,
title={KaLM-Embedding-V2: Superior Training Techniques and Data Inspire A Versatile Embedding Model},
author={Xinping Zhao and Xinshuo Hu and Zifei Shan and Shouzheng Huang and Yao Zhou and Xin Zhang and Zetian Sun and Zhenyu Liu and Dongfang Li and… See the full description on the dataset page: https://huggingface.co/datasets/HIT-TMG/KaLM-embedding-pretrain-data.
