Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01asahi417 /seamless-align-enA-esA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes12k downloads2y agoHugging Face02asahi417 /seamless-align-enA-esA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes12k downloads2y agoHugging Face03lightonai /multilingual-embeddings-pre-training-curated 📚 Collection | 📝 Multilingual Blog | 📝 English Blog Contrastive Multilingual Pre-Training 2.16B query–document pairs across eight languages, plus cross-lingual pairs mDenseOn | mLateOn | DenseOn | LateOn | PyLate | FastPlaid 🎯 TL;DR: The multilingual contrastive pre-training corpus used to train mDenseOn and mLateOn. It extends our curated English data recipe (embeddings-pre-training-curated) to French, German, Italian, Spanish, Portuguese, Swedish, Norwegian, and Arabic… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/multilingual-embeddings-pre-training-curated.texttext-retrieval1B<n<10B6 likes11k downloads2mo agoHugging Face04hysts-bot-data /daily-papers-embeddingstext10K<n<100K10 likes9.7k downloads13h agoHugging Face05asahi417 /seamless-align-enA-jaA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes9.5k downloads2y agoHugging Face06Kandil7 /Athar-Embeddingstabular1M<n<10M2 likes7.6k downloads5mo agoHugging Face07asahi417 /seamless-align-enA-viA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes7k downloads2y agoHugging Face08scaleinvariant /paired-llama-3.2-1b-embeddings-lmsys-chat-1m Paired Llama 3.2 1B Token Embeddings (LMSYS-Chat-1M) This dataset contains paired activations corresponding to single token locations extracted from Meta's Llama 3.2 1B Instruct on conversations from LMSYS-Chat-1M. Embeddings are provided for layers 5 through 14, which capture the most interesting intermediate representations. This dataset was built to study things like: Learning different basis for activations at a given layer Studying if there are cases where position encodes… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/paired-llama-3.2-1b-embeddings-lmsys-chat-1m.tabularfeature-extraction100M<n<1B3 likes6.8k downloads7mo agoHugging Face09asahi417 /seamless-align-enA-zhA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes6.8k downloads2y agoHugging Face10sebasmos /latent-sr-embeddings Latent-SR Embeddings: Precomputed VAE Latents for Medical Image Super-Resolution Precomputed VAE latent embeddings from the paper: "Domain-Specific Latent Representations Improve the Fidelity of Diffusion-Based Medical Image Super-Resolution"Sebastian Cajas, Ashaba Judith, Rahul Gorijavolu, Sahil Kapadia, Hillary Clinton Kasimbazi, Leo Kinyera, Emmanuel Paul Kwesiga, Sri Sri Jaithra Varma Manthena, Luis Filipe Nakayama, Ninsiima Doreen, Leo Anthony Celi.arXiv:2604.12152 (2026)… See the full description on the dataset page: https://huggingface.co/datasets/sebasmos/latent-sr-embeddings.textimage-to-image10K<n<100K1 likes6.8k downloads3mo agoHugging Face11asahi417 /seamless-align-enA-frA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes5.1k downloads2y agoHugging Face12asahi417 /seamless-align-enA-zhA.speaker-embedding.hubert-xltabular100K<n<1M0 likes4.9k downloads2y agoHugging Face13lightonai /embeddings-fine-tuning Overview This dataset is composed of high quality data sources with mined hard negatives. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using this dataset or its curated version. This dataset has originally been created to follow the nv-retrieve setup, that mines the closest negatives to the query in a dataset and filter false negatives if their bi-encoder similarity is higher than a… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning.text10M<n<100M26 likes4.4k downloads3mo agoHugging Face14asahi417 /seamless-align-enA-zhA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes4.2k downloads2y agoHugging Face15justicedao /Caselaw_Access_Project_embeddingsThis is an embeddings dataset for the Caselaw Access Project, created by a user named Endomorphosis. Each caselaw entry is hashed with IPFS / multiformats, so retrieval of the document can be made over the IPFS / filecoin network The ipfs content id "cid" is the primary key that links the dataset to the embeddings, should you want to retrieve from the dataset instead. The dataset has been had embeddings generated with three models: thenlper/gte-small, Alibaba-NLP/gte-large-en-v1.5, and… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/Caselaw_Access_Project_embeddings.text10M<n<100M3 likes4k downloads2y agoHugging Face16asahi417 /seamless-align-deA-enA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes3.9k downloads2y agoHugging Face17asahi417 /seamless-align-enA-hiA.speaker-embedding.hubert-xltabular100K<n<1M0 likes3.3k downloads2y agoHugging Face18asahi417 /seamless-align-enA-jaA.speaker-embedding.xlsr-2btabular10K<n<100K0 likes3.2k downloads2y agoHugging Face19asahi417 /seamless-align-enA-frA.speaker-embedding.w2vbert-600mtabular1M<n<10M0 likes3k downloads2y agoHugging Face20laion /Caselaw_Access_Project_embeddingsOriginal Repository: https://huggingface.co/datasets/justicedao/Caselaw_Access_Project_embeddings/ This is an embeddings dataset for the Caselaw Access Project, created by a user named Endomorphosis. Each caselaw entry is hashed with IPFS / multiformats, so retrieval of the document can be made over the IPFS / filecoin network The ipfs content id "cid" is the primary key that links the dataset to the embeddings, should you want to retrieve from the dataset instead. The dataset has been had… See the full description on the dataset page: https://huggingface.co/datasets/laion/Caselaw_Access_Project_embeddings.textfeature-extraction10M<n<100M2 likes2.9k downloads2y agoHugging Face21asahi417 /seamless-align-enA-jaA.speaker-embedding.hubert-xltabular100K<n<1M1 likes2.8k downloads2y agoHugging Face22chriswolfram /embeddingstext100K<n<1M0 likes2.6k downloads2y agoHugging Face23MongoDB /subset_arxiv_papers_with_embeddingsThis dataset is a curated subset of the original arXiv dataset, each entry enriched with a 256-dimensional embedding vector. The embeddings are generated using OpenAI's "text-embedding-3-small" model. For each data point, the embedding is created by concatenating the text of the title, author(s), and abstract into a single string, which is then processed by the embedding model. This approach captures the semantic essence of each document, facilitating tasks such as similarity search… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/subset_arxiv_papers_with_embeddings.text10K<n<100K2 likes2.6k downloads2y agoHugging Face24asahi417 /seamless-align-enA-hiA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes2.3k downloads2y agoHugging Face25GlobalCampus /openalex-multilingual-embeddings OpenAlex Multilingual Embeddings This dataset contains multilingual text embeddings of all records in OpenAlex with a title or an abstract from the snapshot of 2023-10-20. The dataset was created for the FORAS project to investigate the efficacy of different methods of searching in databases of academic publications. All scripts will be available in a GitHub repository. The project is supported by a grant from the Dutch Research Council (grant no. 406.22.GO.048)… See the full description on the dataset page: https://huggingface.co/datasets/GlobalCampus/openalex-multilingual-embeddings.text100M<n<1B0 likes2.3k downloads3y agoHugging Face26duplexio /emilia-yodas-en-speaker-embeddings Emilia-YODAS English Qwen3-TTS Speaker Embeddings This dataset contains precomputed speaker embeddings for the English subset of Emilia-YODAS. Each row maps an Emilia-YODAS sample ID to one speaker embedding extracted from the corresponding audio. Dataset Details Source dataset: amphion/Emilia-Dataset Source subset: Emilia-YODAS English Embedding model: Qwen/Qwen3-TTS-12Hz-1.7B-Base Embedding shape: (2048,) Embedding dtype: float16 Rows: 4,516,833 Split: train Additional… See the full description on the dataset page: https://huggingface.co/datasets/duplexio/emilia-yodas-en-speaker-embeddings.textfeature-extraction1M<n<10M0 likes2.2k downloads5mo agoHugging Face27asahi417 /seamless-align-enA-viA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes2.1k downloads2y agoHugging Face28asahi417 /seamless-align-enA-hiA.speaker-embedding.xlsr-2btabular10K<n<100K0 likes2k downloads2y agoHugging Face29asahi417 /seamless-align-enA-koA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes2k downloads2y agoHugging Face30HIT-TMG /KaLM-embedding-pretrain-dataThe finetuning dataset is is available at this link:KaLM-Embedding/KaLM-embedding-finetuning-data. Citation If you find these datasets useful, please consider giving a star and citation. @misc{zhao2025kalmembeddingv2, title={KaLM-Embedding-V2: Superior Training Techniques and Data Inspire A Versatile Embedding Model}, author={Xinping Zhao and Xinshuo Hu and Zifei Shan and Shouzheng Huang and Yao Zhou and Xin Zhang and Zetian Sun and Zhenyu Liu and Dongfang Li and… See the full description on the dataset page: https://huggingface.co/datasets/HIT-TMG/KaLM-embedding-pretrain-data.text10M<n<100M22 likes1.8k downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.