Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01shash42 /forecast-news-embeddings Forecast News Embeddings Precomputed LanceDB table used for keyword, semantic, and hybrid retrieval in forecast-sim and future-sim. Snapshot 7,911,857 indexed source articles 16,207,764 text chunks Coverage: 2023-01-11 through 2026-08-31 Snapshot published: 2026-09-18 Lance dataset version: 856 Total artifact size: approximately 303.2 GiB Articles with empty searchable text are not represented. Long articles can produce multiple chunks, so the chunk count is… See the full description on the dataset page: https://huggingface.co/datasets/shash42/forecast-news-embeddings.2 likes18k downloads18d agoHugging Face02asahi417 /seamless-align-enA-viA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes16k downloads2y agoHugging Face03asahi417 /seamless-align-enA-esA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes15k downloads2y agoHugging Face04asahi417 /seamless-align-enA-jaA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes14k downloads2y agoHugging Face05lightonai /multilingual-embeddings-pre-training-curated 📚 Collection | 📝 Multilingual Blog | 📝 English Blog Contrastive Multilingual Pre-Training 2.16B query–document pairs across eight languages, plus cross-lingual pairs mDenseOn | mLateOn | DenseOn | LateOn | PyLate | FastPlaid 🎯 TL;DR: The multilingual contrastive pre-training corpus used to train mDenseOn and mLateOn. It extends our curated English data recipe (embeddings-pre-training-curated) to French, German, Italian, Spanish, Portuguese, Swedish, Norwegian, and Arabic… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/multilingual-embeddings-pre-training-curated.texttext-retrieval1B<n<10B6 likes11k downloads2mo agoHugging Face06asahi417 /seamless-align-enA-esA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes11k downloads2y agoHugging Face07asahi417 /seamless-align-enA-zhA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes10k downloads2y agoHugging Face08hysts-bot-data /daily-papers-embeddingstext10K<n<100K8 likes9.8k downloads50m agoHugging Face09asahi417 /seamless-align-deA-enA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes8.4k downloads2y agoHugging Face10asahi417 /seamless-align-enA-frA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes8.2k downloads2y agoHugging Face11sebasmos /latent-sr-embeddings Latent-SR Embeddings: Precomputed VAE Latents for Medical Image Super-Resolution Precomputed VAE latent embeddings from the paper: "Domain-Specific Latent Representations Improve the Fidelity of Diffusion-Based Medical Image Super-Resolution"Sebastian Cajas, Ashaba Judith, Rahul Gorijavolu, Sahil Kapadia, Hillary Clinton Kasimbazi, Leo Kinyera, Emmanuel Paul Kwesiga, Sri Sri Jaithra Varma Manthena, Luis Filipe Nakayama, Ninsiima Doreen, Leo Anthony Celi.arXiv:2604.12152 (2026)… See the full description on the dataset page: https://huggingface.co/datasets/sebasmos/latent-sr-embeddings.textimage-to-image10K<n<100K1 likes8.1k downloads3mo agoHugging Face12Kandil7 /Athar-Embeddingstabular1M<n<10M2 likes7.6k downloads5mo agoHugging Face13asahi417 /seamless-align-enA-hiA.speaker-embedding.hubert-xltabular100K<n<1M0 likes7.5k downloads2y agoHugging Face14asahi417 /seamless-align-enA-zhA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes7.3k downloads2y agoHugging Face15asahi417 /seamless-align-enA-frA.speaker-embedding.w2vbert-600mtabular1M<n<10M0 likes6.6k downloads2y agoHugging Face16scaleinvariant /paired-llama-3.2-1b-embeddings-lmsys-chat-1m Paired Llama 3.2 1B Token Embeddings (LMSYS-Chat-1M) This dataset contains paired activations corresponding to single token locations extracted from Meta's Llama 3.2 1B Instruct on conversations from LMSYS-Chat-1M. Embeddings are provided for layers 5 through 14, which capture the most interesting intermediate representations. This dataset was built to study things like: Learning different basis for activations at a given layer Studying if there are cases where position encodes… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/paired-llama-3.2-1b-embeddings-lmsys-chat-1m.tabularfeature-extraction100M<n<1B3 likes6.4k downloads7mo agoHugging Face17asahi417 /seamless-align-enA-zhA.speaker-embedding.hubert-xltabular100K<n<1M0 likes6.4k downloads2y agoHugging Face18justicedao /Caselaw_Access_Project_embeddingsThis is an embeddings dataset for the Caselaw Access Project, created by a user named Endomorphosis. Each caselaw entry is hashed with IPFS / multiformats, so retrieval of the document can be made over the IPFS / filecoin network The ipfs content id "cid" is the primary key that links the dataset to the embeddings, should you want to retrieve from the dataset instead. The dataset has been had embeddings generated with three models: thenlper/gte-small, Alibaba-NLP/gte-large-en-v1.5, and… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/Caselaw_Access_Project_embeddings.text10M<n<100M2 likes5.9k downloads2y agoHugging Face19lightonai /embeddings-pre-training-curated Embeddings pre-training curated data This dataset is the English subset of lightonai/embeddings-pre-training, assembled to reproduce the English data recipe described in the mGTE technical report (Zhang et al., 2024). The mGTE paper describes the data sources used to train the GTE family of multilingual text embedding and reranking models, but does not release the data itself. This dataset is our reconstruction of the English portion of that recipe, curated as part of a research… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-pre-training-curated.100M<n<1B15 likes5.1k downloads1mo agoHugging Face20lightonai /embeddings-pre-training Overview This large-scale dataset is designed for pre-training state-of-the-art text embedding models. Its goal is to reproduce and build upon the data recipe described in the mGTE technical report (Zhang et al., 2024), which details the data sources used to train the GTE family of embedding models but does not release the data itself. We assembled this dataset as part of a research effort to understand how data composition affects retrieval model quality. Our experiments… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-pre-training.51 likes5.1k downloads1mo agoHugging Face21kshitijd /platonic-embeddingstimeseries1M<n<10M0 likes5k downloads5mo agoHugging Face22asahi417 /seamless-align-enA-hiA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes4.9k downloads2y agoHugging Face23VDBBench /multimodal-embedding-100M Multimodal Embedding 100M This dataset contains a 100M-row multimodal embedding corpus generated from LAION-style image-text data exported with img2dataset as WebDataset shards. Images were resized to 256 during the WebDataset creation step before embedding generation. The dataset is intended for large-scale vector database ingestion, ANN index construction, nearest-neighbor search, and retrieval benchmark experiments. The dataset is stored as Parquet files and organized to keep… See the full description on the dataset page: https://huggingface.co/datasets/VDBBench/multimodal-embedding-100M.feature-extraction100M<n<1B3 likes4.9k downloads3mo agoHugging Face24JQL-AI /curated_embeddings0 likes4.9k downloads1y agoHugging Face25asahi417 /seamless-align-enA-jaA.speaker-embedding.xlsr-2btabular10K<n<100K0 likes4.8k downloads2y agoHugging Face26hotchpotch /bekko-embedding-v1-unsupervised Bekko Embedding v1 Unsupervised Training Data This is the unsupervised training dataset used for the bekko-embedding-v1 embedding model family. It contains pair and triplet examples for embedding pretraining, published as Hugging Face dataset subsets so that each source/subset/task combination can be loaded independently. For the full training recipe and technical details, see Bekko Embedding: Parameter-Efficient Multilingual Retrieval with Ultra-Compact Encoders. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/bekko-embedding-v1-unsupervised.text-retrieval4 likes4.7k downloads2mo agoHugging Face27asahi417 /seamless-align-enA-viA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes4.7k downloads2y agoHugging Face28asahi417 /seamless-align-enA-jaA.speaker-embedding.hubert-xltabular100K<n<1M0 likes4.5k downloads2y agoHugging Face29asahi417 /seamless-align-enA-hiA.speaker-embedding.xlsr-2btabular10K<n<100K0 likes4.3k downloads2y agoHugging Face30laion /Caselaw_Access_Project_embeddingsOriginal Repository: https://huggingface.co/datasets/justicedao/Caselaw_Access_Project_embeddings/ This is an embeddings dataset for the Caselaw Access Project, created by a user named Endomorphosis. Each caselaw entry is hashed with IPFS / multiformats, so retrieval of the document can be made over the IPFS / filecoin network The ipfs content id "cid" is the primary key that links the dataset to the embeddings, should you want to retrieve from the dataset instead. The dataset has been had… See the full description on the dataset page: https://huggingface.co/datasets/laion/Caselaw_Access_Project_embeddings.textfeature-extraction10M<n<100M2 likes3.9k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.