Team Ai
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01shash42 /forecast-news-embeddings Forecast News Embeddings Precomputed LanceDB table used for keyword, semantic, and hybrid retrieval in forecast-sim and future-sim. Snapshot 7,911,857 indexed source articles 16,207,764 text chunks Coverage: 2023-01-11 through 2026-08-31 Snapshot published: 2026-09-18 Lance dataset version: 856 Total artifact size: approximately 303.2 GiB Articles with empty searchable text are not represented. Long articles can produce multiple chunks, so the chunk count is… See the full description on the dataset page: https://huggingface.co/datasets/shash42/forecast-news-embeddings.2 likes18k downloads22d agoHugging Face02asahi417 /seamless-align-enA-esA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes12k downloads2y agoHugging Face03asahi417 /seamless-align-enA-esA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes12k downloads2y agoHugging Face04lightonai /multilingual-embeddings-pre-training-curated 📚 Collection | 📝 Multilingual Blog | 📝 English Blog Contrastive Multilingual Pre-Training 2.16B query–document pairs across eight languages, plus cross-lingual pairs mDenseOn | mLateOn | DenseOn | LateOn | PyLate | FastPlaid 🎯 TL;DR: The multilingual contrastive pre-training corpus used to train mDenseOn and mLateOn. It extends our curated English data recipe (embeddings-pre-training-curated) to French, German, Italian, Spanish, Portuguese, Swedish, Norwegian, and Arabic… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/multilingual-embeddings-pre-training-curated.texttext-retrieval1B<n<10B6 likes11k downloads2mo agoHugging Face05hysts-bot-data /daily-papers-embeddingstext10K<n<100K10 likes9.7k downloads23h agoHugging Face06asahi417 /seamless-align-enA-jaA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes9.5k downloads2y agoHugging Face07lightonai /embeddings-pre-training-curated Embeddings pre-training curated data This dataset is the English subset of lightonai/embeddings-pre-training, assembled to reproduce the English data recipe described in the mGTE technical report (Zhang et al., 2024). The mGTE paper describes the data sources used to train the GTE family of multilingual text embedding and reranking models, but does not release the data itself. This dataset is our reconstruction of the English portion of that recipe, curated as part of a research… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-pre-training-curated.100M<n<1B15 likes8.2k downloads2mo agoHugging Face08Kandil7 /Athar-Embeddingstabular1M<n<10M2 likes7.6k downloads5mo agoHugging Face09asahi417 /seamless-align-enA-viA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes7k downloads2y agoHugging Face10scaleinvariant /paired-llama-3.2-1b-embeddings-lmsys-chat-1m Paired Llama 3.2 1B Token Embeddings (LMSYS-Chat-1M) This dataset contains paired activations corresponding to single token locations extracted from Meta's Llama 3.2 1B Instruct on conversations from LMSYS-Chat-1M. Embeddings are provided for layers 5 through 14, which capture the most interesting intermediate representations. This dataset was built to study things like: Learning different basis for activations at a given layer Studying if there are cases where position encodes… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/paired-llama-3.2-1b-embeddings-lmsys-chat-1m.tabularfeature-extraction100M<n<1B3 likes6.8k downloads7mo agoHugging Face11asahi417 /seamless-align-enA-zhA.speaker-embedding.w2vbert-600mtabular100K<n<1M0 likes6.8k downloads2y agoHugging Face12sebasmos /latent-sr-embeddings Latent-SR Embeddings: Precomputed VAE Latents for Medical Image Super-Resolution Precomputed VAE latent embeddings from the paper: "Domain-Specific Latent Representations Improve the Fidelity of Diffusion-Based Medical Image Super-Resolution"Sebastian Cajas, Ashaba Judith, Rahul Gorijavolu, Sahil Kapadia, Hillary Clinton Kasimbazi, Leo Kinyera, Emmanuel Paul Kwesiga, Sri Sri Jaithra Varma Manthena, Luis Filipe Nakayama, Ninsiima Doreen, Leo Anthony Celi.arXiv:2604.12152 (2026)… See the full description on the dataset page: https://huggingface.co/datasets/sebasmos/latent-sr-embeddings.textimage-to-image10K<n<100K1 likes6.8k downloads3mo agoHugging Face13asahi417 /seamless-align-enA-frA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes5.1k downloads2y agoHugging Face14asahi417 /seamless-align-enA-zhA.speaker-embedding.hubert-xltabular100K<n<1M0 likes4.9k downloads2y agoHugging Face15kshitijd /platonic-embeddingstimeseries1M<n<10M0 likes4.9k downloads5mo agoHugging Face16lightonai /embeddings-pre-training Overview This large-scale dataset is designed for pre-training state-of-the-art text embedding models. Its goal is to reproduce and build upon the data recipe described in the mGTE technical report (Zhang et al., 2024), which details the data sources used to train the GTE family of embedding models but does not release the data itself. We assembled this dataset as part of a research effort to understand how data composition affects retrieval model quality. Our experiments… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-pre-training.51 likes4.8k downloads2mo agoHugging Face17VDBBench /multimodal-embedding-100M Multimodal Embedding 100M This dataset contains a 100M-row multimodal embedding corpus generated from LAION-style image-text data exported with img2dataset as WebDataset shards. Images were resized to 256 during the WebDataset creation step before embedding generation. The dataset is intended for large-scale vector database ingestion, ANN index construction, nearest-neighbor search, and retrieval benchmark experiments. The dataset is stored as Parquet files and organized to keep… See the full description on the dataset page: https://huggingface.co/datasets/VDBBench/multimodal-embedding-100M.feature-extraction100M<n<1B3 likes4.7k downloads4mo agoHugging Face18hotchpotch /bekko-embedding-v1-unsupervised Bekko Embedding v1 Unsupervised Training Data This is the unsupervised training dataset used for the bekko-embedding-v1 embedding model family. It contains pair and triplet examples for embedding pretraining, published as Hugging Face dataset subsets so that each source/subset/task combination can be loaded independently. For the full training recipe and technical details, see Bekko Embedding: Parameter-Efficient Multilingual Retrieval with Ultra-Compact Encoders. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/bekko-embedding-v1-unsupervised.text-retrieval4 likes4.7k downloads2mo agoHugging Face19lightonai /embeddings-fine-tuning Overview This dataset is composed of high quality data sources with mined hard negatives. It can be used to train a strong retrieval model by itself but is better used after a large-scale contrastive pre-training, for example using this dataset or its curated version. This dataset has originally been created to follow the nv-retrieve setup, that mines the closest negatives to the query in a dataset and filter false negatives if their bi-encoder similarity is higher than a… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-fine-tuning.text10M<n<100M26 likes4.4k downloads3mo agoHugging Face20asahi417 /seamless-align-enA-zhA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes4.2k downloads2y agoHugging Face21JQL-AI /curated_embeddings0 likes4.2k downloads1y agoHugging Face22justicedao /Caselaw_Access_Project_embeddingsThis is an embeddings dataset for the Caselaw Access Project, created by a user named Endomorphosis. Each caselaw entry is hashed with IPFS / multiformats, so retrieval of the document can be made over the IPFS / filecoin network The ipfs content id "cid" is the primary key that links the dataset to the embeddings, should you want to retrieve from the dataset instead. The dataset has been had embeddings generated with three models: thenlper/gte-small, Alibaba-NLP/gte-large-en-v1.5, and… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/Caselaw_Access_Project_embeddings.text10M<n<100M3 likes4k downloads2y agoHugging Face23asahi417 /seamless-align-deA-enA.speaker-embedding.xlsr-2btabular100K<n<1M0 likes3.9k downloads2y agoHugging Face24allenai /olmoearth-paper-embeddings OlmoEarth — Foundation-Model Embeddings for Paper Table 2 This dataset contains pre-extracted embeddings from 26 Earth-observation foundation models evaluated on the 24 downstream tasks that make up Table 2 of the OlmoEarth paper: OlmoEarth: Stable Latent Image Modeling for Multimodal Earth Observation AI2, 2025. arXiv:2511.13655. For every supported (model, task) pair we ran the model's encoder over the task's train / validation / test splits with the paper-best… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmoearth-paper-embeddings.geospatialfeature-extraction10M<n<100M8 likes3.5k downloads4mo agoHugging Face25asahi417 /seamless-align-enA-hiA.speaker-embedding.hubert-xltabular100K<n<1M0 likes3.3k downloads2y agoHugging Face26asahi417 /seamless-align-enA-jaA.speaker-embedding.xlsr-2btabular10K<n<100K0 likes3.2k downloads2y agoHugging Face27asahi417 /seamless-align-enA-frA.speaker-embedding.w2vbert-600mtabular1M<n<10M0 likes3k downloads2y agoHugging Face28laion /Caselaw_Access_Project_embeddingsOriginal Repository: https://huggingface.co/datasets/justicedao/Caselaw_Access_Project_embeddings/ This is an embeddings dataset for the Caselaw Access Project, created by a user named Endomorphosis. Each caselaw entry is hashed with IPFS / multiformats, so retrieval of the document can be made over the IPFS / filecoin network The ipfs content id "cid" is the primary key that links the dataset to the embeddings, should you want to retrieve from the dataset instead. The dataset has been had… See the full description on the dataset page: https://huggingface.co/datasets/laion/Caselaw_Access_Project_embeddings.textfeature-extraction10M<n<100M2 likes2.9k downloads2y agoHugging Face29asahi417 /seamless-align-enA-jaA.speaker-embedding.hubert-xltabular100K<n<1M1 likes2.8k downloads2y agoHugging Face30chriswolfram /embeddingstext100K<n<1M0 likes2.6k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. Team Ai does not host these files.