datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
forecast-news-embeddings
Forecast News Embeddings
Precomputed LanceDB table used for keyword, semantic, and hybrid retrieval in
forecast-sim and future-sim.
Snapshot
7,911,857 indexed source articles
16,207,764 text chunks
Coverage: 2023-01-11 through 2026-08-31
Snapshot published: 2026-09-18
Lance dataset version: 856
Total artifact size: approximately 303.2 GiB
Articles with empty searchable text are not represented. Long articles can
produce multiple chunks, so the chunk count is… See the full description on the dataset page: https://huggingface.co/datasets/shash42/forecast-news-embeddings.seamless-align-enA-viA.speaker-embedding.xlsr-2bseamless-align-enA-esA.speaker-embedding.w2vbert-600mseamless-align-enA-jaA.speaker-embedding.w2vbert-600mmultilingual-embeddings-pre-training-curated
📚 Collection | 📝 Multilingual Blog | 📝 English Blog
Contrastive Multilingual Pre-Training
2.16B query–document pairs across eight languages, plus cross-lingual pairs
mDenseOn |
mLateOn |
DenseOn |
LateOn |
PyLate |
FastPlaid
🎯 TL;DR: The multilingual contrastive pre-training corpus used to train mDenseOn and mLateOn. It extends our curated English data recipe (embeddings-pre-training-curated) to French, German, Italian, Spanish, Portuguese, Swedish, Norwegian, and Arabic… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/multilingual-embeddings-pre-training-curated.seamless-align-enA-esA.speaker-embedding.xlsr-2bseamless-align-enA-zhA.speaker-embedding.w2vbert-600mdaily-papers-embeddingsseamless-align-deA-enA.speaker-embedding.xlsr-2bseamless-align-enA-frA.speaker-embedding.xlsr-2blatent-sr-embeddings
Latent-SR Embeddings: Precomputed VAE Latents for Medical Image Super-Resolution
Precomputed VAE latent embeddings from the paper:
"Domain-Specific Latent Representations Improve the Fidelity of Diffusion-Based Medical Image Super-Resolution"Sebastian Cajas, Ashaba Judith, Rahul Gorijavolu, Sahil Kapadia, Hillary Clinton Kasimbazi, Leo Kinyera, Emmanuel Paul Kwesiga, Sri Sri Jaithra Varma Manthena, Luis Filipe Nakayama, Ninsiima Doreen, Leo Anthony Celi.arXiv:2604.12152 (2026)… See the full description on the dataset page: https://huggingface.co/datasets/sebasmos/latent-sr-embeddings.Athar-Embeddingsseamless-align-enA-hiA.speaker-embedding.hubert-xlseamless-align-enA-zhA.speaker-embedding.xlsr-2bseamless-align-enA-frA.speaker-embedding.w2vbert-600mpaired-llama-3.2-1b-embeddings-lmsys-chat-1m
Paired Llama 3.2 1B Token Embeddings (LMSYS-Chat-1M)
This dataset contains paired activations corresponding to single token locations extracted from Meta's Llama 3.2 1B Instruct on conversations from LMSYS-Chat-1M.
Embeddings are provided for layers 5 through 14, which capture the most interesting intermediate representations.
This dataset was built to study things like:
Learning different basis for activations at a given layer
Studying if there are cases where position encodes… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/paired-llama-3.2-1b-embeddings-lmsys-chat-1m.seamless-align-enA-zhA.speaker-embedding.hubert-xlCaselaw_Access_Project_embeddingsThis is an embeddings dataset for the Caselaw Access Project, created by a user named Endomorphosis.
Each caselaw entry is hashed with IPFS / multiformats, so retrieval of the document can be made over the IPFS / filecoin network
The ipfs content id "cid" is the primary key that links the dataset to the embeddings, should you want to retrieve from the dataset instead.
The dataset has been had embeddings generated with three models: thenlper/gte-small, Alibaba-NLP/gte-large-en-v1.5, and… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/Caselaw_Access_Project_embeddings.embeddings-pre-training-curated
Embeddings pre-training curated data
This dataset is the English subset of lightonai/embeddings-pre-training, assembled to reproduce the English data recipe described in the mGTE technical report (Zhang et al., 2024).
The mGTE paper describes the data sources used to train the GTE family of multilingual text embedding and reranking models, but does not release the data itself. This dataset is our reconstruction of the English portion of that recipe, curated as part of a research… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-pre-training-curated.embeddings-pre-training
Overview
This large-scale dataset is designed for pre-training state-of-the-art text embedding models. Its goal is to reproduce and build upon the data recipe described in the mGTE technical report (Zhang et al., 2024), which details the data sources used to train the GTE family of embedding models but does not release the data itself.
We assembled this dataset as part of a research effort to understand how data composition affects retrieval model quality. Our experiments… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/embeddings-pre-training.platonic-embeddingsseamless-align-enA-hiA.speaker-embedding.w2vbert-600mmultimodal-embedding-100M
Multimodal Embedding 100M
This dataset contains a 100M-row multimodal embedding corpus generated from LAION-style image-text data exported with img2dataset as WebDataset shards. Images were resized to 256 during the WebDataset creation step before embedding generation. The dataset is intended for large-scale vector database ingestion, ANN index construction, nearest-neighbor search, and retrieval benchmark experiments.
The dataset is stored as Parquet files and organized to keep… See the full description on the dataset page: https://huggingface.co/datasets/VDBBench/multimodal-embedding-100M.curated_embeddingsseamless-align-enA-jaA.speaker-embedding.xlsr-2bbekko-embedding-v1-unsupervised
Bekko Embedding v1 Unsupervised Training Data
This is the unsupervised training dataset used for the bekko-embedding-v1 embedding model family. It contains pair and triplet examples for embedding pretraining, published as Hugging Face dataset subsets so that each source/subset/task combination can be loaded independently.
For the full training recipe and technical details, see Bekko Embedding: Parameter-Efficient Multilingual Retrieval with Ultra-Compact Encoders.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/bekko-embedding-v1-unsupervised.seamless-align-enA-viA.speaker-embedding.w2vbert-600mseamless-align-enA-jaA.speaker-embedding.hubert-xlseamless-align-enA-hiA.speaker-embedding.xlsr-2bCaselaw_Access_Project_embeddingsOriginal Repository:
https://huggingface.co/datasets/justicedao/Caselaw_Access_Project_embeddings/
This is an embeddings dataset for the Caselaw Access Project, created by a user named Endomorphosis.
Each caselaw entry is hashed with IPFS / multiformats, so retrieval of the document can be made over the IPFS / filecoin network
The ipfs content id "cid" is the primary key that links the dataset to the embeddings, should you want to retrieve from the dataset instead.
The dataset has been had… See the full description on the dataset page: https://huggingface.co/datasets/laion/Caselaw_Access_Project_embeddings.
