embedding
forecast-news-embeddings
Forecast News Embeddings
Precomputed LanceDB table used for keyword, semantic, and hybrid retrieval in
forecast-sim and future-sim.
Snapshot
7,911,857 indexed source articles
16,207,764 text chunks
Coverage: 2023-01-11 through 2026-08-31
Snapshot published: 2026-09-18
Lance dataset version: 856
Total artifact size: approximately 303.2 GiB
Articles with empty searchable text are not represented. Long articles can
produce multiple chunks, so the chunk count is… See the full description on the dataset page: https://huggingface.co/datasets/shash42/forecast-news-embeddings.seamless-align-enA-viA.speaker-embedding.xlsr-2bseamless-align-enA-esA.speaker-embedding.w2vbert-600mseamless-align-enA-jaA.speaker-embedding.w2vbert-600mseamless-align-enA-zhA.speaker-embedding.w2vbert-600mmultilingual-embeddings-pre-training-curated
📚 Collection | 📝 Multilingual Blog | 📝 English Blog
Contrastive Multilingual Pre-Training
2.16B query–document pairs across eight languages, plus cross-lingual pairs
mDenseOn |
mLateOn |
DenseOn |
LateOn |
PyLate |
FastPlaid
🎯 TL;DR: The multilingual contrastive pre-training corpus used to train mDenseOn and mLateOn. It extends our curated English data recipe (embeddings-pre-training-curated) to French, German, Italian, Spanish, Portuguese, Swedish, Norwegian, and Arabic… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/multilingual-embeddings-pre-training-curated.
