datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multilingual-embeddings-pre-training-curated
📚 Collection | 📝 Multilingual Blog | 📝 English Blog
Contrastive Multilingual Pre-Training
2.16B query–document pairs across eight languages, plus cross-lingual pairs
mDenseOn |
mLateOn |
DenseOn |
LateOn |
PyLate |
FastPlaid
🎯 TL;DR: The multilingual contrastive pre-training corpus used to train mDenseOn and mLateOn. It extends our curated English data recipe (embeddings-pre-training-curated) to French, German, Italian, Spanish, Portuguese, Swedish, Norwegian, and Arabic… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/multilingual-embeddings-pre-training-curated.openalex-multilingual-embeddings
OpenAlex Multilingual Embeddings
This dataset contains multilingual text embeddings of all records in OpenAlex with a title or an abstract from the snapshot of 2023-10-20.
The dataset was created for the FORAS project to investigate the efficacy of
different methods of searching in databases of academic publications. All scripts will be available in a GitHub repository.
The project is supported by a grant from the Dutch Research Council (grant no. 406.22.GO.048)… See the full description on the dataset page: https://huggingface.co/datasets/GlobalCampus/openalex-multilingual-embeddings.
