datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikipedia-40-langs-with-embeddings-embeddinggemma2
Wikipedia 40 Languages with EmbeddingGemma-2 Retrieval Embeddings
Precomputed teacher embeddings from google/embeddinggemma-2
for offline embedding distillation, built on alibayram/wikipedia-40-langs.
It is the distillation corpus for alibayram/embeddingmagibu2,
following the pipeline of arXiv:2605.29992.
Status: splits are added as they finish (test → validation → train).
Live progress and ETA: progress.json. Files under progress/ are partial
chunks of the split being encoded… See the full description on the dataset page: https://huggingface.co/datasets/alibayram/wikipedia-40-langs-with-embeddings-embeddinggemma2.bharat-government-documents-embeddinggemma-300m
Bharat Government Documents — EmbeddingGemma-300m embeddings
All 330,688 rows of ankitjh4/bharat-government-documents
(documents.csv.gz, all 19 original columns kept) plus an embedding column: a 768-d float32,
L2-normalized vector from google/embeddinggemma-300m.
How the embeddings were made
Input text: title: {title} | text: {text} (title: none when empty), EmbeddingGemma's document prompt.
Truncated to the model max of 2048 tokens (longer documents are… See the full description on the dataset page: https://huggingface.co/datasets/anudit/bharat-government-documents-embeddinggemma-300m.WebFAQ-SWE-emb-EmbeddingGemma300m-v1doctorsearch_chicago_250812_1630_enriched_uchicago_embeddinggemma-300m_with_embeddingstldr_vs_abstract_google_embeddinggemma-300m
