datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
embeddinggemma-2-GGUF-metrics
embeddinggemma-2 GGUF, everything behind the numbers
This dataset holds the measurements, logs and inputs behind
AtomicChat/embeddinggemma-2-GGUF.
What is here
Path
What it is
results.json
Every measured file: ours, Unsloth's, AutoRound's (webmp3/Sakura-EmbeddingGemma-2-AutoRound-GGUF) and ggml-org's Q8_0. Each row has the size and, per eval set and width (768, 256), the mean and p99 of 1 - cosine to BF16, the same-top-result rate with its 95% interval… See the full description on the dataset page: https://huggingface.co/datasets/AtomicChat/embeddinggemma-2-GGUF-metrics.wikipedia-40-langs-with-embeddings-embeddinggemma2
Wikipedia 40 Languages with EmbeddingGemma-2 Retrieval Embeddings
Precomputed teacher embeddings from google/embeddinggemma-2
for offline embedding distillation, built on alibayram/wikipedia-40-langs.
It is the distillation corpus for alibayram/embeddingmagibu2,
following the pipeline of arXiv:2605.29992.
Status: splits are added as they finish (test → validation → train).
Live progress and ETA: progress.json. Files under progress/ are partial
chunks of the split being encoded… See the full description on the dataset page: https://huggingface.co/datasets/alibayram/wikipedia-40-langs-with-embeddings-embeddinggemma2.trama-embeddinggemma
EmbeddingGemma 300M for LiteRT (copy used by the Trama browser)
This repository hosts an unmodified copy of two files published by Google in
litert-community/embeddinggemma-300m,
so that the Trama Android browser can download them, only when the user
explicitly enables on-device semantic search. Everything then runs on the phone;
no text ever leaves it.
File
Size (bytes)
SHA-256
embeddinggemma-300M_seq256_mixed-precision.tflite
179131736… See the full description on the dataset page: https://huggingface.co/datasets/Bitfarmy/trama-embeddinggemma.bharat-government-documents-embeddinggemma-300m
Bharat Government Documents — EmbeddingGemma-300m embeddings
All 330,688 rows of ankitjh4/bharat-government-documents
(documents.csv.gz, all 19 original columns kept) plus an embedding column: a 768-d float32,
L2-normalized vector from google/embeddinggemma-300m.
How the embeddings were made
Input text: title: {title} | text: {text} (title: none when empty), EmbeddingGemma's document prompt.
Truncated to the model max of 2048 tokens (longer documents are… See the full description on the dataset page: https://huggingface.co/datasets/anudit/bharat-government-documents-embeddinggemma-300m.WebFAQ-SWE-emb-EmbeddingGemma300m-v1corpus_embeddings_embeddinggemma-300mdoctorsearch_chicago_250812_1630_enriched_uchicago_embeddinggemma-300m_with_embeddingstldr_vs_abstract_google_embeddinggemma-300m
