Team Ai
Modelpublic

stefra/embeddinggemma2-glance-v1

sourceHugging Faceapache-2.0updated 3d agoView on Hugging Face
0likes19downloads
Model Card

embeddinggemma2-glance-v1

Architettura GLANCE Gemma: EmbeddingGemma 2 addestrato con statement tuning multi-domanda, su testo e immagini insieme. Dato uno stato (un testo, una o due immagini, o testo e immagini) e uno o piu' statement, restituisce la probabilita' che ogni statement sia vero.

Le immagini diventano token dentro la stessa sequenza del testo; gli statement stanno nella stessa sequenza dello stato ma una maschera di attenzione a blocchi li rende indipendenti: il punteggio di uno statement e' identico a quello che si otterrebbe valutandolo da solo, e lo stato (immagini comprese) si calcola una volta sola.

Uso

Basta transformers (con il supporto a EmbeddingGemma 2): il codice del modello e' incluso nel repository (trust_remote_code=True).

python
from PIL import Image
from transformers import AutoModel

model = AutoModel.from_pretrained("stefra/embeddinggemma2-glance-v1", trust_remote_code=True)

model.predict("The vase is broken.", ["The vase is intact.", "Something got damaged."])  # testo
model.predict(Image.open("gatto.jpg"), ["There is a cat.", "The cat is black."])          # immagine
model.predict({"images": ["sx.jpg", "dx.jpg"]}, ["The left image has more dogs."])       # due immagini
model.predict({"text": "Lot 12, 1950s", "images": ["vaso.jpg"]}, ["The vase is antique."])  # testo + immagine
model.predict([("I love it.", ["It is positive."]), ({"images": ["a.jpg"]}, ["It is a dog."])])  # batch

Una stringa nuda e' sempre testo; le immagini sono PIL.Image, array numpy o bytes, e dentro {"images": [...]} anche percorsi e URL. Classificazione zero-shot: uno statement per classe e vince il piu' probabile.

Precisione: EmbeddingGemma 2 non supporta float16. Su GPU con bf16 nativo (A100, L4, H100) il modello si carica in bfloat16, altrimenti in float32; si puo' scegliere con device= e dtype=.

Metriche

Validazione (sorgenti di training)

sorgentenaccuracyF1ROC-AUCBrierECEtop-1
absa5070.9330.9340.9850.0500.027–
ade4810.9040.9050.9530.0810.054–
amazon_reviews4880.8710.8630.9380.1010.071–
app_reviews4970.7750.7690.8590.1640.104–
banking775100.9610.9610.9890.0340.016–
coco5100.9040.9020.9740.0650.037–
coco_pairs5790.6840.7260.8120.1630.062–
complaints5100.9510.9510.9930.0380.030–
dbpedia5100.9900.9901.0000.0090.008–
dpr3960.5150.4920.5030.2510.016–
entity_matching3420.9120.9110.9770.0650.033–
fewnerd5100.9180.9170.9680.0660.031–
gqa5960.7850.7830.8830.1400.058–
massive5100.9470.9480.9870.0450.033–
mintaka1260.8570.7910.9480.0880.073–
mnli5100.7570.7660.8260.1700.033–
nlvr21020.6860.6730.7450.2110.118–
paws2560.5270.2190.5900.2520.090–
piqa3400.5650.5910.5620.2560.057–
product_catalog5070.8600.8610.9300.1080.059–
qasc2550.9490.9190.9850.0360.033–
qqp2790.8710.8090.9480.0930.065–
race4310.6240.5890.6510.2380.055–
samsum4701.0001.0001.0000.0000.003–
sciq2550.8510.7890.9370.0980.040–
snli5100.8430.8440.9320.1070.057–
snli_ve6000.8150.8190.8950.1310.030–
squad5020.7090.7210.7730.1910.038–
tweet_irony4530.6310.7080.6920.2240.070–
tweet_offensive5100.7880.7920.8770.1550.093–
tweet_sentiment5100.7590.7670.8310.1730.074–
tweet_stance4350.7720.7600.8730.1460.042–
winogrande5100.5160.5500.5200.2500.021–
yahoo_answers5100.8650.8640.9440.0990.050–
yelp_polarity5100.9550.9550.9930.0360.023–
ALL image23870.7890.7960.8990.1300.035–
ALL text131400.8180.8150.9170.1190.029–
ALL155270.8140.8120.9140.1210.029–

Held-out (zero-shot)

sorgentenaccuracyF1ROC-AUCBrierECEtop-1
ag_news30000.7870.7780.8810.1620.122–
emotion30000.7500.7580.8000.1850.073–
food101303000.6710.0550.9210.2360.3330.397
oxford_pets111000.2460.0640.6990.5660.6740.247
rotten_tomatoes30000.7350.7370.8200.1860.100–
ALL image414000.5570.0590.8590.3250.424–
ALL text90000.7580.7570.8330.1780.094–
ALL504000.5930.2800.7470.2980.346–