Team Ai
Modelpublic

apple/coreml-mobileclip

sourceHugging Faceotherupdated 2y agoView on Hugging Face
55likes2.5kdownloads
Model Card

MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training

MobileCLIP was introduced in MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training (CVPR 2024), by Pavan Kumar Anasosalu Vasu, Hadi Pouransari, Fartash Faghri, Raviteja Vemulapalli, Oncel Tuzel.

This repository contains the text and image encoders of all variants of MobileCLIP exported to Core ML. These Core ML models can be plugged-into the demo app provided in the official MobileCLIP repo

<img src="mce_example.gif" width="240" height="540" />

Highlights

  • —Our smallest variant MobileCLIP-S0 obtains similar zero-shot performance as OpenAI's ViT-B/16 model while being 4.8x faster and 2.8x smaller.
  • —MobileCLIP-S2 obtains better avg zero-shot performance than SigLIP's ViT-B/16 model while being 2.3x faster and 2.1x smaller, and trained with 3x less seen samples.
  • —MobileCLIP-B(LT) attains zero-shot ImageNet performance of 77.2% which is significantly better than recent works like DFN and SigLIP with similar architectures or even OpenAI's ViT-L/14@336.

Checkpoints

Model# Seen <BR>Samples (B)# Params (M) <BR> (img + txt)Latency (ms) <BR> (img + txt)IN-1k Zero-Shot <BR> Top-1 Acc. (%)Avg. Perf. (%) <BR> on 38 datasets
MobileCLIP-S01311.4 + 42.41.5 + 1.667.858.1
MobileCLIP-S11321.5 + 63.42.5 + 3.372.661.3
MobileCLIP-S21335.7 + 63.43.6 + 3.374.463.7
MobileCLIP-B1386.3 + 63.410.4 + 3.376.865.2
MobileCLIP-B (LT)3686.3 + 63.410.4 + 3.377.265.8

Download

Install huggingface-cli

bash
brew install huggingface-cli
bash
huggingface-cli download --local-dir models apple/coreml-mobileclip

Citation

[MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training](https://arxiv.org/pdf/2311.17049.pdf). (CVPR 2024) Pavan Kumar Anasosalu Vasu, Hadi Pouransari, Fartash Faghri, Raviteja Vemulapalli, Oncel Tuzel.

bibtex
@InProceedings{mobileclip2024,
  author = {Pavan Kumar Anasosalu Vasu, Hadi Pouransari, Fartash Faghri, Raviteja Vemulapalli, Oncel Tuzel},
  title = {MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training},
  booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
  month = {June},
  year = {2024},
}