blckrvrfx/edge-multimodal-embeddings
0
π Edge Multimodal Embeddings
Production-grade multimodal embedding system that runs on edge devices. Handles text, images, video, audio, PDFs, code, and 40+ file formats β all under 4GB VRAM.
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β EdgeMultimodalEmbedder β
β β
β ββββββββββββ ββββββββββββ ββββββββββββ ββββββββββββ β
β β Text β β Vision β β Audio β β Video β β
β β (Nomic) β β (Nomic) β β (CLAP) β β(VideoMAE)β β
β β 768d β β 768d β β 512d β β 384d β β
β ββββββ¬ββββββ ββββββ¬ββββββ ββββββ¬ββββββ ββββββ¬ββββββ β
β β β β β β
β ββββββββββββββββ΄βββββββ¬βββββββ΄βββββββββββββββ β
β β β
β ββββββββββββββββ΄βββββββββββββββ β
β β Projection Heads β 512d β β
β ββββββββββββββββ¬βββββββββββββββ β
β β β
β L2-Normalized 512d β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ¨ Key Features
π Model Stack
Text β Image share the same 768d space via Nomic. No projection needed for text-image retrieval!
ποΈ Hardware Profiles
from edge_multimodal_embeddings import EdgeMultimodalEmbedder
# Quick start β auto-selects profile
embedder = EdgeMultimodalEmbedder.from_profile("edge_gpu")
# Profiles:
# "phone" β 384d, INT4, 1.5GB budget, model swapping, MiniLM text
# "edge_gpu" β 512d, INT8, 3.5GB budget, all models loaded
# "workstation" β 768d, FP16, 8GB budget, no quantization
# "raspberry_pi" β 256d, INT4, 1GB budget, model swapping, no audio
# "jetson" β 512d, FP16, 3GB budgetπ Quick Start
Install
pip install torch transformers Pillow numpy
# Optional: pip install onnxruntime onnx librosa soundfile opencv-python pymupdfEmbed Anything
from edge_multimodal_embeddings import EdgeMultimodalEmbedder
embedder = EdgeMultimodalEmbedder.from_profile("edge_gpu")
# Text
text_emb = embedder.embed("The quick brown fox")
# Image (auto-detected from file extension)
img_emb = embedder.embed("photo.jpg")
# Video
vid_emb = embedder.embed("clip.mp4")
# Audio
audio_emb = embedder.embed("song.wav")
# PDF (extracts text + images, embeds both)
pdf_emb = embedder.embed("paper.pdf")
# Code
code_emb = embedder.embed("main.py", modality="document")
# Cross-modal similarity
similarity = embedder.similarity(text_emb, img_emb)
print(f"Text-Image similarity: {similarity:.4f}")Batch Embedding
# Mixed modality batch
inputs = ["Hello world", "photo.jpg", "clip.mp4"]
embeddings = embedder.embed_batch(inputs)
# β shape: (3, 512)Phone Deployment
# Step 1: Export on workstation
embedder = EdgeMultimodalEmbedder.from_profile("phone")
embedder.export_onnx("./phone_models")
# Step 2: Run on phone (only needs onnxruntime + numpy)
from edge_multimodal_embeddings.runtime import EdgeRuntime
runtime = EdgeRuntime("./phone_models", num_threads=4)
emb = runtime.embed_text("Find photos of cats")
img_emb = runtime.embed_image("cat.jpg")
sim = runtime.similarity(emb, img_emb)π§ Custom Configuration
from edge_multimodal_embeddings.core.config import (
EmbedderConfig, ModelConfig, QuantizationMode
)
config = EmbedderConfig(
unified_dim=512,
device="cuda",
quantization=QuantizationMode.INT8_DYNAMIC,
max_memory_mb=3500,
lazy_load=True,
model_swapping=False,
max_batch_size=16,
# Swap in different models
text_model=ModelConfig(
model_id="sentence-transformers/all-MiniLM-L6-v2",
embedding_dim=384,
max_input_size=512,
),
# Disable modalities you don't need
audio_model=ModelConfig("laion/clap-htsat-unfused", 512, enabled=False),
video_model=ModelConfig("MCG-NJU/videomae-small-finetuned-kinetics", 384, enabled=False),
)
embedder = EdgeMultimodalEmbedder(config)π Supported File Formats
Plus: PIL Images, numpy arrays, torch tensors, base64 strings, URLs, raw bytes.
π§ Architecture Details
Why This Model Stack?
- Nomic text + vision share a native 768d embedding space β no projection needed for textβimage retrieval, highest accuracy per byte.
- CLAP is the only Apache-2.0 audio-text model with proper zero-shot capabilities.
- VideoMAE-Small is remarkably tiny (84MB) while achieving 79% top-1 on Kinetics-400.
- Projection heads are single Linear + LayerNorm layers (~0.5MB each) that unify all spaces to 512d.
Quantization Strategy
Based on MobileQuant (2408.13933) and EdgeVL (2403.04908):
- Dynamic INT8 (default for edge): Weights quantized, activations computed in FP32. <1% cosine similarity loss. No calibration needed.
- Static INT8: Both weights and activations quantized. Needs calibration data. Best throughput.
- INT4: 8Γ memory reduction. 1-3% quality loss. For extreme edge (phones, RPi).
- Per-tensor quantization for mobile NPU compatibility (not per-token).
Memory Management
- Lazy loading: Models loaded on first use, not at init
- LRU swapping: When memory budget is exceeded, least-recently-used model is evicted
- Adaptive batching: Automatically reduces batch size on OOM and retries
- Real-time monitoring: Track memory usage via
embedder.get_status()
π± Mobile/Edge Deployment
Android (ONNX Runtime Mobile)
// build.gradle
implementation 'com.microsoft.onnxruntime:onnxruntime-mobile:1.17.0'
// Load model
val session = env.createSession(modelBytes, SessionOptions().apply {
setIntraOpNumThreads(4)
addNnapi() // Use Android Neural Networks API
})iOS (CoreML)
// Use ONNX β CoreML conversion or CoreML EP
let config = MLModelConfiguration()
config.computeUnits = .all // CPU + GPU + ANERaspberry Pi / ARM Linux
pip install onnxruntime # ARM64 wheel available
python -c "from edge_multimodal_embeddings.runtime import EdgeRuntime; ..."π Benchmarks
π CLI
# Embed
edge-embed embed "Hello world"
edge-embed embed photo.jpg
edge-embed embed --modality video clip.mp4
# Compare
edge-embed compare "a cat" "a kitten"
# Benchmark
edge-embed benchmark -n 100
# Export
edge-embed export --format onnx -o ./exported/
# System info
edge-embed infoπ License
Apache-2.0. All component models are Apache-2.0 or MIT licensed.
π Credits
- Nomic AI β nomic-embed-text/vision-v1.5
- LAION β CLAP audio embeddings
- MCG-NJU β VideoMAE
- MobileQuant β Mobile quantization research
- EdgeVL β Edge visual-language models
- MobileCLIP β Architecture inspiration
