b4ph/mlcd-mteb-cifar-eval
MLCD vs CLIP on MTEB CIFAR-10/100: integration and evaluation Evaluation results accompanying the MTEB integration of two MLCD image encoders (PR #5406, resolving issue #2571). Two DeepGlint-AI MLCD encoders were integrated into MTEB, verified against the reference implementation, and evaluated on the official MTEB CIFAR-10/CIFAR-100 image-classification tasks alongside size-matched OpenAI CLIP baselines. What was measured Official MTEB image classification: 5… See the full description on the dataset page: https://huggingface.co/datasets/b4ph/mlcd-mteb-cifar-eval.
MLCD vs CLIP on MTEB CIFAR-10/100: integration and evaluation
Evaluation results accompanying the MTEB integration of two MLCD image encoders (PR #5406, resolving issue #2571).
Two DeepGlint-AI MLCD encoders were integrated into MTEB, verified against the reference implementation, and evaluated on the official MTEB CIFAR-10/CIFAR-100 image-classification tasks alongside size-matched OpenAI CLIP baselines.
What was measured
Official MTEB image classification: 5 experiments x 16 train samples per class, full 10k test split, fp32, RTX 5090. Each model's embedding is the pooled CLS token after post_layernorm (the public MLCD checkpoints contain only the vision tower; weight-inspection verified).
Accuracy, 5-repetition mean:
MLCD beat its size-matched CLIP baseline in all four predeclared comparisons (paired bootstrap, 10,000 draws, seed 20260907; Bonferroni-adjusted 98.75% two-sided intervals all above zero): CIFAR-10 +4.19 / +2.84 pp, CIFAR-100 +9.39 / +10.47 pp. These are checkpoint comparisons (training data and recipe also differ), on two related small-image datasets. They do not isolate the MLCD training objective and do not establish broad vision-task superiority.
Files
cell_scores.csv: one row per model/task: aggregate accuracy, all 5 per-experiment accuracies, wall time, peak VRAM.comparisons.csv: the 4 paired comparisons with 95% and Bonferroni 98.75% bootstrap intervals.timing.csv: matched-conditions throughput and VRAM (3 warmup batches, 5 counterbalanced blocks, fixed score-blind subset, batch 64).
What this is not
The authors' published linear-probe scores (86.8 base / 93.69 large on CIFAR-100) are not reproduced and could not be: their probe used a trained embedding head that is absent from the public HF checkpoints. These numbers are independent evaluations under MTEB's few-shot protocol, not reproductions.
Reproducibility
Pinned model revisions, dataset revisions, code fingerprints, per-experiment predictions, and the full evidence chain are archived in the contributor's project directory; the adapter and runner are included in the linked PR.
