Team Ai
Datasetpublic

b4ph/mlcd-mteb-cifar-eval

MLCD vs CLIP on MTEB CIFAR-10/100: integration and evaluation Evaluation results accompanying the MTEB integration of two MLCD image encoders (PR #5406, resolving issue #2571). Two DeepGlint-AI MLCD encoders were integrated into MTEB, verified against the reference implementation, and evaluated on the official MTEB CIFAR-10/CIFAR-100 image-classification tasks alongside size-matched OpenAI CLIP baselines. What was measured Official MTEB image classification: 5… See the full description on the dataset page: https://huggingface.co/datasets/b4ph/mlcd-mteb-cifar-eval.

sourceHugging Faceapache-2.0updated 29d agoView on Hugging Face
0likes85downloads
Dataset Card

MLCD vs CLIP on MTEB CIFAR-10/100: integration and evaluation

Evaluation results accompanying the MTEB integration of two MLCD image encoders (PR #5406, resolving issue #2571).

Two DeepGlint-AI MLCD encoders were integrated into MTEB, verified against the reference implementation, and evaluated on the official MTEB CIFAR-10/CIFAR-100 image-classification tasks alongside size-matched OpenAI CLIP baselines.

What was measured

Official MTEB image classification: 5 experiments x 16 train samples per class, full 10k test split, fp32, RTX 5090. Each model's embedding is the pooled CLS token after post_layernorm (the public MLCD checkpoints contain only the vision tower; weight-inspection verified).

Accuracy, 5-repetition mean:

TaskMLCD baseCLIP baseMLCD largeCLIP large
CIFAR-1094.2990.1098.2895.44
CIFAR-10077.1867.7989.5079.04

MLCD beat its size-matched CLIP baseline in all four predeclared comparisons (paired bootstrap, 10,000 draws, seed 20260907; Bonferroni-adjusted 98.75% two-sided intervals all above zero): CIFAR-10 +4.19 / +2.84 pp, CIFAR-100 +9.39 / +10.47 pp. These are checkpoint comparisons (training data and recipe also differ), on two related small-image datasets. They do not isolate the MLCD training objective and do not establish broad vision-task superiority.

Files

  • —cell_scores.csv: one row per model/task: aggregate accuracy, all 5 per-experiment accuracies, wall time, peak VRAM.
  • —comparisons.csv: the 4 paired comparisons with 95% and Bonferroni 98.75% bootstrap intervals.
  • —timing.csv: matched-conditions throughput and VRAM (3 warmup batches, 5 counterbalanced blocks, fixed score-blind subset, batch 64).

What this is not

The authors' published linear-probe scores (86.8 base / 93.69 large on CIFAR-100) are not reproduced and could not be: their probe used a trained embedding head that is absent from the public HF checkpoints. These numbers are independent evaluations under MTEB's few-shot protocol, not reproductions.

Reproducibility

Pinned model revisions, dataset revisions, code fingerprints, per-experiment predictions, and the full evidence chain are archived in the contributor's project directory; the adapter and runner are included in the linked PR.