Team Ai
Modelpublic

Asher-1/map_anything_vggt

sourceHugging Faceupdated 9d agoView on Hugging Face
0likes453downloads
Model Card

Github: https://github.com/Asher-1/map-anything-ggml

MODEL CARDS — all six supported models (ckpts, GGUF, accuracy, speed)

This directory (cpp_ggml/models/) holds the official torch checkpoints (pytorch/<model>/, conversion-only) and all 32 GGUF files (gguf/, the C++ runtime input: 6 architectures, each f32/f16/q80 plus ONE measured K-quant — q6K for omega/pi3x/mapanything/dust3r, q5K for pi3/vggt-1b (the tier that wins on that backbone's official-protocol data). This document covers per-model use cases, size/VRAM, official-protocol accuracy and measured speed for **every** supported model, with the measured charts embedded inline. All figures (raw links) and all referenced documents (blob links) use **absolute GitHub URLs** (branch `main`), so the document renders anywhere — no checkout required: `https://github.com/Asher-1/map-anything-ggml/raw/main/cppggml/benchmarks/charts/...`

1. The six models at a glance

Accuracy = official MapAnything ETH3D protocol (130 sets, seed 777, 2 views, two-sided vs the official PyTorch; lower is better). Speed = CUDA latency vs the official torch fp32 baseline (RTX 4090). Gate = quantization x backend parity matrix.

ModelFamily / paradigmpointmaps↓depth↓ATE↓rot°↓metric scale↓speed vs torchgate
vggt-omega (default)vggt (N-view)0.03020.02080.00650.560.762CUDA 0.85x4q x 3b all PASS (+ text cos 1.0)
pi3xpi3 + metric scale0.05330.03690.01622.110.251CUDA 1.43-1.48x12/12
mapanythingunified (N-view, AAT)0.05530.04690.01240.920.162CUDA 1.79-2.28x12/12
pi3pi3 (N-view)0.06500.04850.01721.820.733CUDA ~1.0x15/15
vggt-1bvggt (N-view)0.06740.05410.02082.280.788CUDA 1.60-1.69x15/15
dust3rpair-wise ancestor0.10080.10330.05092.770.897CUDA 2.02-2.06x12/12

How to read the figures embedded in every card below:

  • —e2e_latency_bar — steady-state P50 latency (log axis): official torch fp32 vs cpp f32/f16/q80/q6K/q5_K, per backend;
  • —quant_pareto_3d — file size vs pose error vs depth error per quant;
  • —parity_scatter — cpp f16 vs torch f32 per-point parity (random frames);
  • —pose_error_heatmap — gate-matrix error heatmap (quants x backends);
  • —recon_depth/pointcloud/metrics_comparison — real-scene ETH3D courtyard reconstruction (GT / PyTorch / cpp quants side by side, avg_dis scale alignment, view0-frame point clouds).

2. Quick selection

Your scenarioPickWhy
>= 6 GB VRAM, accuracy first512-f16flagship resolution, closest to PyTorch on every metric
4-6 GB VRAM / best value512-q8_0half the size, fastest (87 ms), near-f16 accuracy
~1 GB VRAM / balanced sweet spot512-q6_K1.05 GB; ETH3D depth AbsRel matches torch f32 exactly, ATE even beats it (AUC5 76.5)
Metric scale + camera posesmapanything-f16metric_scale 0.162 far ahead; also the fastest large model on CUDA
Fastest CUDA inferencedust3r-q8_084.9 ms pair-wise (1.75x vs official torch)
Reproduce the paper's 416-reproduce numbers416-reproduce-f16official reproduction ckpt
Low-res inputs / text-alignment research256-text-f16256-resolution ckpt (C++ builds the TextAlignmentHead branch and emits .text_embedding.bin)
Two-view pointmaps / family-ancestor researchdust3r-f16pair-wise (S=2), no pose head, best quant robustness
No GPUany + --backend cpuq8_0 fastest (12.3 s on omega-512)

3. The model cards

3.1 vggt-omega-1b-512 (flagship, default)

  • —ckpt: pytorch/vggt-omega/vggt_omega_1b_512.pt (official facebook/VGGT-Omega release)
  • —Default inference resolution: 512 (balanced mode adapts the token budget to each input's aspect ratio)
  • —Use cases: general scene reconstruction (indoor/outdoor, 2-100 views), depth estimation, camera pose estimation
  • —GGUF: f32 4.36 GB / f16 2.18 GB / q80 1.26 GB / q6K 1.05 GB
  • —Official protocol accuracy (ETH3D, 130 sets, official dataset + 8 metrics):
MetricPyTorch f32f16q8_0q6_KOfficial ref*
Pose AUC@5 up (x100)79.2379.0877.0876.4679.53
Pose ATE RMSE down0.006550.006620.006480.006440.00999
Depth AbsRel down0.02080.02090.02100.02070.0204
Point AbsRel down0.03020.03020.03050.03020.0263

\* Official reproduction.md (retrained ckpt, 2-100 views); slightly wider protocol, order-of-magnitude reference only.

  • —Speed (512x512x2, steady-state pure inference, RTX 4090):
FormatPyTorch f32f32f16q8_0q6_K
CUDA65.0ms109.076.270.775.5
Vulkan——94.797.3101.5
CPU—1868217.8s12.3s—

(2026-10-02 re-measure, 10 repeats, torch baseline re-run in the same session; CPU rows keep the 09-24 exclusive-CPU numbers)

omega latency

omega quant pareto

omega parity

omega pose heatmap

omega recon depth

omega recon pointcloud

omega recon metrics

3.2 vggt-omega-1b-416-reproduce (paper reproduction)

  • —ckpt: pytorch/vggt-omega/vggt_omega_1b_416_reproduce.pt
  • —Default resolution: 416
  • —Use case: reproducing the official paper/tech-report evaluation numbers (this is the ckpt behind reproduction.md)
  • —GGUF: f32 4.36 GB / f16 2.18 GB / q80 1.26 GB / q6K 1.05 GB
  • —Official protocol: see `benchmarks/results/vggt-omega/eval_eth3d_416_*.md` and `eval_eth3d_matrix.md`; accuracy families match PyTorch.

3.3 vggt-omega-1b-256-text (low-res + alignment head)

  • —ckpt: pytorch/vggt-omega/vggt_omega_1b_256_text.pt
  • —Default resolution: 256
  • —Use cases: low-resolution inputs, VRAM-constrained devices, text-alignment research
  • —Note: since 2026-09-24 the C++ graph implements the official TextAlignmentHead (enabled by the GGUF vggt.enable_text_alignment flag): the language-aligned embedding is emitted as .text_embedding.bin (2048-dim, L2-normalized). Gate: 256-text f16 PASS with text cosine = 1.000000 vs the official torch head (see `benchmarks/results/gate_matrix.md`)
  • —GGUF: f32 5.15 GB / f16 2.58 GB / q80 1.47 GB / q6K 1.22 GB (the text head adds ~0.4 GB of f32 extras over the 512 variant)
  • —Official protocol: `benchmarks/results/vggt-omega/eval_eth3d_text_*.md` (the text_torch column previously had a wrong protocol; all columns now use the real 256-text weights)

3.4 vggt-1b (official VGGT-1B)

  • —ckpt: pytorch/vggt1b/model.pt (official facebook/VGGT-1B)
  • —Default resolution: 518 (patch 14)
  • —Use cases: N-view reconstruction with the classic VGGT heads; the only model besides omega emitting a dedicated world pointmap + conf (.points.bin / .points_conf.bin) alongside pose/depth
  • —GGUF: f32 4.54 GB / f16 2.43 GB / q80 1.43 GB / q5K 1.04 GB
  • —Gate: 12/12 PASS (f32/f16/q80/q5K x CPU/CUDA/Vulkan; f16 pose max 0.0006-0.0025, depth medrel 0.06-0.29%; q5K kept — its AUC5 63.7 is the best of the three measured quants; q6K removed 2026-10-01 — the weakest tier on 5 of 7 metrics, dominated by q80/q5K — [`benchmarks/results/gatematrix.md`](https://github.com/Asher-1/map-anything-ggml/blob/main/cppggml/benchmarks/results/gatematrix.md))
  • —Official protocol (two-sided, 130 sets): `benchmarks/results/vggt-1b/bench_official_eth3d_vggt.md`
  • —Real-scene reconstruction (courtyard, window [5,0], 518x336):
metrictorchf16q8_0q5_K
AbsRel0.0211580.0209730.0215420.015032
d10.9825000.9829110.9822490.996840
chamfer0.0134760.0134420.0136870.012912
pose rot diff (deg)—0.0140.0220.043
pose trans diff (m)—0.00020.00030.0016
  • —Speed (518x518x2, P50, RTX 4090; re-measured 2026-09-24, exclusive):
Backendtorch fp32f32f16q8_0q5_K
CUDA284.0ms242.2 (1.17x)177.1 (1.60x)168.1 (1.69x)171.7 (1.65x)
Vulkan—221.2 (1.28x)193.8 (1.47x)194.3 (1.46x)198.0 (1.43x)
CPU12738ms12929 (0.99x)11834 (1.08x)12214 (1.04x)13832 (0.92x)

vggt1b recon depth

vggt1b recon pointcloud

vggt1b recon metrics

vggt1b latency

vggt1b quant pareto

vggt1b parity

3.5 pi3

  • —ckpt: pytorch/pi3/ (official yyfz233/Pi3, model.safetensors)
  • —Default resolution: 518 (patch 14)
  • —Use cases: N-view reconstruction with the simplified pi3 head set; poses are row-major c2w 4x4 (.pose.bin), depth = local_points camera-z
  • —GGUF: f32 3.66 GB / f16 1.84 GB / q80 981 MB / q5K 639 MB
  • —Gate: 12/12 PASS (f32/f16/q80/q5K x CPU/CUDA/Vulkan; f16 localpoints medrel 0.0004-0.0037; q5K is pi3's best depth/rot/ATE/ pointmaps tier — 0.046483 / 1.086°; q6K removed 2026-10-01 — no best metric on the 130 sets, squeezed between q5K and q80)
  • —Official protocol (130 sets, two-sided): metric-level parity — `benchmarks/results/pi3/bench_official_eth3d_pi3.md`
  • —Paper protocol (13 scenes, Acc/Comp/NC + Umeyama/ICP): torch/cpp deltas <= 0.01 on every metric; compmed 0.1251/0.1236 pinned to the paper's 0.128 — [`benchmarks/results/pi3/evalpi3papereth3d.md`](https://github.com/Asher-1/map-anything-ggml/blob/main/cppggml/benchmarks/results/pi3/evalpi3papereth3d.md)
  • —Real-scene reconstruction (courtyard, window [5,0], 518x336):
metrictorchf16q8_0q5_K
AbsRel0.0137300.0135430.0140920.014039
d10.9991740.9991590.9991660.999069
chamfer0.0108940.0107210.0109940.010034
pose rot diff (deg)—0.0000.1211.209
pose trans diff (m)—0.0020.0050.015
  • —Speed (518x518x2, P50, RTX 4090; CPU re-measured exclusively 2026-09-24):
Backendtorch fp32f32f16q8_0q5_K
CUDA237.4ms348.5242.0226.4 (1.05x)234.0
Vulkan—420.3368.7358.2400.5
CPU13744ms10774 (1.28x)10249 (1.34x)11821 (1.16x)11932 (1.15x)

pi3 recon depth

pi3 recon pointcloud

pi3 recon metrics

pi3 latency

pi3 quant pareto

pi3 parity

3.6 pi3x (pi3 + metric scale)

  • —ckpt: pytorch/pi3x/ (official yyfz233/Pi3X, model.safetensors)
  • —Default resolution: 518 (patch 14)
  • —Use cases: pi3 plus a metric-scale head (.scale.bin); balanced alternative with a pinned paper protocol
  • —GGUF: f32 3.79 GB / f16 1.94 GB / q80 1.08 GB / q6K 894 MB
  • —Gate: 12/12 PASS (f32/f16/q80/q6K x CPU/CUDA/Vulkan; thresholds calibrated to pi3x's own f16-KV noise floor, q6K lp medrel 0.033-0.055). q5K was removed on 2026-10-01 — its ConvHead chain is sensitive to 4.5-bit weights and the fleet confirmation shows q6K wins AUC5 by +3.7 points at the same size class (q8_0 is the recommended quant)
  • —Official protocol (130 sets, 518x336, two-sided; per-quant cpp-only numbers from results/FLEETQUANTCONFIRMATION.md):
Metrictorch f32cpp f16cpp q8_0cpp q6_K
pointmapsabsrel0.0533400.0534480.0537490.053740
zdepthabs_rel0.0369220.0369270.0370980.036885
poseatermse0.0162320.0162800.0162430.016311
poseauc5 (x100)66.9267.0867.3866.15
roterrdeg2.1082.1132.1212.121
metricscaleabs_rel0.2507330.2510580.2524370.251801
  • —Paper protocol (13 scenes): torch/cpp deltas <= 0.0021 on every metric
  • —Real-scene reconstruction (courtyard, window [5,0], 518x336):
metrictorchf16q8_0q6_K
AbsRel0.0157790.0157740.0163170.015737
d10.9993680.9993680.9993680.999346
chamfer0.0118140.0117350.0118970.011826
fscore0.9816320.9819980.9819980.981659
pose rot diff (deg)—0.0280.0880.190
pose trans diff (m)—0.0050.0320.011
  • —Speed (518x518x2, P50, RTX 4090 — faster than official torch across all three backends; re-measured 2026-09-24, exclusive):
Backendtorch fp32torch TF32f32f16q8_0q6_K
CUDA290.3ms242.2252.5202.5 (1.43x)195.7 (1.48x)202.1 (1.44x)
Vulkan——237.4229.8 (1.26x)232.9 (1.25x)237.8 (1.22x)
CPU17526ms—13986 (1.25x)12959 (1.35x)14891 (1.18x)—

pi3x recon depth

pi3x recon pointcloud

pi3x recon metrics

pi3x latency

pi3x quant pareto

pi3x parity

3.7 mapanything (unified, AAT)

  • —ckpt: pytorch/mapanything/ (official facebook/map-anything, model.safetensors + config.json for from_pretrained)
  • —Default resolution: 518 (patch 14)
  • —Use cases: unified N-view model with alternating-attention tokens; emits rays + non-ambiguous mask (.mask.bin) + metric scale (.scale.bin) on top of the pi3-family contract — the pick for metric scale and camera poses
  • —GGUF: f32 4.44 GB / f16 2.81 GB / q80 1.26 GB / q6K 1.05 GB
  • —Gate: 9/9 PASS (f16/q80/q6K x CPU/CUDA/Vulkan, f16 rays medrel ~0.001-0.003; the f32 GGUF exists for the latency baseline column; q5K retired 2026-10-01, HF-only)
  • —Official protocol (MapAnything paper protocol = its paper protocol, 130 sets, two-sided): deltas <= 0.3% on every metric (pointmaps 0.05334/0.05345, AUC@5 66.9/67.1, metricscale 0.2507/0.2511) — [`benchmarks/results/mapanything/benchofficialeth3dmapanything.md`](https://github.com/Asher-1/map-anything-ggml/blob/main/cppggml/benchmarks/results/mapanything/benchofficialeth3dmapanything.md)
  • —Real-scene reconstruction (courtyard): AbsRel 0.077170/0.077287/0.077505 (f16/q80/q6K) vs torch 0.077106 (delta < 0.13%); pose rot diff 0.010-0.030°
  • —Speed (518x518x2, P50, RTX 4090; re-measured 2026-09-24, exclusive):
Backendtorch fp32f32f16q8_0q6_K
CUDA244.0ms174.2 (1.40x)136.0 (1.79x)106.9 (2.28x)114.0 (2.14x)
Vulkan—145.6137.8 (1.77x)139.8 (1.75x)146.9 (1.66x)
CPU12814ms11515 (1.11x)11047 (1.16x)10813 (1.19x)—

mapanything recon depth

mapanything recon pointcloud

mapanything recon metrics

mapanything latency

mapanything quant pareto

mapanything parity

3.8 dust3r-512_dpt (pair-wise ancestor, M5)

  • —ckpt: pytorch/dust3r/ (model.safetensors + config.json; official naver/DUSt3RViTLargeBaseDecoder512dpt, HF PyTorchModelHubMixin layout, no .pth in the repo)
  • —Default resolution: 512 (patch 16; pair-wise — exactly S=2 inputs)
  • —Use cases: two-view pointmap regression, the family's common ancestor; both pointmaps are emitted in view1's camera frame (s0 = head1 self-view, s1 = head2 other-view); conf = 1+exp; no pose (the model has no pose head — the official poses come from the out-of-network global-alignment optimizer, not ported; the CLI emits no .pose.bin, the bench recovers poses via closed-form Procrustes)
  • —GGUF: f32 2.18 GB / f16 1.17 GB / q80 699 MB / q6K 604 MB
  • —Gate (fixed 512x512 S=2 frames, torch f32 ref): 12/12 PASS (f32/f16/q80/q6K x CPU/CUDA/Vulkan); lp medrel 0.0002 (f16/f32) -> 0.0016 (q80) -> 0.0029-0.0039 (q6K); the most quant-robust architecture in the repo. q5K was removed on 2026-10-01 — the fleet confirmation shows q6K beats it on EVERY official metric (depth/ATE/ AUC5/rot/pointmaps/scale; details below) ([`benchmarks/results/gatematrix.md`](https://github.com/Asher-1/map-anything-ggml/blob/main/cppggml/benchmarks/results/gatematrix.md))
  • —Official protocol (130 sets x 2 views, seed 777, 512x336, two-sided; pose via closed-form Procrustes — the official protocol itself is num_views=2):
metrictorch f32cpp f16rel delta
pointmapsabsrel0.1008100.1008550.04%
zdepthabs_rel0.1032790.1033100.03%
poseatermse0.0509410.0509550.03%
roterrdeg2.7732.7261.7%
rotauc30 (x100)56.4456.360.14%
raydirserr_deg2.42152.42130.01%

Full table: `benchmarks/results/dust3r/bench_official_eth3d_dust3r.md`

Per-quant cpp-only (130 sets, `benchmarks/results/FLEET_QUANT_CONFIRMATION.md`): q6_K wins depth/ATE/AUC5/rot/pointmaps/scale over both q8_0 and q5_K (zdepth 0.102975 vs q80 0.103174 / q5K 0.103523; ATE 0.050718 vs 0.050953 / 0.051103; AUC5 8.00 vs 7.38 / 7.54) — the 604 MB q6K is the best accuracy point of the whole tier ladder; q5_K was removed on 2026-10-01 (dominated).

  • —Real-scene reconstruction (courtyard, window [5,0], 512x336): cpp f16 vs torch AbsRel 0.132052 vs 0.132059 (delta 0.005%); pose delta rot 0.000° / trans 0.00003 m. q80/q6K pose deltas 0.020° / 0.024°; q6_K also edges out every tier on AbsRel/chamfer (0.131862/0.081195).
  • —Speed (512x512x2, RTX 4090; CPU re-measured exclusively 2026-09-24; `benchmarks/results/dust3r/speed_dust3r.md`):
Backendtorch fp32f32f16q8_0q6_K
CUDA148.2ms84.2 (1.76x)72.7 (2.04x)71.8 (2.06x)73.2 (2.02x)
Vulkan—66.964.1 (2.31x)65.4 (2.27x)69.7 (2.13x)
CPU5559ms5233 (1.06x)5068 (1.10x)5732 (0.97x)—

dust3r recon depth

dust3r recon pointcloud

dust3r recon metrics

dust3r parity

dust3r latency

dust3r quant pareto

4. Quantization formats

FormatRelative sizePositioningAccuracy
f322x f16reference baseline (numeric oracle)bit-locked to torch weights
f161xaccuracy-first defaultpose 0.0234 / depth 0.17% (gate)
q8_00.57xnear-f16 accuracy, fastestpose 0.0141 / depth 0.82%
q6_K0.48x~1 GB sweet spot — ETH3D depth/ATE at torch-f32 levelpose 0.1160 / depth 0.53% (gate)

\ Gate numbers = CUDA parity on the canonical matrix frames (2026-09-30 re-run, gate_summary.json). On the official ETH3D 130 sets: q6_K matches torch f32 depth AbsRel exactly (0.020748 vs 0.020752) and beats it on ATE (0.006436 vs 0.006549)*; q80/f16 stay <= 1% relative. K-quant assignments are strictly data-driven (per-model 130-set verdicts, results/FLEETQUANT_CONFIRMATION.md — the quant response is NOT monotonic in bit width, so each model ships exactly ONE measured K tier):

  • —q6_K for omega (2026-09-30): dominates q5K (+22% size buys AUC5 71.7 -> 76.5 and 3-6x tighter parity floors; the q5K loss is inherent to 4.5-bit weight rounding, identical in any runtime), same judgment as the q4_K removal;
  • —q6_K for pi3x / dust3r (2026-10-01): wins AUC5 by +3.7 on pi3x and sweeps every official metric on dust3r;
  • —q6_K for mapanything (2026-10-01): wins depth/rot/pointmaps/AUC5 (q5_K only edges metric scale by 1%);
  • —q5_K for pi3 (2026-10-01): best depth/rot/ATE/pointmaps tier (0.046483 / 0.018289 / 1.086° / 0.065244); pi3's q6_K had NO best metric and was removed (locally + HF);
  • —q5_K for vggt-1b (2026-10-01): best AUC5 (63.7 vs q80 63.1); its q6K was the weakest tier on 5 of 7 metrics and was removed (locally + HF). Removed tiers remain reproducible from the git history of RESULTS.md / FLEETQUANTCONFIRMATION.md / the converters.
  • —q4K was removed on 2026-09-18 (not a valid Pareto point), see [`benchmarks/RESULTS.md`](https://github.com/Asher-1/map-anything-ggml/blob/main/cppggml/benchmarks/RESULTS.md) §6.
  • —Quantization applies only to 2-D Linear weights; LayerNorm/conv/RoPE stay floating point.

quant pareto

5. Output directory conventions (vs the official demo)

SideCommandOutput directory/files
Official PyTorch demodemo_gradio.py (default config)demo_outputs/input_images_<ts>/predictions.npz + GLB
C++ CLIvggt-cli --images ... --out-prefix PP.pose.bin (S,9) / P.depth.bin (S,H,W) / P.depth_conf.bin (S,H,W) / P.meta.json
Official npz packagingscripts/make_predictions_npz.py --prefix PP.predictions.npz (official keys incl. extrinsic/intrinsic/worldpointsfrom_depth)
Dual-runtime comparisonscripts/compare_official_demo.py<out>/torch/predictions.npz vs <out>/cpp/out.*.bin + COMPARISON.md

Key mapping: pose_enc ↔ .pose.bin, depth[...,0] ↔ .depth.bin, depth_conf ↔ .depth_conf.bin; extrinsic/intrinsic/world_points_from_depth are derived from pose_enc+depth with the official formulas (built into both helper scripts).

6. One-click reproduction

bash
./run_mapggml.sh demo                 # newcomers: download -> convert -> build -> infer
./run_mapggml.sh gate                 # dual-runtime parity gate
./run_mapggml.sh bench --backend cuda # steady-state latency matrix
PYTHONPATH=/tmp/torch_cuda_lib:<repo> python3 cpp_ggml/scripts/compare_official_demo.py \
    --images a.jpg b.jpg --out-dir /tmp/demo_compare   # official demo dual-runtime compare