Asher-1/map_anything_vggt
Github: https://github.com/Asher-1/map-anything-ggml
MODEL CARDS — all six supported models (ckpts, GGUF, accuracy, speed)
This directory (cpp_ggml/models/) holds the official torch checkpoints (pytorch/<model>/, conversion-only) and all 32 GGUF files (gguf/, the C++ runtime input: 6 architectures, each f32/f16/q80 plus ONE measured K-quant — q6K for omega/pi3x/mapanything/dust3r, q5K for pi3/vggt-1b (the tier that wins on that backbone's official-protocol data). This document covers per-model use cases, size/VRAM, official-protocol accuracy and measured speed for **every** supported model, with the measured charts embedded inline. All figures (raw links) and all referenced documents (blob links) use **absolute GitHub URLs** (branch `main`), so the document renders anywhere — no checkout required: `https://github.com/Asher-1/map-anything-ggml/raw/main/cppggml/benchmarks/charts/...`
1. The six models at a glance
Accuracy = official MapAnything ETH3D protocol (130 sets, seed 777, 2 views, two-sided vs the official PyTorch; lower is better). Speed = CUDA latency vs the official torch fp32 baseline (RTX 4090). Gate = quantization x backend parity matrix.
How to read the figures embedded in every card below:
e2e_latency_bar— steady-state P50 latency (log axis): official torch fp32 vs cpp f32/f16/q80/q6K/q5_K, per backend;quant_pareto_3d— file size vs pose error vs depth error per quant;parity_scatter— cpp f16 vs torch f32 per-point parity (random frames);pose_error_heatmap— gate-matrix error heatmap (quants x backends);recon_depth/pointcloud/metrics_comparison— real-scene ETH3D courtyard reconstruction (GT / PyTorch / cpp quants side by side, avg_dis scale alignment, view0-frame point clouds).
2. Quick selection
3. The model cards
3.1 vggt-omega-1b-512 (flagship, default)
- ckpt:
pytorch/vggt-omega/vggt_omega_1b_512.pt(official facebook/VGGT-Omega release) - Default inference resolution: 512 (balanced mode adapts the token budget to each input's aspect ratio)
- Use cases: general scene reconstruction (indoor/outdoor, 2-100 views), depth estimation, camera pose estimation
- GGUF: f32 4.36 GB / f16 2.18 GB / q80 1.26 GB / q6K 1.05 GB
- Official protocol accuracy (ETH3D, 130 sets, official dataset + 8 metrics):
\* Official reproduction.md (retrained ckpt, 2-100 views); slightly wider protocol, order-of-magnitude reference only.
- Speed (512x512x2, steady-state pure inference, RTX 4090):
(2026-10-02 re-measure, 10 repeats, torch baseline re-run in the same session; CPU rows keep the 09-24 exclusive-CPU numbers)







3.2 vggt-omega-1b-416-reproduce (paper reproduction)
- ckpt:
pytorch/vggt-omega/vggt_omega_1b_416_reproduce.pt - Default resolution: 416
- Use case: reproducing the official paper/tech-report evaluation numbers (this is the ckpt behind reproduction.md)
- GGUF: f32 4.36 GB / f16 2.18 GB / q80 1.26 GB / q6K 1.05 GB
- Official protocol: see `benchmarks/results/vggt-omega/eval_eth3d_416_*.md` and `eval_eth3d_matrix.md`; accuracy families match PyTorch.
3.3 vggt-omega-1b-256-text (low-res + alignment head)
- ckpt:
pytorch/vggt-omega/vggt_omega_1b_256_text.pt - Default resolution: 256
- Use cases: low-resolution inputs, VRAM-constrained devices, text-alignment research
- Note: since 2026-09-24 the C++ graph implements the official TextAlignmentHead (enabled by the GGUF
vggt.enable_text_alignmentflag): the language-aligned embedding is emitted as.text_embedding.bin(2048-dim, L2-normalized). Gate: 256-text f16 PASS with text cosine = 1.000000 vs the official torch head (see `benchmarks/results/gate_matrix.md`) - GGUF: f32 5.15 GB / f16 2.58 GB / q80 1.47 GB / q6K 1.22 GB (the text head adds ~0.4 GB of f32 extras over the 512 variant)
- Official protocol: `benchmarks/results/vggt-omega/eval_eth3d_text_*.md` (the text_torch column previously had a wrong protocol; all columns now use the real 256-text weights)
3.4 vggt-1b (official VGGT-1B)
- ckpt:
pytorch/vggt1b/model.pt(official facebook/VGGT-1B) - Default resolution: 518 (patch 14)
- Use cases: N-view reconstruction with the classic VGGT heads; the only model besides omega emitting a dedicated world pointmap + conf (
.points.bin/.points_conf.bin) alongside pose/depth - GGUF: f32 4.54 GB / f16 2.43 GB / q80 1.43 GB / q5K 1.04 GB
- Gate: 12/12 PASS (f32/f16/q80/q5K x CPU/CUDA/Vulkan; f16 pose max 0.0006-0.0025, depth medrel 0.06-0.29%; q5K kept — its AUC5 63.7 is the best of the three measured quants; q6K removed 2026-10-01 — the weakest tier on 5 of 7 metrics, dominated by q80/q5K — [`benchmarks/results/gatematrix.md`](https://github.com/Asher-1/map-anything-ggml/blob/main/cppggml/benchmarks/results/gatematrix.md))
- Official protocol (two-sided, 130 sets): `benchmarks/results/vggt-1b/bench_official_eth3d_vggt.md`
- Real-scene reconstruction (courtyard, window [5,0], 518x336):
- Speed (518x518x2, P50, RTX 4090; re-measured 2026-09-24, exclusive):






3.5 pi3
- ckpt:
pytorch/pi3/(official yyfz233/Pi3,model.safetensors) - Default resolution: 518 (patch 14)
- Use cases: N-view reconstruction with the simplified pi3 head set; poses are row-major c2w 4x4 (
.pose.bin), depth = local_points camera-z - GGUF: f32 3.66 GB / f16 1.84 GB / q80 981 MB / q5K 639 MB
- Gate: 12/12 PASS (f32/f16/q80/q5K x CPU/CUDA/Vulkan; f16 localpoints medrel 0.0004-0.0037; q5K is pi3's best depth/rot/ATE/ pointmaps tier — 0.046483 / 1.086°; q6K removed 2026-10-01 — no best metric on the 130 sets, squeezed between q5K and q80)
- Official protocol (130 sets, two-sided): metric-level parity — `benchmarks/results/pi3/bench_official_eth3d_pi3.md`
- Paper protocol (13 scenes, Acc/Comp/NC + Umeyama/ICP): torch/cpp deltas <= 0.01 on every metric; compmed 0.1251/0.1236 pinned to the paper's 0.128 — [`benchmarks/results/pi3/evalpi3papereth3d.md`](https://github.com/Asher-1/map-anything-ggml/blob/main/cppggml/benchmarks/results/pi3/evalpi3papereth3d.md)
- Real-scene reconstruction (courtyard, window [5,0], 518x336):
- Speed (518x518x2, P50, RTX 4090; CPU re-measured exclusively 2026-09-24):






3.6 pi3x (pi3 + metric scale)
- ckpt:
pytorch/pi3x/(official yyfz233/Pi3X,model.safetensors) - Default resolution: 518 (patch 14)
- Use cases: pi3 plus a metric-scale head (
.scale.bin); balanced alternative with a pinned paper protocol - GGUF: f32 3.79 GB / f16 1.94 GB / q80 1.08 GB / q6K 894 MB
- Gate: 12/12 PASS (f32/f16/q80/q6K x CPU/CUDA/Vulkan; thresholds calibrated to pi3x's own f16-KV noise floor, q6K lp medrel 0.033-0.055). q5K was removed on 2026-10-01 — its ConvHead chain is sensitive to 4.5-bit weights and the fleet confirmation shows q6K wins AUC5 by +3.7 points at the same size class (q8_0 is the recommended quant)
- Official protocol (130 sets, 518x336, two-sided; per-quant cpp-only numbers from results/FLEETQUANTCONFIRMATION.md):
- Paper protocol (13 scenes): torch/cpp deltas <= 0.0021 on every metric
- Real-scene reconstruction (courtyard, window [5,0], 518x336):
- Speed (518x518x2, P50, RTX 4090 — faster than official torch across all three backends; re-measured 2026-09-24, exclusive):






3.7 mapanything (unified, AAT)
- ckpt:
pytorch/mapanything/(official facebook/map-anything,model.safetensors+config.jsonforfrom_pretrained) - Default resolution: 518 (patch 14)
- Use cases: unified N-view model with alternating-attention tokens; emits rays + non-ambiguous mask (
.mask.bin) + metric scale (.scale.bin) on top of the pi3-family contract — the pick for metric scale and camera poses - GGUF: f32 4.44 GB / f16 2.81 GB / q80 1.26 GB / q6K 1.05 GB
- Gate: 9/9 PASS (f16/q80/q6K x CPU/CUDA/Vulkan, f16 rays medrel ~0.001-0.003; the f32 GGUF exists for the latency baseline column; q5K retired 2026-10-01, HF-only)
- Official protocol (MapAnything paper protocol = its paper protocol, 130 sets, two-sided): deltas <= 0.3% on every metric (pointmaps 0.05334/0.05345, AUC@5 66.9/67.1, metricscale 0.2507/0.2511) — [`benchmarks/results/mapanything/benchofficialeth3dmapanything.md`](https://github.com/Asher-1/map-anything-ggml/blob/main/cppggml/benchmarks/results/mapanything/benchofficialeth3dmapanything.md)
- Real-scene reconstruction (courtyard): AbsRel 0.077170/0.077287/0.077505 (f16/q80/q6K) vs torch 0.077106 (delta < 0.13%); pose rot diff 0.010-0.030°
- Speed (518x518x2, P50, RTX 4090; re-measured 2026-09-24, exclusive):






3.8 dust3r-512_dpt (pair-wise ancestor, M5)
- ckpt:
pytorch/dust3r/(model.safetensors+config.json; official naver/DUSt3RViTLargeBaseDecoder512dpt, HF PyTorchModelHubMixin layout, no .pth in the repo) - Default resolution: 512 (patch 16; pair-wise — exactly S=2 inputs)
- Use cases: two-view pointmap regression, the family's common ancestor; both pointmaps are emitted in view1's camera frame (s0 = head1 self-view, s1 = head2 other-view); conf = 1+exp; no pose (the model has no pose head — the official poses come from the out-of-network global-alignment optimizer, not ported; the CLI emits no .pose.bin, the bench recovers poses via closed-form Procrustes)
- GGUF: f32 2.18 GB / f16 1.17 GB / q80 699 MB / q6K 604 MB
- Gate (fixed 512x512 S=2 frames, torch f32 ref): 12/12 PASS (f32/f16/q80/q6K x CPU/CUDA/Vulkan); lp medrel 0.0002 (f16/f32) -> 0.0016 (q80) -> 0.0029-0.0039 (q6K); the most quant-robust architecture in the repo. q5K was removed on 2026-10-01 — the fleet confirmation shows q6K beats it on EVERY official metric (depth/ATE/ AUC5/rot/pointmaps/scale; details below) ([`benchmarks/results/gatematrix.md`](https://github.com/Asher-1/map-anything-ggml/blob/main/cppggml/benchmarks/results/gatematrix.md))
- Official protocol (130 sets x 2 views, seed 777, 512x336, two-sided; pose via closed-form Procrustes — the official protocol itself is num_views=2):
Full table: `benchmarks/results/dust3r/bench_official_eth3d_dust3r.md`
Per-quant cpp-only (130 sets, `benchmarks/results/FLEET_QUANT_CONFIRMATION.md`): q6_K wins depth/ATE/AUC5/rot/pointmaps/scale over both q8_0 and q5_K (zdepth 0.102975 vs q80 0.103174 / q5K 0.103523; ATE 0.050718 vs 0.050953 / 0.051103; AUC5 8.00 vs 7.38 / 7.54) — the 604 MB q6K is the best accuracy point of the whole tier ladder; q5_K was removed on 2026-10-01 (dominated).
- Real-scene reconstruction (courtyard, window [5,0], 512x336): cpp f16 vs torch AbsRel 0.132052 vs 0.132059 (delta 0.005%); pose delta rot 0.000° / trans 0.00003 m. q80/q6K pose deltas 0.020° / 0.024°; q6_K also edges out every tier on AbsRel/chamfer (0.131862/0.081195).
- Speed (512x512x2, RTX 4090; CPU re-measured exclusively 2026-09-24; `benchmarks/results/dust3r/speed_dust3r.md`):






4. Quantization formats
\ Gate numbers = CUDA parity on the canonical matrix frames (2026-09-30 re-run, gate_summary.json). On the official ETH3D 130 sets: q6_K matches torch f32 depth AbsRel exactly (0.020748 vs 0.020752) and beats it on ATE (0.006436 vs 0.006549)*; q80/f16 stay <= 1% relative. K-quant assignments are strictly data-driven (per-model 130-set verdicts, results/FLEETQUANT_CONFIRMATION.md — the quant response is NOT monotonic in bit width, so each model ships exactly ONE measured K tier):
- q6_K for omega (2026-09-30): dominates q5K (+22% size buys AUC5 71.7 -> 76.5 and 3-6x tighter parity floors; the q5K loss is inherent to 4.5-bit weight rounding, identical in any runtime), same judgment as the q4_K removal;
- q6_K for pi3x / dust3r (2026-10-01): wins AUC5 by +3.7 on pi3x and sweeps every official metric on dust3r;
- q6_K for mapanything (2026-10-01): wins depth/rot/pointmaps/AUC5 (q5_K only edges metric scale by 1%);
- q5_K for pi3 (2026-10-01): best depth/rot/ATE/pointmaps tier (0.046483 / 0.018289 / 1.086° / 0.065244); pi3's q6_K had NO best metric and was removed (locally + HF);
- q5_K for vggt-1b (2026-10-01): best AUC5 (63.7 vs q80 63.1); its q6K was the weakest tier on 5 of 7 metrics and was removed (locally + HF). Removed tiers remain reproducible from the git history of RESULTS.md / FLEETQUANTCONFIRMATION.md / the converters.
- q4K was removed on 2026-09-18 (not a valid Pareto point), see [`benchmarks/RESULTS.md`](https://github.com/Asher-1/map-anything-ggml/blob/main/cppggml/benchmarks/RESULTS.md) §6.
- Quantization applies only to 2-D Linear weights; LayerNorm/conv/RoPE stay floating point.

5. Output directory conventions (vs the official demo)
Key mapping: pose_enc ↔ .pose.bin, depth[...,0] ↔ .depth.bin, depth_conf ↔ .depth_conf.bin; extrinsic/intrinsic/world_points_from_depth are derived from pose_enc+depth with the official formulas (built into both helper scripts).
6. One-click reproduction
./run_mapggml.sh demo # newcomers: download -> convert -> build -> infer
./run_mapggml.sh gate # dual-runtime parity gate
./run_mapggml.sh bench --backend cuda # steady-state latency matrix
PYTHONPATH=/tmp/torch_cuda_lib:<repo> python3 cpp_ggml/scripts/compare_official_demo.py \
--images a.jpg b.jpg --out-dir /tmp/demo_compare # official demo dual-runtime compare