Team Ai
Datasetpublic

avithal/repro-sf-mamba

Reproduction: SF-Mamba: Rethinking State Space Model for Vision Independent reproduction of ICML 2026 paper R9AUrEgZEq (arXiv 2603.16423) by Yoshimura, Hayashi, Hoshino, Wang, Ohashi (Sony). No official code or checkpoints are released ("We will release the source code after publication"). SF-Mamba builds on the public NVlabs/MambaVision backbone (MambaVisionMixer(d_state=8, d_conv=3, expand=1), SSM on dim/2 channels). What is reproduced Claim Approach… See the full description on the dataset page: https://huggingface.co/datasets/avithal/repro-sf-mamba.

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes36downloads
Dataset Card

Reproduction: SF-Mamba: Rethinking State Space Model for Vision

Independent reproduction of ICML 2026 paper R9AUrEgZEq (arXiv 2603.16423) by Yoshimura, Hayashi, Hoshino, Wang, Ohashi (Sony).

No official code or checkpoints are released ("We will release the source code after publication"). SF-Mamba builds on the public NVlabs/MambaVision backbone (MambaVisionMixer(d_state=8, d_conv=3, expand=1), SSM on dim/2 channels).

What is reproduced

ClaimApproachScript
C1 auxiliary patch swapping (Sec 3.2, Eq 2-3)mechanism test: tiny unidirectional Mamba vs +aux-swap vs bidirectional on a synthetic future-information taskclaim1_swap.py
C2 batch folding + periodic state reset, 110–180% SSM speedup (Sec 3.3, Fig 3-4)(a) fp64 reference proof of fold+reset ⇔ unfolded equality; (b) upstream mamba_ssm selective-scan CUDA kernel benchmark at the paper's exact Fig-4 configs on A100, sweeping B2claim2_bench.py
C3 ImageNet acc/throughput of SF-Mamba-T/S/Baccuracy: not reproducible (no code/weights; 300-epoch ×3 training infeasible). Throughput side: benchmark public MambaVision-T/S/B + Swin-T/ConvNeXt-T on A100 batch 128 and compare with the paper's Table 1claim34_throughput.py
C4 tradeoff vs Vim-S / VMamba-Tbaseline numbers cross-checked against official releases + our A100 orderingclaim34_throughput.py + literature
C5 COCO / ADE20Kdocumented; training infeasible here (8-GPU mmdet/mmseg schedules)logbook page

Rerun

bash
# CPU-ok correctness parts
python claim2_bench.py --out claim2.json          # fp64 fold+reset equivalence
python claim1_swap.py --steps 1500 --seeds 3      # mechanism test (GPU faster)

# A100 job (HF Jobs)
hf jobs run --flavor a100-large --timeout 3600 -v ./:/repro \
  -v hf://buckets/<user>/<bucket>:/data \
  pytorch/pytorch:2.4.0-cuda12.1-cudnn9-devel bash /repro/job.sh