ARotting/micro-mamba-memory
MicroMamba MicroMamba is a small-compute test of input-dependent state-space dynamics. Each sequence contains distracting symbols, a few marked symbols, and a final query asking for one marked item by ordinal position. Solving the task requires selective storage and retrieval rather than ordinary next-token statistics. The model uses a compact Mamba-inspired block with: a causal depthwise convolution; learned stable diagonal state dynamics; input-dependent discretization, input… See the full description on the dataset page: https://huggingface.co/datasets/ARotting/micro-mamba-memory.
MicroMamba
MicroMamba is a small-compute test of input-dependent state-space dynamics. Each sequence contains distracting symbols, a few marked symbols, and a final query asking for one marked item by ordinal position. Solving the task requires selective storage and retrieval rather than ordinary next-token statistics.
The model uses a compact Mamba-inspired block with:
- a causal depthwise convolution;
- learned stable diagonal state dynamics;
- input-dependent discretization, input, and readout terms;
- gated residual output.
The benchmark retains two controls: a state-space model whose dynamics do not depend on the current input and a GRU with comparable scale. This is a pedagogical Mamba-inspired implementation, not a bit-exact reproduction of the official Mamba kernel or its large-scale language-model results.
Verified results
All variants trained on the same 12,000 length-48 sequences and were evaluated on 4,000 independently generated sequences at each length.
On this controlled task, input-dependent state dynamics improved in-distribution accuracy by 44.63 points over fixed dynamics and 18.48 points over the larger GRU. The result is specific to this synthetic selective-memory benchmark.
Reproduce
uv run python projects/micro-mamba/train.py