Team Ai
Datasetpublic

ARotting/micro-mamba-memory

MicroMamba MicroMamba is a small-compute test of input-dependent state-space dynamics. Each sequence contains distracting symbols, a few marked symbols, and a final query asking for one marked item by ordinal position. Solving the task requires selective storage and retrieval rather than ordinary next-token statistics. The model uses a compact Mamba-inspired block with: a causal depthwise convolution; learned stable diagonal state dynamics; input-dependent discretization, input… See the full description on the dataset page: https://huggingface.co/datasets/ARotting/micro-mamba-memory.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes20downloads
Dataset Card

MicroMamba

MicroMamba is a small-compute test of input-dependent state-space dynamics. Each sequence contains distracting symbols, a few marked symbols, and a final query asking for one marked item by ordinal position. Solving the task requires selective storage and retrieval rather than ordinary next-token statistics.

The model uses a compact Mamba-inspired block with:

  • —a causal depthwise convolution;
  • —learned stable diagonal state dynamics;
  • —input-dependent discretization, input, and readout terms;
  • —gated residual output.

The benchmark retains two controls: a state-space model whose dynamics do not depend on the current input and a GRU with comparable scale. This is a pedagogical Mamba-inspired implementation, not a bit-exact reproduction of the official Mamba kernel or its large-scale language-model results.

Verified results

All variants trained on the same 12,000 length-48 sequences and were evaluated on 4,000 independently generated sequences at each length.

VariantParametersLength 48Length 96 zero-shot
Selective SSM4,59487.85%87.23%
Fixed-dynamics SSM3,31443.23%31.05%
GRU control7,14669.38%70.00%

On this controlled task, input-dependent state dynamics improved in-distribution accuracy by 44.63 points over fixed dynamics and 18.48 points over the larger GRU. The result is specific to this synthetic selective-memory benchmark.

Reproduce

powershell
uv run python projects/micro-mamba/train.py