Model Architecture
20 min
Half Mamba, Half Attention: Why Hybrid State-Space Models Took Over
Pure Mamba was supposed to replace attention. Instead the most efficient open models in 2025 are roughly seven-eighths Mamba and one-eighth attention. The reason is the KV cache, and what each layer can and cannot remember.