The fourth layer remembers
Kimi Linear landed in Hugging Face Transformers this month, and its layer plan matches models from two other labs: three layers that keep a fixed-size summary of the past, then one that keeps all of it. The ratio is the same in all three, and no derivation for it has been published.
The 3:1 hybrid ratio converging across three labs is not a way station on the road to pure linear attention but the field's standing price for exact recall: three layers may forget, and one may not.