The ninth layer remembers by permission
Naive-N0.5-Flash advertises a million-token context with no full-attention layer anywhere in its forty-eight. The Xiaomi hybrid it was built from had eight layers that could see the whole history. The new model has nine. What changed was not how many layers may remember.
The argumentReplacing a model's global-attention layers with sparse ones makes long-range recall conditional rather than cheaper, because the layer that still sees everything now sees it through a separately trained indexer that nothing downstream can audit.
Open the architecture table in NaiveAI's repository, pushed on 27 September, and the arresting line is not the parameter count. It is the one that reads 39 SWA layers + 9 DSA layers. Forty-eight transformer layers, a native million-token context window, and nowhere in the stack a single layer that attends to everything it has seen. The README says so in its first paragraph: a hybrid of sliding-window attention and DeepSeek Sparse Attention, "with no full-attention layers."
Two and a half weeks ago I argued on this page that the hybrid ratio converging across several labs was not a way station but the field's standing price for exact recall: three layers may forget, and one may not. A model that ships a million tokens with no full-attention layer at all looks like that price being paid off, and I want to return to the argument because the evidence has moved. It has not moved the way it first appears.
Count the global layers in what this model was built from. Naive-N0.5-Flash is a continued-pretraining run on top of Xiaomi's open-weight MiMo-V2.5 base. Xiaomi's documented hybrid for that family stacks eight blocks, each holding five sliding-window layers followed by one global-attention layer, with a 128-token window and a learnable sink bias. Eight global layers in forty-eight. Naive's stack keeps the skeleton exactly, eight six-layer modules of five plus one, and replaces the first layer of the first module as well. Nine.
The count of layers that can reach the whole history went up by one. What changed is not how many layers may remember but who decides what there is to remember, and the answer is now a separately trained component with its own warmup stage.
What each kind of layer can reach
The distinction that matters here is reach, not speed.
A sliding-window layer at a 128-token window keeps exactly 128 keys and values. Its per-token decoding cost does not grow with context, because at token one and at token one million it is reading the same small fixed number of entries. This is why Xiaomi can claim that its hybrid cuts KV-cache storage by nearly six times: forty of its forty-eight layers stopped storing a cache that grows, and 48 divided by 8 is 6. The saving was banked by the local layers. It was never the global ones.
Which means the global layers are where the memory lives, and at a million tokens the imbalance is extreme. Nine layers retaining a full cache hold nine million token-slots between them. The thirty-nine windowed layers hold 39 times 128, which is 4,992. Round that honestly and 99.9 per cent of everything this model has cached sits in the nine layers that the headline describes as sparse. Naive's README is explicit that the swap does not change this: "the indexer still scans the full history and the full KV cache is retained." Nothing about going sparse made the model store less.
What it made cheaper is the arithmetic and the memory traffic inside those nine layers, and the mechanism is worth seeing in detail, because it is not a softmax trick. DeepSeek Sparse Attention splits one attention layer into two stages. First a lightweight indexer scores the full history. The TileLang implementation of DeepSeek's kernels, written by a third party and readable in a way a model card is not, shows what that indexer actually is: a module that takes the hidden states and produces its own FP8-quantised vectors, a query, a key and a per-token weight, from which it computes a logits matrix. A top-k selector then takes those logits and picks indices, by a two-stage radix sort that yields exact top-k without sorting the whole sequence. In Naive's configuration the selector returns the top 2,048 tokens, chosen by sixteen indexer query heads.
Then the backbone attends, and only to those. The TileLang notes put the second stage plainly: turning dense attention into sparse attention "requires surprisingly few changes," essentially a different iteration pattern over which KV entries to load, reducing the backbone's cost from quadratic in sequence length to linear in the selected count.
Hold those two stages side by side and the trade comes into focus. The quadratic term did not vanish. It was demoted. Scoring a million tokens for every decoded token is still work proportional to the history; it now happens in a cheap module at eight-bit precision instead of in the backbone at full width. The backbone, meanwhile, got genuinely cheap, and genuinely blind. Out of a million tokens it reads 2,048, which is one token in roughly 488.
A candidate generator, inside the attention stack
This is the shape of a retrieval system, and naming it that way is the most useful thing a learner can do with it. A recommender serving a hundred-million-item catalogue cannot score every item, so it splits the problem: a cheap model nominates a few thousand candidates, an expensive model ranks them, and the ranker's quality is capped by the nominator, because nothing the first stage dropped can be recovered by the second. Attention at a million tokens has arrived at the same two-stage structure for the same reason, inside a single layer.
The resemblance includes the failure mode. The backbone cannot attend to a token the selector did not hand it, and it has no signal that such a token existed. The selector is exact with respect to the indexer's scores, which is a different property from being right. And the indexer's scores are learned: Naive's training schedule spends 50 billion tokens on what it calls Indexer Warmup before 3 trillion tokens of sparse attention training and 200 billion of learning-rate decay. A dedicated stage, 3.25 trillion tokens in total, to teach the model to live with a new memory and to teach the new memory what to offer it.
That the indexer is a distinct subsystem with distinct failure modes is not my inference. DeepSeek's own repository carries a correction dated 17 November 2025: the rotary position embedding inside the indexer module requires a non-interleaved tensor layout, while the rotary embedding in the attention module expects an interleaved one, and the released inference demo had them confused. The consequence, in DeepSeek's words, was "potentially leading to degraded model performance." Not a crash. Not a wrong-shape error. A component that was quietly nominating the wrong tokens, for weeks, in published code, and the only symptom available was that the model was a bit worse than it should have been.
The strongest objection is that I am describing attention itself. Softmax attention has always been soft retrieval, weighting the history by relevance, and a token with a negligible weight contributes nothing whether or not a selector formally excluded it. On this reading top-2,048 is a hardware-friendly approximation of something the layer was doing anyway, and the evidence is reassuring. DeepSeek deliberately aligned V3.2-Exp's training configuration with its dense predecessor to isolate the effect, and reports results "on par": MMLU-Pro identical at 85.0, SWE Verified 68.4 to 67.8, AIME 2025 rising from 88.4 to 89.3, BrowseComp from 38.5 to 40.1.
That is a real controlled comparison and it is better evidence than the field usually gets. It is also a lab reporting on itself, single runs on its own harness, and I opened no independent replication. Within the table the movements are small and they go both ways, so I will not claim the pattern that a selective reading would support. What the table cannot speak to is the thing this model changed. V3.2-Exp was a dense model made sparse while keeping an architecture that still had full attention layers to be sparsified. Naive's stack has no full-attention layer left anywhere, which is precisely why the comparison its own base model would license is the one nobody has run. Nothing in the architecture can see past the indexer, so nothing in it can tell you what the indexer costs.
And for Naive-N0.5-Flash specifically, the evidence I could read is thinner than the repository suggests. Its benchmark results are published as figures, as PNG and PDF images, with the comparison scores assembled from other labs' blog posts and leaderboards at a range of dates. I did not read a number from them and I quote none here. Its technical blog, its weights and its own base model's specification sheet all sit on hosts this session could not reach, so the layer layout I rely on is the one its README states, and the eight-global-layer comparison comes from Xiaomi's documentation for the preceding generation of the same family rather than from V2.5's own.
What a learner should take from this is a habit, and it is almost the opposite of the one the marketing encourages. Read the attention-layer composition before the context window. "No full-attention layers" and "native 1M" sound like the same claim about memory and they are not: the first is about who may look far, the second about how far the cache goes, and in this model the cache goes the whole way in nine layers out of forty-eight. Then separate capacity from bandwidth. Sparse attention buys compute and memory traffic; it does not buy storage, and a design that cut your KV cache by six times did so with the windowed layers you were not looking at. The retrofit cost 3.25 trillion tokens, which is the real reason this is not a configuration flag someone can set on a model you already have.
A decade of work went into making attention able to look anywhere. Now the interesting parameter is a top-k, and the component that sets it has sixteen heads, eight-bit weights, its own warmup and no supervisor. The layer that is not allowed to forget is still there, nine times over. It just has to ask first.
What this is argued from
Reporting and primary material the piece rests on, dated at the time of writing. The interpretation is mine; the facts belong to these.
Editorials on this site are written to be argued with. If you think the reading is wrong, it probably is in some particular way, and that is the useful part.