First Light: MiMo-V2.6-Pro's architecture in 3D

Samuel Batista

Samuel Batista / October 01, 2026

5 min read • ––– views

After Kimi K3, here's a look inside Xiaomi's MiMo-V2.6-Pro. Follow a prompt from the numbers behind each token, to how those tokens relate, all the way to MiMo's choice of what to say.

The ring shows the text model in MiMo-V2.6-Pro-RL: 70 processing blocks, each gathering information through attention before passing it through a neural network. Most blocks read only the nearest 128 tokens; ten can look across the entire context. In 69 blocks, a router selects 8 of 384 smaller networks, called experts, for each token. These are the model's architectural settings, not measurements from the animation.

The answer, scores, expert choices and numerical values are illustrative. This isn't a recording of MiMo generating an answer. Use Study mode to pause at each caption and advance at your own pace.

Questions you might have

What are the weights, stars and moving lights?

The machinery holds learned weights, which stay fixed during the conversation. The stars represent the conversation's tokens. Each block's column stores numerical records of those tokens, called keys and values: keys help attention find matches, and values supply the information it combines. Together, the stored records are the KV cache.

Each moving beam represents 6,144 numbers being updated as a token passes through the model. The animation calls these numbers the token's working memory, but they aren't a separate storage device. The columns keep records for later tokens to read; the beam carries the current token's changing representation.

Why do only the prompt's 7 tokens take the first lap?

Each time MiMo processes a token, every one of its 70 blocks files a record of it in that block's column. Later tokens read those records rather than processing the token again, so each token only needs one lap.

Our conversation already had 3,072 tokens from earlier messages, and their records were filed when those messages were processed. The software running MiMo can keep those records in memory after it replies, so when the new message arrives, the earlier tokens are already done. Only the prompt's 7 new tokens need a lap. If the records were cleared in the meantime, for example to make room for other people's conversations, all 3,079 tokens would have to go through again.

Paste a 500,000-token document into a new chat and prefill, the stage that processes the prompt before the answer begins, has to work through all of it, up to MiMo's limit of 1,048,576 tokens. In principle the whole prompt could take one lap together. In practice, the software running the model, such as SGLang, splits a long prompt into chunks of a few thousand tokens, one lap per chunk. That keeps memory in bounds and lets the server keep other people's answers flowing between chunks. Either way, the rule from the animation holds: each token reads only the tokens before it, whether they came in an earlier chunk or earlier in its own.

Length has a cost. In the ten global blocks, every token reads everything before it, so their attention work grows roughly with the square of the prompt's length. That's why a long prompt takes a while before the answer's first word appears, and why keeping earlier turns' records saves so much time.

Does a short attention window mean MiMo forgets older tokens?

Not everywhere. The 60 sliding-window blocks stop reading records outside their recent window, while the ten global blocks can still use older records within the context. Keeping a record available doesn't guarantee the model will retrieve the right detail. The context limit of 1,048,576 tokens is a capacity setting, not a promise of perfect recall.

What does the dimmer inside a neural-network unit represent?

Each unit in a block's neural network makes two weighted sums from its inputs. It applies SiLU (Sigmoid Linear Unit), a smooth mathematical function, to one sum and multiplies the result by the other. The dimmer shows how one branch controls the other, but the analogy has limits: the multiplier can be negative or greater than one. This gated calculation, known as SwiGLU, appears in the implementation.

Are the experts assigned subjects?

No one labels them "math" or "writing." Training shapes both their weights and the router that selects them. The selected experts process the same token's numbers, and the router decides how much each one contributes to the combined result: the blend shares in the animation, which add up to one. They're separate from the learned weights stored inside each expert. Xiaomi reports 1.02 trillion total parameters and 42 billion active per token: many learned numbers to draw on without using all of them on every step.

How can MiMo write several tokens in one pass?

A small drafting model proposes seven tokens after the latest known token. MiMo checks the draft by processing eight positions together: the latest known token plus the seven guesses. At each position it compares the guess with the token it would have chosen in its place. In the animation, two guesses match before one fails, so it keeps those two, adds its own correction and discards the rest. This is speculative decoding, using the release's DFlash implementation.

I illustrate choosing the most likely token at each step. Real generation can instead sample from the probabilities; Xiaomi recommends sampling settings. Seven guesses don't guarantee seven accepted tokens or a sevenfold speedup.

Made with Claude Opus 5.5, GPT-6 Astra, MiMo-V2.6-Pro, three.js and Web Audio API.