Most layers do not build a KV cache
KDA carries history in a fixed-size recurrent state. Only 24 Gated MLA layers retain token-growing global KV state.
A 2.8-trillion-parameter model whose compute path is easier to understand as three orthogonal routing systems: KDA and MLA route information through time, Attention Residuals route it through depth, and Stable LatentMoE routes it through channels.
Based on the Kimi K3 technical report and released configuration, with computation paths traced through TorchTitan. The diagrams show logical tensor shapes before parallelism divides them across GPUs.
The decoder repeats a four-layer rhythm: three linear-attention KDA layers, then one global Gated MLA layer. Twenty-three complete groups account for 92 layers; a final MLA layer makes global interaction the last operation in the 93-layer stack.
KDA carries history in a fixed-size recurrent state. Only 24 Gated MLA layers retain token-growing global KV state.
Block AttnRes replaces uniform residual accumulation with learned mixing over eight layer-block summaries plus the embedding.
The routed expert bank holds 97.94% of the parameters, while each token selects only 16 of 896 experts.
MoonViT-V2 is trained from scratch under the same next-token objective, and projected visual tokens enter the shared decoder stream.
The heavy-bordered boxes are matrix multiplies or attention kernels. Choose a representative layer, then click an operation to see its weights, output tensor, and backward dependencies.
T is the number of packed tokens in this logical invocation. Several documents can occupy those tokens; cumulative sequence offsets preserve their boundaries. R = 16T counts token-expert assignments across the whole expert group. A GPU receives only its share of R.
Select a component to follow its actual data path. Dimensions below use the released Kimi K3 configuration.
KDA applies channel-wise forgetting before a delta-rule write. Query, key, and value each pass through a width-4 depthwise causal convolution and Swish; query and key are L2-normalized. A low-rank branch creates one retention value for every head and key channel.
The lower-bounded decay keeps a 16-token tile's reciprocal cumulative decay below e80, within BF16 range. That lets diagonal as well as off-diagonal tiles use dense Tensor Core matrix multiplies.
Every fourth layer performs unrestricted global attention, and layer 93 adds one final global pass. Keys and values are reconstructed from a 512-wide latent. K3 uses NoPE: the extra shared 64-wide key slice is retained, but no rotary transform is applied.
The shared path stays at model width. The routed path first compresses 7168 channels to 3584, selects 16 experts from 896, normalizes their aggregate, then projects back to 7168. This halves the width carried through expert dispatch and expert matrices.
Quantile Balancing estimates each expert's target quantile from globally reduced histograms. Its bias changes assignment only; the mixture weights and router gradients use the original scores. The bias update takes effect on the next step.
Instead of forcing all prior computation through one accumulated residual stream, each sublayer forms a learned softmax mixture of the embedding, completed 12-layer block sums, and the current block's partial sum.
Full AttnRes would retain every layer output. Block AttnRes reduces the memory and pipeline-communication term from O(layers × width) to O(blocks × width), while preserving learned access to distant depth.
The 401M-parameter vision tower is trained from scratch with next-token prediction. Images and videos share its parameters; spatial and temporal attention are factorized, then a 2×2 pixel shuffle cuts the visual token count by four before projection into the language width.
The routed expert matrices dominate the checkpoint. Toggle the accounting lens to see why the model can activate about 104.2B parameters while storing 2.78T.
KDA state is constant with sequence length; MLA cache grows one compressed entry per token. The calculator is structural rather than a serving-memory promise: it omits allocator overhead, convolution history, quantization metadata, and tensor-parallel sharding.
The report's system is as architectural as the network itself. It combines parallel dimensions, balanced expert placement, and tensor-granular activation storage policies to keep useful work on the critical path.
Dynamic redundant experts guarantee a perfectly balanced plan with at most E/R redundant experts per rank. Static shapes remove per-layer host synchronization; fused permutation writes directly into the remote expert-grouped layout.
Each rank computes a local transition and a zero-initialized state. Because these summaries compose associatively, an all-gather plus prefix scan reconstructs every rank's incoming state with communication independent of sequence length.
Recompute, block-wise FP8 quantization, local offload, and remote offload become pluggable storage policies. AttnRes is checkpointed, and dispatch is recomputed during MoE backward instead of retained.
Training combines PP with virtual stages, EP, ZeRO-1, Pipeline ZeRO-2, and CP. Dynamic vision CP balances large images, while most vision forward and backward work is scheduled into pipeline bubbles.
The exact structural count shown here is 2,779,931,738,208 parameters from the TorchTitan module shapes. The routed bank contributes 2,722,740,830,208. The 104,189,515,872 active-text-parameter reconstruction counts 16/896 of that bank, all non-routed text-path weights, and the output head; it excludes the input embedding lookup and vision pathway. The released checkpoint uses MXFP4 for routed expert weights and MXFP8 input activations during quantization-aware post-training; parameter counts remain logical counts, not stored bytes.
The report also describes one MTP layer later fine-tuned into an EAGLE-3-style draft model. The released backbone configuration sets
num_nextn_predict_layers=0, so that draft is not included in the 2.78T tally above.