Sequence · depth · width

An infra-oriented diagram of Kimi K3

A 2.8-trillion-parameter model whose compute path is easier to understand as three orthogonal routing systems: KDA and MLA route information through time, Attention Residuals route it through depth, and Stable LatentMoE routes it through channels.

Based on the Kimi K3 technical report and released configuration, with computation paths traced through TorchTitan. The diagrams show logical tensor shapes before parallelism divides them across GPUs.

2.78Ttotal parameters
104.2Bactivated parameters
93decoder layers
69 + 24KDA + Gated MLA
16 / 896routed experts active
1Mmaximum context

The model at a glance

The decoder repeats a four-layer rhythm: three linear-attention KDA layers, then one global Gated MLA layer. Twenty-three complete groups account for 92 layers; a final MLA layer makes global interaction the last operation in the 93-layer stack.

Native vision input MoonViT-V2
14×14 patches
images + video
27 layers · width 1024
spatial + temporal attention
2×2 pixel shuffle
MLP projector → 7168
Decoder backbone · width 7168 23 repeated groups, then one final global layer
KDA
KDA
KDA
Gated MLA
Layer 1 uses a dense FFN; layers 2-93 use Stable LatentMoE · 896 choose 16 + 2 shared
Block Attention Residuals retrieve from the embedding, completed block sums, and the current partial block
Output Final AttnRes
RMSNorm
163,840-way LM head
KDA layer Gated MLA layer dense FFN outline
Layer 1 KDA + the model's only dense FFN; all later layers use Stable LatentMoE.

Most layers do not build a KV cache

KDA carries history in a fixed-size recurrent state. Only 24 Gated MLA layers retain token-growing global KV state.

Depth has an attention mechanism too

Block AttnRes replaces uniform residual accumulation with learned mixing over eight layer-block summaries plus the embedding.

Almost all weights are dormant per token

The routed expert bank holds 97.94% of the parameters, while each token selects only 16 of 896 experts.

Vision is native, not attached afterward

MoonViT-V2 is trained from scratch under the same next-token objective, and projected visual tokens enter the shared decoder stream.

Follow the kernels, one layer at a time

The heavy-bordered boxes are matrix multiplies or attention kernels. Choose a representative layer, then click an operation to see its weights, output tensor, and backward dependencies.

GEMM / attention Norm / elementwise / layout Expert dispatch / combine Backward needs this output

T is the number of packed tokens in this logical invocation. Several documents can occupy those tokens; cumulative sequence offsets preserve their boundaries. R = 16T counts token-expert assignments across the whole expert group. A GPU receives only its share of R.

One layer, five views

Select a component to follow its actual data path. Dimensions below use the released Kimi K3 configuration.

Kimi Delta Attention: fixed-state sequence mixing

KDA applies channel-wise forgetting before a delta-rule write. Query, key, and value each pass through a width-4 depthwise causal convolution and Swish; query and key are L2-normalized. A low-rank branch creates one retention value for every head and key channel.

St = (I - βtktktT) Diag(αt) St-1 + βtktvtT
  • Layers69
  • Heads96
  • Head / state shape128 / 128×128
  • Convolution width4
  • Log-decay boundgmin = -5
  • Parameters per KDA443.74M
q, k, v7168 → 96×128, ShortConv, Swish; q and k then L2Norm
α retention7168 → 128 → 96×128; scaled sigmoid bounds log-decay to (-5, 0)
β write strength7168 → 96, followed by sigmoid
Chunk inputT × 7168
→
Parallel tilesdense Tensor Core math inside a chunk
→
Recurrent state96 × 128 × 128, fixed with context length
→
Head RMSNorm + gatefull-rank 7168 → 12288 gate
→
Output projection12288 → 7168

The lower-bounded decay keeps a 16-token tile's reciprocal cumulative decay below e80, within BF16 range. That lets diagonal as well as off-diagonal tiles use dense Tensor Core matrix multiplies.

Gated MLA: periodic global interaction

Every fourth layer performs unrestricted global attention, and layer 93 adds one final global pass. Keys and values are reconstructed from a 512-wide latent. K3 uses NoPE: the extra shared 64-wide key slice is retained, but no rotary transform is applied.

yt = Wo[Sigmoid(Wgxt) ⊙ attention(qt, k≤t, v≤t)]
  • Layers24
  • Query heads96
  • Query rank1536
  • KV latent512
  • q/k and v per head192 and 128
  • Parameters per MLA232.20M
Query branch7168 → 1536 → 96×192
KV branch7168 → 512 latent + 64 shared key; latent → 96×(128 key + 128 value)
Output gate7168 → 96×128; sigmoid modulation before Wo
xT × 7168
→
Low-rank Q / KV1536 query latent; 512 KV latent
→
Global attention96 heads; causal; NoPE
→
FP32 output tilepaper's training-kernel correction for rounding bias
→
Gate + Wo96×128 → 7168

Stable LatentMoE: extreme width behind a narrow waist

The shared path stays at model width. The routed path first compresses 7168 channels to 3584, selects 16 experts from 896, normalizes their aggregate, then projects back to 7168. This halves the width carried through expert dispatch and expert matrices.

u = Σi∈Top16(x) piEi(Wdownx),   y = ΣEshared(x) + WupRMSNorm(u)
  • Routed experts896
  • Active per token16 (1/56)
  • Latent width3584
  • Expert hidden width3072
  • Shared experts2 at width 7168
  • Expert bank per layer29.595B
Router + Quantile BalancingSigmoid scores, bias-adjusted Top-16 selection, then normalization using the original unbiased scores
896 routed expertsEach expert has three matrices and 33.03M parameters; SiTU-GLU caps both multiplicative branches
2 shared expertsFull-width 7168 → 3072 paths run for every token and capture common transformations
xT × 7168
→
Wdown7168 → 3584
→
Dispatch Top-16896 experts · SiTU-GLU
→
Weighted sum + RMSNormstabilizes routed scale
→
Wup + shared path3584 → 7168

Quantile Balancing estimates each expert's target quantile from globally reduced histograms. Its bias changes assignment only; the mixture weights and router gradients use the original scores. The bias update takes effect on the next step.

Block AttnRes: retrieval over depth

Instead of forcing all prior computation through one accumulated residual stream, each sublayer forms a learned softmax mixture of the embedding, completed 12-layer block sums, and the current block's partial sum.

hl = Σi softmax(wlT RMSNorm(bi)) · bi
  • Layer blocks8
  • Nominal block size12 layers
  • Final block9 layers
  • Sources incl. embeddingup to 9
  • Storage scalingO(blocks × width)
  • Learned parameters2.67M
Embedding b0always retained as a source
+
Completed blocksb1 ... bn-1
+
Current partial sumfor later layers in a block
→
RMSNorm keys + pseudo-queryone scalar score per source
→
Softmax mixtureinput to attention or FFN

Full AttnRes would retain every layer output. Block AttnRes reduces the memory and pipeline-communication term from O(layers × width) to O(blocks × width), while preserving learned access to distant depth.

MoonViT-V2: native visual tokens

The 401M-parameter vision tower is trained from scratch with next-token prediction. Images and videos share its parameters; spatial and temporal attention are factorized, then a 2×2 pixel shuffle cuts the visual token count by four before projection into the language width.

  • Vision layers27
  • Vision width1024
  • Attention heads12
  • Patch size14
  • Vision tower401.2M
  • Tower + projector447.4M
Image / videoup to 3584×3584 images
→
Patch embed14×14 patches, width 1024
→
MoonViT-V227 bias-free transformer layers
→
2×2 pixel shuffle4× fewer tokens, 4096 channels
→
MLP projector4096 → 7168

Where 2.78 trillion parameters live

The routed expert matrices dominate the checkpoint. Toggle the accounting lens to see why the model can activate about 104.2B parameters while storing 2.78T.

2.780T
Stored: 92 layers × 896 experts × 3 matrices × 3584 × 3072 = 2.7227T routed-expert parameters, or 97.94% of the model.

Why hybrid attention changes long-context memory

KDA state is constant with sequence length; MLA cache grows one compressed entry per token. The calculator is structural rather than a serving-memory promise: it omits allocator overhead, convolution history, quantization metadata, and tensor-parallel sharding.

24 MLA caches
69 KDA states
K3 hybrid total
93-layer MLA baseline

MLA = N × 24 × (512 latent + 64 shared-key) × bytes; KDA = 69 × 96 × 128 × 128 × state bytes.

Training the 3T-class model

The report's system is as architectural as the network itself. It combines parallel dimensions, balanced expert placement, and tensor-granular activation storage policies to keep useful work on the critical path.

Router output
current layer and microbatch
Online plan
place bounded redundant experts
Zero-copy dispatch
write directly into expert order
Static expert work
exactly S×K tokens per rank
Combine + reduce
return redundant gradients home

MoonEP: balance first

Dynamic redundant experts guarantee a perfectly balanced plan with at most E/R redundant experts per rank. Static shapes remove per-layer host synchronization; fused permutation writes directly into the remote expert-grouped layout.

KDA Context Parallelism

Each rank computes a local transition and a zero-initialized state. Because these summaries compose associatively, an all-gather plus prefix scan reconstructs every rank's incoming state with communication independent of sequence length.

One activation abstraction

Recompute, block-wise FP8 quantization, local offload, and remote offload become pluggable storage policies. AttnRes is checkpointed, and dispatch is recomputed during MoE backward instead of retained.

Use pipeline bubbles

Training combines PP with virtual stages, EP, ZeRO-1, Pipeline ZeRO-2, and CP. Dynamic vision CP balances large images, while most vision forward and backward work is scheduled into pipeline bubbles.

Paper system versus TorchTitan: MoonEP is the system described in the Kimi K3 report. TorchTitan's current K3 path implements the model and offers its own FSDP/EP dispatch choices, including HybridEP; that is not a reproduction of MoonEP's redundant-expert planner, Pipeline ZeRO-2, or the report's complete PP/VP/CP schedule.

Sources and counting notes

  1. Kimi K3 Technical Report: architecture, equations, training recipe, MoonEP, long-context systems, and reported headline counts.
  2. Kimi K3 model card and released configuration: layer indices, dimensions, quantization metadata, and tokenizer IDs.
  3. TorchTitan Kimi K3 implementation: tensor shapes and structural parameter count.
  4. Diagram presentation inspired by Edward Yang's DeepSeek architecture diagram. The Kimi K3 computation graph and dimensions are specific to this model.

The exact structural count shown here is 2,779,931,738,208 parameters from the TorchTitan module shapes. The routed bank contributes 2,722,740,830,208. The 104,189,515,872 active-text-parameter reconstruction counts 16/896 of that bank, all non-routed text-path weights, and the output head; it excludes the input embedding lookup and vision pathway. The released checkpoint uses MXFP4 for routed expert weights and MXFP8 input activations during quantization-aware post-training; parameter counts remain logical counts, not stored bytes.

The report also describes one MTP layer later fine-tuned into an EAGLE-3-style draft model. The released backbone configuration sets num_nextn_predict_layers=0, so that draft is not included in the 2.78T tally above.