Kimi K3 studies / 03

Working the roofline
for Kimi K3 on GB200

We made it fit. Now ask which work, bytes, and dependencies set the pace.

Study 01 followed the tensor shapes. Study 02 budgeted the live copies. This study turns those shapes into a performance model: first a GEMM roofline, then a distributed communication ledger, then a check against a real training run.

512GB200 GPUs / FP32 master and state
2,500 TF/snominal dense BF16 ceiling per GPU
8 TB/snominal HBM bandwidth per GPU
370 tok/s/GPUmeasured clean 4K reference

Default: PP8 / CP1 / TP1 / dense FSDP64 / EP64, FP32 master weights, gradients and Muon momentum, BF16 model GEMMs, full layer checkpointing, one 4K row per microbatch, 64 microbatches, ordinary 1F1B, no CUDA graphs. Hardware ceilings and placement assumptions are not measured kernel or network rates.

A ceiling is not a benchmark

Three roofs, three different byte counts

A kernel can be limited by HBM. Distributed work can instead be limited by NVLink or the scale-out network. Divide work by the bandwidth of the resource that actually carries its bytes. In particular, a fast NVLink roof does not make cross-domain traffic disappear.

Resource Default ceiling How to read it
BF16 Tensor Cores 2,500 TFLOP/s / GPU NVIDIA lists 360 PFLOP/s for 72 GPUs with sparsity. Dense = 360 / 72 / 2 = 2.5 PFLOP/s. Expert sparsity is not 2:4 weight sparsity.
HBM 8 TB/s / GPU 576 TB/s / 72. This is bandwidth, not the 184.31-GiB capacity reported by the training runtime.
NVLink 900 GB/s / GPU, one direction Half of the advertised 1.8-TB/s bidirectional endpoint bandwidth. Actual collective bandwidth, contention and power caps can lower it.
Scale-out network 50 GB/s / GPU, illustrative A 400-Gbit/s-equivalent, one-direction budget. It is not a claim about this cluster's NIC allocation or measured throughput. Edit it for the real system.
One GEMMt ≥ max(F / P, B / βHBM)

F = multiply-add FLOPs; B = modeled operand bytes. Arithmetic intensity I = F / B.

The ridgeI* = P / βHBM = 312.5 FLOP/B

Left of the ridge, fewer HBM bytes help. Right of it, the compute ceiling matters. A point under a roof need not attain that roof.

Change a shape, watch a roof move

The roofline lab

The lab is an accounting model, not a launch-config generator. Every run has 512 GPUs and TP1. Changing rows or microbatches changes the global batch; changing CP changes the number of independent data replicas. Memory feasibility still needs its own budget.

Tokens / update

Rows / local expert

Arithmetic + resource sketch

Forward BF16 GEMMs: the HBM roofline
Kimi K3 forward GEMM arithmetic intensity and attainable throughput A logarithmic HBM roofline. Buttons below select a GEMM and explain its shape and position.
These are modeled roof points, not measured kernel dots. B = 2(MK + KN + MN), with BF16 inputs, weights and output, one operand pass and no cross-GEMM cache reuse. Tiling, padding, FP32 accumulation/casts, launch overhead and grouped-GEMM efficiency are omitted. Router points show its matrix contraction only.
Take 6ND apart

Stored parameters are not activated FLOPs

For Y = XW, with X [M,K] and W [K,N], forward costs 2MNK FLOPs. Backward has two more contractions of the same size: dX = dYWT and dW = XTdY. Three contractions give 6MNK. Summed over learned matrices, this becomes 6NactiveDtokens. Here N is activated matrix weights, not the entire 2.780T checkpoint.

ForwardXMK WKNYMN / 2MNK
Input gradientdYMN WNKdXMK / 2MNK
Weight gradientXKM dYMNdWKN / 2MNK
Routed expert weights activated / token
  = 92 layers * 16 experts * 3 matrices * 3584 * 3072
  = 48,620,371,968

All learned active matrices = 104,174,092,288
Learned-matrix training work = 625.04 GFLOP / token

The familiar 104.1895B active-parameter count also includes small non-matmul weights. It is useful for a conventional metric, but those weights do not each cause a dense matmul. The input embedding lookup and text-only vision path are excluded from per-token matrix compute; the LM head is included.

Attention adds work without adding a weight matrix

Core Algorithmic training work / token What is not in the formula
69 KDA layers 69 × 18 × 96 × 128 × 128 = 1.9535 GFLOP Three recurrent state contractions, forward plus gradients. The actual chunked kernel also has transforms, triangular solves, FP32 work and memory traffic.
24 causal MLA layers 24 × 3 × 96 × (192 + 128) × (Seff + 1) Seff = Σ sdoc2 / Σ sdoc. This counts causal QK and PV contractions, not norms, softmax or gating.
FlashAttention replay 24 × 96 × 192 × (Seff + 1) The QK contraction rematerialized by attention backward. Counted for the runtime sketch, not conventional model FLOPs.
Full layer checkpoint One extra forward for decoder matrices and attention cores The model-wide output-head GEMM is not included in the layer checkpoint replay. Full recompute also repeats MoE communication.

Equal documents make Seff easy to see: one 4K document gives 4,096; four independent 1K documents give 1,024. Two 2K rows have the same tokenwise projection work as one 4K row, but roughly half the MLA contractions. Flattening [rows, context] to T does not remove document boundaries. CP shortens local token arrays, not the logical context attended by MLA.

Per-token work at the current shapeAlgorithmic accounting / not hardware counters

Model FLOPs, issued-work estimates, MFU and HFU are different quantities

The fixed training results below follow TorchTitan's convention: 6 × 104,189,515,872 + KDA contractions + non-causal MLA contractions, using the configured row context. At 4K that is 645.210 GFLOP/token. The runtime sketch instead uses causal pairs, learned matrices only, FlashAttention replay and the selected layer recomputation.

MFU = model FLOP/s divided by the chosen nominal dense hardware peak. It should not reward checkpoint recomputation. HFU includes actual hardware work, but an algorithmic replay estimate is not a measured HFU: custom-kernel implementation work is still missing. More recompute can make a job look compute-bound while making time-to-result worse.

Follow each byte to its destination

EP is payload and control plane

K3 compresses the routed stream from 7,168 to 3,584 channels before dispatch. A BF16 latent row is 7,168 bytes, not 14,336. The shared experts stay on the 7,168-wide dense path and are not dispatched. With full checkpointing, the route is traversed three times: forward, recomputed forward, and the two reverse communication operations in backward.

Forwarddispatch + combine
Checkpoint replaydispatch + combine again
Backwardreverse dispatch + reverse combine
No checkpoint: 4 payload exchanges / layer / microbatch
Full checkpoint: 6 payload exchanges / layer / microbatch

Assignment budget, uniform routing:
  remote GPU copies / token = top_k * (1 - 1 / EP)
  cross-domain copies      = top_k * (1 - local_EP_ranks / EP)

Payload bytes = exchanges * local_tokens * latent_width * 2 * remote_copies

One-copy-per-assignment is a conservative payload budget. If the dispatcher sends a token once per destination rank and sums results before returning them, multiple chosen experts on that rank can share a copy. The second accounting option models this ideal deduplication under uniform sampling without replacement; it does not assert that every backend or phase implements it. Metadata, alignment and routing weights are additional.

The extra HybridEP all-gather is not an expert weight gather

The 4K reference exchanges a Boolean routing map [4,096, 896] before dispatch: 3.5 MiB locally, 224 MiB gathered across EP64, of which 220.5 MiB came from other ranks. The original and recomputed forward each do it. On a 12-MoE-layer stage, 64 microbatches × 12 layers × 2 = 1,536 metadata all-gathers per update. Small inputs can thus imply a substantial global output and a repeated dependency.

Router / top-kRouting-map all-gather Scan / permutationPayload dispatchExpert GEMMs

FSDP counts depend on the lifecycle, not just the acronym

At PP8/EP64, routed experts have expert FSDP1: there is no routed-weight FSDP gather. Dense weights are sharded over 64 ranks. A BF16 ring all-gather receives 2P(1 − 1/n) bytes per rank; an FP32 reduce-scatter has the corresponding 4P(1 − 1/n) budget. Multiply by the actual number of collectives.

Lifecycle model Weight AG rounds / update Gradient RS rounds / update Tradeoff
4K reference-like M + 1 = 65 1 65 gathers per layer observed in the interior-stage trace. Extrapolating that multiplier to other lab shapes is illustrative.
Textbook reshard 2M = 128 M = 64 Gather before forward and backward; reduce each microbatch. Lower residency can cost much more communication.
Retain views across the schedule 1 1 An optimistic lifecycle alternative. Keeps BF16 views and deferred gradients live; requires a different memory budget and runtime support.

For a hierarchical collective, a domain with a of n shards need only import P(1 − a/n) parameters once, then share them locally. Dividing its import by a models the per-rank share of an a-NIC aggregate budget. This is an ideal domain-level network bound, not the measured wire volume of NCCL's chosen algorithm. NIC sharing or a different rank layout changes it.

Per-stage communication ledger

Traffic / update Endpoint budget NVLink service NIC-domain budget Network service

Budgets are balanced one-direction endpoint traffic or amortized one-direction domain import; opposite-direction capacity is not added to the bandwidth denominator. The two projections are not disjoint byte categories to add. Different collectives on the same resource compete; overlap does not grant each its own full NVLink or NIC bandwidth.

Resource service at the current shapeMaximum over PP ranks / ideal rates

Long-context CP scenarios keep the logical attention length, DP count and local-token shape straight, but CP attention communication is not priced here. KDA state/prefix exchange, MLA KV/head redistribution and backend-specific CP collectives require their own trace-derived ledger. The 64K preset is a shape and compute comparison, not a complete long-context step-time prediction.

The useful FLOPs still need a schedule

Ordinary 1F1B has a bubble, not a ban on overlap

Under balanced stage costs, an ideal ordinary 1F1B schedule has compute-bubble overhead (PP − 1) / M relative to useful scheduled work. At PP8 and M64 that is 10.94% overhead, or 9.86% of the compute-plus-bubble interval. Each optimizer update fills and drains its own pipeline; a previous warmup update does not remove those structural bubbles.

Current ideal compute / fill-and-drain sketch
Scheduled workBubble

More microbatches help amortize the bubble, but there is a tradeoff. At fixed global batch, each microbatch gets smaller, potentially lowering GEMM efficiency. In this lab the row shape stays fixed when M changes, so the global batch grows. Science determines whether that larger update is appropriate.

The PP packet is bigger than [T, 7168]

K3 sends (hiddenTD, deltaTND). With one stage per rank, a microbatch visits each receiving rank for the first time, so the residual cache cannot omit the earlier blocks. At the default 12-layer split, outgoing block counts are 2, 4, 6, 8, 10, 12 and 14. The final hop carries one hidden stream plus 14 residual streams: 15 × 56 MiB = 840 MiB forward, and a matching gradient payload backward.

Interleaving with multiple virtual stages can shrink a stage's gathered working set, revisit a rank with a populated residual cache, and offer different scheduling opportunities. It also adds stage boundaries and communication. It is not modeled by simply dividing the 1F1B bubble or memory by VP. The lab deliberately keeps one contiguous stage per rank.

Once per update, not once per microbatch

Muon is a matrix workload of its own

FP32 momentum is persistent training history. Newton–Schulz uses a BF16 working matrix. For one expert matrix [3,584, 3,072], orient the working matrix as X [r,c] with r = 3,072 and c = 3,584. A standard iteration forms A = XXT, B = bA + ccoefA2, then updates X using BX and a scaled X. The leading contractions cost 4r2c + 2r3 FLOPs per iteration.

Gram matrixX XT[3072, 3072]
PolynomialA2[3072, 3072]
Matrix updateB X[3072, 3584]

Five iterations cost about 0.966 TFLOP per logical expert matrix. An interior PP rank with 12 MoE layers and 14 experts per layer has 12 × 14 × 3 = 504 such matrices: about 487 TFLOP per update, before the smaller non-routed matrices, redistribution, normalization, pointwise work and communication. That is a workload estimate, not an optimizer timing measurement.

Fused gate/up storage still contains two logical expert matrices. Muon orthogonalizes those matrices separately; using one wider fused storage shape would give a different Newton–Schulz cost and a different update.

Current selected stage

Fixed 4K measured reference1.50 s / 88.56 s = 1.69%

Exposed optimizer-phase interval, including clipping/checks/update/scheduler. This is not the duration of a CPU DistMuon annotation.

A long CPU optimizer span can include waiting for the pipeline's last backward or a gradient-norm collective. The reference's complete exposed optimizer tail was small compared with its clean step; it does not support blaming the throughput gap on Muon alone. More tokens per update amortize one optimizer update, but also change optimization and memory requirements.

Put the ceiling next to evidence

A roofline gap is a question, not a diagnosis

The fixed 512-GB200 HybridEP references below completed 12 updates with FP32 resident/master weights and optimizer states, BF16 compute, full checkpointing, PP8/FSDP64/EP64, 64 microbatches, and no CUDA graphs. Clean throughput excludes the profiled update and trace-export disturbance. These measurements stay fixed when you change the lab.

Reference Rows × context Tokens/s/GPU Implied clean update Model TFLOP/s/GPU Nominal dense BF16 MFU

MFU here uses a fixed 2,500-TFLOP/s dense denominator and the documented TorchTitan estimator, not a 5,000-TFLOP/s sparse ceiling. The 4K data are packed C4, so using a full 4K document for every row in the calculator is an attention-work upper scenario, not an audited document-length histogram. The learned-routing reference also changes HybridEP to blocking mode; its slowdown is not a controlled routing-imbalance-only measurement.

A concrete discrepancy: metadata wait versus byte service

In the inspected middle-stage trace, all 1,536 same-shape routing-map all-gathers were effectively exposed relative to useful model compute. Their kernel residency times varied far more than a constant-byte model predicts. NCCL residency includes peer waiting and contention; it is not a direct bandwidth sample.

Routing-map all-gather Duration Interpretation
Default nominal byte-service sketch 220.5 MiB / 900 GB/s, under the assumed all-local EP placement. Not measured.
Observed median 1.398 ms Identical payload, one profiled rank/update.
Observed P95 36.600 ms A large long tail; not explained just by message size.
Observed maximum 277.746 ms Includes progress/waiting, not only active transfer.

Metadata calls overlapping FSDP all-gathers had a 31.464-ms median, versus 1.029 ms without that overlap. This is evidence of correlated communication delay, not proof of one cause: rank arrival skew, contention and earlier compute stalls can coincide. Summing overlapping NCCL bars would double-count wall time; measure interval unions and the wait before the consumer can start.

Calibrate the compute roof

Benchmark the actual grouped expert shapes and dense GEMMs at the real power cap. Separately time KDA, MLA, SiTU-GLU, norms and routing; peak BF16 contraction rates do not price these kernels.

Validate the fabric mapping

Map ranks to NVLink domains and NIC budgets. EP64 fitting inside NVL72 is a placement possibility, not proof that this allocation achieved it.

Match the collectives

Use process-group, sequence number, dtype and element count. Routing metadata and FSDP parameters can have the same NCCL kernel-family name.

Measure the dependency gap

Find when the router map is ready, when the all-gather finishes, and when scan/dispatch starts. Across ranks, compare arrival skew and the downstream PP receive.

The lab omits CP-specific communication, optimizer service from its main step sketch, backend chunk work, pointwise kernels, CPU/launch cost, collective startup, imbalance, allocator effects and non-ideal overlap. Its resource-plus-bubble number is an optimistic diagnostic sketch, not a speedup promise, validated prediction, or complete training speed of light.

Sources and reproducible accounting

  1. Kimi K3 technical report and released configuration: 69 KDA / 24 MLA layers, expert and latent dimensions, attention heads.
  2. NVIDIA GB200 NVL72 specifications: dense versus sparse Tensor Core rates, HBM and NVLink. Rates use decimal bytes; allocation sizes use binary MiB/GiB.
  3. TorchTitan FLOP accounting, Kimi K3 implementation, and DistMuon: conventional model FLOPs, parameter and optimizer shapes. The lifecycle and profile observations are specific to the reference recipe.
  4. PyTorch pipeline parallelism and FSDP: scheduling, mixed precision and sharding semantics.
  5. Approach inspired by Edward Yang's DeepSeek-V3 roofline study. Background: the original roofline paper (Berkeley PDF) and All About Rooflines.

The calculation source is available as plain JavaScript, with shared structural accounting used by the memory study. Export the current stage ledger to inspect the inputs and numbers; the export contains assumptions, not profiler measurements. Fixed reference numbers are summarized from the associated training notes without internal run URLs, dataset paths or host identifiers.