F = multiply-add FLOPs; B = modeled operand bytes. Arithmetic intensity I = F / B.
Working the roofline
for Kimi K3 on GB200
We made it fit. Now ask which work, bytes, and dependencies set the pace.
Study 01 followed the tensor shapes. Study 02 budgeted the live copies. This study turns those shapes into a performance model: first a GEMM roofline, then a distributed communication ledger, then a check against a real training run.
Default: PP8 / CP1 / TP1 / dense FSDP64 / EP64, FP32 master weights, gradients and Muon momentum, BF16 model GEMMs, full layer checkpointing, one 4K row per microbatch, 64 microbatches, ordinary 1F1B, no CUDA graphs. Hardware ceilings and placement assumptions are not measured kernel or network rates.
Three roofs, three different byte counts
A kernel can be limited by HBM. Distributed work can instead be limited by NVLink or the scale-out network. Divide work by the bandwidth of the resource that actually carries its bytes. In particular, a fast NVLink roof does not make cross-domain traffic disappear.
| Resource | Default ceiling | How to read it |
|---|---|---|
| BF16 Tensor Cores | 2,500 TFLOP/s / GPU | NVIDIA lists 360 PFLOP/s for 72 GPUs with sparsity. Dense = 360 / 72 / 2 = 2.5 PFLOP/s. Expert sparsity is not 2:4 weight sparsity. |
| HBM | 8 TB/s / GPU | 576 TB/s / 72. This is bandwidth, not the 184.31-GiB capacity reported by the training runtime. |
| NVLink | 900 GB/s / GPU, one direction | Half of the advertised 1.8-TB/s bidirectional endpoint bandwidth. Actual collective bandwidth, contention and power caps can lower it. |
| Scale-out network | 50 GB/s / GPU, illustrative | A 400-Gbit/s-equivalent, one-direction budget. It is not a claim about this cluster's NIC allocation or measured throughput. Edit it for the real system. |
Left of the ridge, fewer HBM bytes help. Right of it, the compute ceiling matters. A point under a roof need not attain that roof.
The roofline lab
The lab is an accounting model, not a launch-config generator. Every run has 512 GPUs and TP1. Changing rows or microbatches changes the global batch; changing CP changes the number of independent data replicas. Memory feasibility still needs its own budget.
Stored parameters are not activated FLOPs
For Y = XW, with X [M,K] and W [K,N], forward costs 2MNK FLOPs. Backward has two more contractions of the same size: dX = dYWT and dW = XTdY. Three contractions give 6MNK. Summed over learned matrices, this becomes 6NactiveDtokens. Here N is activated matrix weights, not the entire 2.780T checkpoint.
Routed expert weights activated / token = 92 layers * 16 experts * 3 matrices * 3584 * 3072 = 48,620,371,968 All learned active matrices = 104,174,092,288 Learned-matrix training work = 625.04 GFLOP / token
The familiar 104.1895B active-parameter count also includes small non-matmul weights. It is useful for a conventional metric, but those weights do not each cause a dense matmul. The input embedding lookup and text-only vision path are excluded from per-token matrix compute; the LM head is included.
Attention adds work without adding a weight matrix
| Core | Algorithmic training work / token | What is not in the formula |
|---|---|---|
| 69 KDA layers | 69 × 18 × 96 × 128 × 128 = 1.9535 GFLOP | Three recurrent state contractions, forward plus gradients. The actual chunked kernel also has transforms, triangular solves, FP32 work and memory traffic. |
| 24 causal MLA layers | 24 × 3 × 96 × (192 + 128) × (Seff + 1) | Seff = Σ sdoc2 / Σ sdoc. This counts causal QK and PV contractions, not norms, softmax or gating. |
| FlashAttention replay | 24 × 96 × 192 × (Seff + 1) | The QK contraction rematerialized by attention backward. Counted for the runtime sketch, not conventional model FLOPs. |
| Full layer checkpoint | One extra forward for decoder matrices and attention cores | The model-wide output-head GEMM is not included in the layer checkpoint replay. Full recompute also repeats MoE communication. |
Equal documents make Seff easy to see: one 4K document gives 4,096; four independent 1K documents give 1,024. Two 2K rows have the same tokenwise projection work as one 4K row, but roughly half the MLA contractions. Flattening [rows, context] to T does not remove document boundaries. CP shortens local token arrays, not the logical context attended by MLA.
Model FLOPs, issued-work estimates, MFU and HFU are different quantities
The fixed training results below follow TorchTitan's convention: 6 × 104,189,515,872 + KDA contractions + non-causal MLA contractions, using the configured row context. At 4K that is 645.210 GFLOP/token. The runtime sketch instead uses causal pairs, learned matrices only, FlashAttention replay and the selected layer recomputation.
MFU = model FLOP/s divided by the chosen nominal dense hardware peak. It should not reward checkpoint recomputation. HFU includes actual hardware work, but an algorithmic replay estimate is not a measured HFU: custom-kernel implementation work is still missing. More recompute can make a job look compute-bound while making time-to-result worse.
EP is payload and control plane
K3 compresses the routed stream from 7,168 to 3,584 channels before dispatch. A BF16 latent row is 7,168 bytes, not 14,336. The shared experts stay on the 7,168-wide dense path and are not dispatched. With full checkpointing, the route is traversed three times: forward, recomputed forward, and the two reverse communication operations in backward.
No checkpoint: 4 payload exchanges / layer / microbatch Full checkpoint: 6 payload exchanges / layer / microbatch Assignment budget, uniform routing: remote GPU copies / token = top_k * (1 - 1 / EP) cross-domain copies = top_k * (1 - local_EP_ranks / EP) Payload bytes = exchanges * local_tokens * latent_width * 2 * remote_copies
One-copy-per-assignment is a conservative payload budget. If the dispatcher sends a token once per destination rank and sums results before returning them, multiple chosen experts on that rank can share a copy. The second accounting option models this ideal deduplication under uniform sampling without replacement; it does not assert that every backend or phase implements it. Metadata, alignment and routing weights are additional.
The extra HybridEP all-gather is not an expert weight gather
The 4K reference exchanges a Boolean routing map [4,096, 896] before dispatch: 3.5 MiB locally, 224 MiB gathered across EP64, of which 220.5 MiB came from other ranks. The original and recomputed forward each do it. On a 12-MoE-layer stage, 64 microbatches × 12 layers × 2 = 1,536 metadata all-gathers per update. Small inputs can thus imply a substantial global output and a repeated dependency.
FSDP counts depend on the lifecycle, not just the acronym
At PP8/EP64, routed experts have expert FSDP1: there is no routed-weight FSDP gather. Dense weights are sharded over 64 ranks. A BF16 ring all-gather receives 2P(1 − 1/n) bytes per rank; an FP32 reduce-scatter has the corresponding 4P(1 − 1/n) budget. Multiply by the actual number of collectives.
| Lifecycle model | Weight AG rounds / update | Gradient RS rounds / update | Tradeoff |
|---|---|---|---|
| 4K reference-like | M + 1 = 65 | 1 | 65 gathers per layer observed in the interior-stage trace. Extrapolating that multiplier to other lab shapes is illustrative. |
| Textbook reshard | 2M = 128 | M = 64 | Gather before forward and backward; reduce each microbatch. Lower residency can cost much more communication. |
| Retain views across the schedule | 1 | 1 | An optimistic lifecycle alternative. Keeps BF16 views and deferred gradients live; requires a different memory budget and runtime support. |
For a hierarchical collective, a domain with a of n shards need only import P(1 − a/n) parameters once, then share them locally. Dividing its import by a models the per-rank share of an a-NIC aggregate budget. This is an ideal domain-level network bound, not the measured wire volume of NCCL's chosen algorithm. NIC sharing or a different rank layout changes it.
| Traffic / update | Endpoint budget | NVLink service | NIC-domain budget | Network service |
|---|
Budgets are balanced one-direction endpoint traffic or amortized one-direction domain import; opposite-direction capacity is not added to the bandwidth denominator. The two projections are not disjoint byte categories to add. Different collectives on the same resource compete; overlap does not grant each its own full NVLink or NIC bandwidth.
Long-context CP scenarios keep the logical attention length, DP count and local-token shape straight, but CP attention communication is not priced here. KDA state/prefix exchange, MLA KV/head redistribution and backend-specific CP collectives require their own trace-derived ledger. The 64K preset is a shape and compute comparison, not a complete long-context step-time prediction.
Ordinary 1F1B has a bubble, not a ban on overlap
Under balanced stage costs, an ideal ordinary 1F1B schedule has compute-bubble overhead (PP − 1) / M relative to useful scheduled work. At PP8 and M64 that is 10.94% overhead, or 9.86% of the compute-plus-bubble interval. Each optimizer update fills and drains its own pipeline; a previous warmup update does not remove those structural bubbles.
More microbatches help amortize the bubble, but there is a tradeoff. At fixed global batch, each microbatch gets smaller, potentially lowering GEMM efficiency. In this lab the row shape stays fixed when M changes, so the global batch grows. Science determines whether that larger update is appropriate.
The PP packet is bigger than [T, 7168]
K3 sends (hiddenTD, deltaTND). With one stage per rank, a microbatch visits each receiving rank for the first time, so the residual cache cannot omit the earlier blocks. At the default 12-layer split, outgoing block counts are 2, 4, 6, 8, 10, 12 and 14. The final hop carries one hidden stream plus 14 residual streams: 15 × 56 MiB = 840 MiB forward, and a matching gradient payload backward.
Interleaving with multiple virtual stages can shrink a stage's gathered working set, revisit a rank with a populated residual cache, and offer different scheduling opportunities. It also adds stage boundaries and communication. It is not modeled by simply dividing the 1F1B bubble or memory by VP. The lab deliberately keeps one contiguous stage per rank.
Muon is a matrix workload of its own
FP32 momentum is persistent training history. Newton–Schulz uses a BF16 working matrix. For one expert matrix [3,584, 3,072], orient the working matrix as X [r,c] with r = 3,072 and c = 3,584. A standard iteration forms A = XXT, B = bA + ccoefA2, then updates X using BX and a scaled X. The leading contractions cost 4r2c + 2r3 FLOPs per iteration.
Five iterations cost about 0.966 TFLOP per logical expert matrix. An interior PP rank with 12 MoE layers and 14 experts per layer has 12 × 14 × 3 = 504 such matrices: about 487 TFLOP per update, before the smaller non-routed matrices, redistribution, normalization, pointwise work and communication. That is a workload estimate, not an optimizer timing measurement.
Fused gate/up storage still contains two logical expert matrices. Muon orthogonalizes those matrices separately; using one wider fused storage shape would give a different Newton–Schulz cost and a different update.
Exposed optimizer-phase interval, including clipping/checks/update/scheduler. This is not the duration of a CPU DistMuon annotation.
A long CPU optimizer span can include waiting for the pipeline's last backward or a gradient-norm collective. The reference's complete exposed optimizer tail was small compared with its clean step; it does not support blaming the throughput gap on Muon alone. More tokens per update amortize one optimizer update, but also change optimization and memory requirements.
A roofline gap is a question, not a diagnosis
The fixed 512-GB200 HybridEP references below completed 12 updates with FP32 resident/master weights and optimizer states, BF16 compute, full checkpointing, PP8/FSDP64/EP64, 64 microbatches, and no CUDA graphs. Clean throughput excludes the profiled update and trace-export disturbance. These measurements stay fixed when you change the lab.
| Reference | Rows × context | Tokens/s/GPU | Implied clean update | Model TFLOP/s/GPU | Nominal dense BF16 MFU |
|---|
MFU here uses a fixed 2,500-TFLOP/s dense denominator and the documented TorchTitan estimator, not a 5,000-TFLOP/s sparse ceiling. The 4K data are packed C4, so using a full 4K document for every row in the calculator is an attention-work upper scenario, not an audited document-length histogram. The learned-routing reference also changes HybridEP to blocking mode; its slowdown is not a controlled routing-imbalance-only measurement.
A concrete discrepancy: metadata wait versus byte service
In the inspected middle-stage trace, all 1,536 same-shape routing-map all-gathers were effectively exposed relative to useful model compute. Their kernel residency times varied far more than a constant-byte model predicts. NCCL residency includes peer waiting and contention; it is not a direct bandwidth sample.
| Routing-map all-gather | Duration | Interpretation |
|---|---|---|
| Default nominal byte-service sketch | 220.5 MiB / 900 GB/s, under the assumed all-local EP placement. Not measured. | |
| Observed median | 1.398 ms | Identical payload, one profiled rank/update. |
| Observed P95 | 36.600 ms | A large long tail; not explained just by message size. |
| Observed maximum | 277.746 ms | Includes progress/waiting, not only active transfer. |
Metadata calls overlapping FSDP all-gathers had a 31.464-ms median, versus 1.029 ms without that overlap. This is evidence of correlated communication delay, not proof of one cause: rank arrival skew, contention and earlier compute stalls can coincide. Summing overlapping NCCL bars would double-count wall time; measure interval unions and the wait before the consumer can start.
Calibrate the compute roof
Benchmark the actual grouped expert shapes and dense GEMMs at the real power cap. Separately time KDA, MLA, SiTU-GLU, norms and routing; peak BF16 contraction rates do not price these kernels.
Validate the fabric mapping
Map ranks to NVLink domains and NIC budgets. EP64 fitting inside NVL72 is a placement possibility, not proof that this allocation achieved it.
Match the collectives
Use process-group, sequence number, dtype and element count. Routing metadata and FSDP parameters can have the same NCCL kernel-family name.
Measure the dependency gap
Find when the router map is ready, when the all-gather finishes, and when scan/dispatch starts. Across ranks, compare arrival skew and the downstream PP receive.
The lab omits CP-specific communication, optimizer service from its main step sketch, backend chunk work, pointwise kernels, CPU/launch cost, collective startup, imbalance, allocator effects and non-ideal overlap. Its resource-plus-bubble number is an optimistic diagnostic sketch, not a speedup promise, validated prediction, or complete training speed of light.
Sources and reproducible accounting
- Kimi K3 technical report and released configuration: 69 KDA / 24 MLA layers, expert and latent dimensions, attention heads.
- NVIDIA GB200 NVL72 specifications: dense versus sparse Tensor Core rates, HBM and NVLink. Rates use decimal bytes; allocation sizes use binary MiB/GiB.
- TorchTitan FLOP accounting, Kimi K3 implementation, and DistMuon: conventional model FLOPs, parameter and optimizer shapes. The lifecycle and profile observations are specific to the reference recipe.
- PyTorch pipeline parallelism and FSDP: scheduling, mixed precision and sharding semantics.
- Approach inspired by Edward Yang's DeepSeek-V3 roofline study. Background: the original roofline paper (Berkeley PDF) and All About Rooflines.
The calculation source is available as plain JavaScript, with shared structural accounting used by the memory study. Export the current stage ledger to inspect the inputs and numbers; the export contains assumptions, not profiler measurements. Fixed reference numbers are summarized from the associated training notes without internal run URLs, dataset paths or host identifiers.