Ring Attention keeps query blocks fixed on each GPU while key/value blocks move around a communication ring

Sequence Parallelism III: Ring Attention for Context That Does Not Fit

Sequence Parallelism III: Ring Attention for Context That Does Not Fit Megatron SP reduces replicated activation memory around tensor-parallel blocks. DeepSpeed Ulysses uses All-to-All to turn sequence shards into head shards for attention. Ring Attention changes the unit of work again. It asks each rank to keep a block of queries fixed, then circulate key/value blocks around a ring until every query block has seen every key/value block it needs. That is the core idea in Ring Attention with Blockwise Transformers for Near-Infinite Context (arXiv:2310.01889). ...

May 26, 2025 · 8 min · Duo An
Ulysses sequence-shards activations and uses All-to-All to make each rank own all tokens for one attention head

Sequence Parallelism II: DeepSpeed Ulysses and All-to-All Attention

Sequence Parallelism II: DeepSpeed Ulysses and All-to-All Attention Megatron SP is a careful memory optimization around an existing tensor-parallel block. DeepSpeed Ulysses starts from a different question. What if each device owns a sequence slice most of the time, but attention temporarily wants each device to own a head slice instead? The answer in DeepSpeed Ulysses is an All-to-All transpose (arXiv:2309.14509). Before attention, every rank has all heads for a subset of tokens. After All-to-All, every rank has all tokens for a subset of heads. That one layout change lets local attention run per head while the rest of the layer can remain sequence-sharded. ...

May 19, 2025 · 8 min · Duo An
Megatron-LM tensor parallel MLP with column and row splits

Tensor Parallelism in Megatron-LM: Splitting Layers, Not Stacks

Tensor Parallelism in Megatron-LM: Splitting Layers, Not Stacks Pipeline parallelism splits a model by depth. Tensor parallelism splits the math inside one layer. Megatron-LM made tensor parallelism practical for Transformers by choosing split points that preserve local GEMMs and put collectives only where the algebra requires them. The original Megatron paper introduced the core intra-layer pattern (arXiv:1909.08053). The later Megatron-LM systems paper showed how that TP axis composes with data and pipeline parallelism at cluster scale (arXiv:2104.04473). ...

May 18, 2025 · 8 min · Duo An
DeepSpeed-Megatron MoE initialization flow creating expert-parallel groups during model wrapping

MoE Internals: DeepSpeed-Megatron Expert Parallel Implementation

MoE Internals: DeepSpeed-Megatron Expert Parallel Implementation DeepSpeed-Megatron MoE is easy to misread if you only inspect one initialization function. Megatron builds the usual TP, PP, and DP topology. DeepSpeed then wraps the model, finds MoE modules, and gives those modules expert-parallel groups. The topology is split across the training framework and the MoE layer implementation. This post connects the pieces: deepspeed.initialize(), EP groups, expert-DP groups, MoELayer, and the difference between a DeepSpeed MoE wrapper and a Megatron-integrated SwitchMLP-style module. For routing, capacity, and All-to-All fundamentals, start with MoE Parallelism Principles. ...

May 17, 2025 · 8 min · Duo An
Tensor parallelism leaves LayerNorm and Dropout activations replicated while Megatron SP shards them along sequence

Sequence Parallelism I: Megatron SP Cuts Activation Memory Along the Sequence

Sequence Parallelism I: Megatron SP Cuts Activation Memory Along the Sequence Tensor parallelism is usually introduced as a way to make matrix multiplications fit. That is true, but it hides a second problem. After the weights are split, many activations are still replicated on every tensor-parallel rank. For short contexts this is tolerable. For long contexts it becomes one of the reasons training throughput collapses into activation checkpointing. Megatron sequence parallelism, usually shortened to Megatron SP, is a targeted fix from Reducing Activation Recomputation in Large Transformer Models (arXiv:2205.05198). It does not replace tensor parallelism. It keeps Megatron’s column-parallel and row-parallel linear layers, then shards the sequence-local regions that tensor parallelism had left replicated. The trick is small enough to miss and important enough to change the memory budget of a whole Transformer block. ...

May 12, 2025 · 8 min · Duo An