Duo's Tech Blog
  • Archive
  • Search
  • Tags
  • Categories

Tags

  • All-Reduce 1
  • All-to-All 1
  • AMP 1
  • Attention 2
  • BF16 1
  • Capacity Planning 1
  • Communication 2
  • Context Parallelism 1
  • CoT 1
  • CUDA 4
  • Curriculum 1
  • Data Parallel 1
  • Data Parallelism 1
  • Data Pipeline 2
  • DDP 3
  • DeepSpeed 5
  • Diffusion 1
  • Distillation 1
  • Distributed Systems 1
  • Distributed Training 11
  • Dynamo 1
  • Embeddings 1
  • Expert Parallel 2
  • FlashAttention 1
  • GPipe 1
  • GPU 3
  • Grouped GEMM 1
  • GShard 1
  • Inductor 1
  • Kernel 2
  • Kernel Fusion 1
  • LLM 20
  • Long Context 1
  • Loss Scaling 1
  • MegaScale 1
  • Megatron 8
  • Megatron-LM 2
  • Memory 2
  • MFU 1
  • Mixed Precision 1
  • Modal 2
  • Model Parallel 1
  • MoE 4
  • Multimodal 2
  • NCCL 2
  • Nsight 1
  • Optimization 1
  • Optimizer 2
  • Paper Reading 1
  • Parallelism 6
  • Performance 10
  • Pipeline Parallel 1
  • Pipeline Parallelism 1
  • Profiling 2
  • PyTorch 4
  • Qwen 1
  • Roofline 1
  • Scalability 1
  • Scaling Laws 1
  • Sequence Parallelism 1
  • SigLIP 1
  • SiQ-VL 1
  • Streaming 1
  • Systems 6
  • Tensor Parallel 2
  • Tensor Parallelism 1
  • torch.compile 1
  • Training 7
  • Transformer 2
  • Triton 1
  • Ulysses 1
  • VLM 4
  • ZeRO 2
© 2026 Duo's Tech Blog ยท Powered by Hugo & PaperMod