Capacity planning pipeline from intent to parallelism

From Scaling Laws to Cluster Size: Capacity Planning That Survives Contact With GPUs

From Scaling Laws to Cluster Size: Capacity Planning That Survives Contact With GPUs Scaling laws tell you how loss should move with tokens and parameters. Clusters tell you what you can actually buy. Capacity planning is the bridge — and most teams skip it until the first failed launch burns a week of calendar time and a pile of GPU-hours. This post is a reproducible planning pipeline: tokens → steps → FLOPs → GPU-hours → forced parallelism. The arithmetic lives in playground/capacity_plan.py; the figures below are generated from its JSON output. Not exact. Accurate enough to catch 10× fantasy plans. ...

February 2, 2025 · 4 min · Duo An