A team is fine-tuning a 13B-parameter Llama model with NVIDIA NeMo on a node of eight A100 80GB GPUs. They want the optimizer state to be partitioned across data-parallel ranks so that per-GPU memory drops, while keeping the model replicas synchronized. Which distributed strategy should they select in the NeMo training configuration?
The distributed optimizer in Megatron Core implements ZeRO Stage 1 semantics by sharding the optimizer states, such as Adam first and second moments plus the master weights, across the data-parallel ranks. Each rank keeps a full model replica for the forward and backward pass, so no extra model-parallel communication is introduced, and per-GPU memory falls in proportion to the data-parallel size.
Why this answer
Sharding optimizer state across data-parallel ranks is exactly what the distributed optimizer in Megatron Core does, providing ZeRO Stage 1 behavior with full model replicas per rank. Tensor and pipeline parallelism change how parameters and layers are split, and recomputation targets activations. Only optimizer-state partitioning reduces the Adam moments and master weights each GPU must hold without adding model-parallel communication.
Exam trap
The trap here is assuming that any multi-GPU parallelism mode reduces optimizer memory, when only data-parallel optimizer sharding actually partitions the Adam state.