A team is training a large language model using SageMaker with multiple GPUs. They need to reduce training time by splitting the model across devices due to memory constraints. Which distributed training strategy should they use?
Trap 1: SageMaker Distributed Data Parallel (SMDDP)
SMDDP replicates the full model on every GPU and shards only the training data, so each device must still hold all parameters, gradients and optimiser states; it cannot relieve memory pressure when the model itself exceeds a single device. It is the right choice when per-device memory suffices and throughput is the bottleneck.
Trap 2: Data parallelism
Data parallelism keeps a complete model replica on each GPU and splits only the batch, so per-device memory consumption is unchanged and the model still cannot fit. It is tempting because it is the default multi-GPU approach and scales throughput well when each device can already hold the full model.
Trap 3: SageMaker Distributed Model Parallel (SMDMP)
SMDMP is correct but the option name is incomplete; however, model parallelism is the general strategy.
- A
SageMaker Distributed Data Parallel (SMDDP)
Why it fails: SMDDP replicates the full model on every GPU and shards only the training data, so each device must still hold all parameters, gradients and optimiser states; it cannot relieve memory pressure when the model itself exceeds a single device. It is the right choice when per-device memory suffices and throughput is the bottleneck.
- B
Data parallelism
Why it fails: Data parallelism keeps a complete model replica on each GPU and splits only the batch, so per-device memory consumption is unchanged and the model still cannot fit. It is tempting because it is the default multi-GPU approach and scales throughput well when each device can already hold the full model.
- C
SageMaker Distributed Model Parallel (SMDMP)
Why it fails: SMDMP is correct but the option name is incomplete; however, model parallelism is the general strategy.
- D
Model parallelism
Model parallelism partitions the model's layers across GPUs, so each device holds only a fraction of the parameters. This directly addresses the memory constraint in the stem, whereas data parallelism replicates the full model on every device.