A team is training a large language model on 8 NVIDIA A100 GPUs using PyTorch's DistributedDataParallel (DDP). Profiling shows that all GPUs are frequently idle, waiting for gradient synchronization. The network interconnect between nodes is a 100 Gb Ethernet with TCP/IP, and the model has 13 billion parameters. What is the most effective optimization to reduce the idle time?
InfiniBand with RDMA provides higher bandwidth and lower latency than TCP/IP over Ethernet, reducing the all-reduce time for gradient synchronization. For a 13B parameter model, gradient tensors are large, and the communication overhead dominates. NCCL over RDMA bypasses the CPU and kernel network stack, significantly improving throughput and lowering idle time.
Why this answer
The idle time is caused by slow gradient synchronization over TCP/IP Ethernet. InfiniBand with RDMA and NCCL provides the necessary bandwidth and low latency to reduce all-reduce time, directly addressing the bottleneck. Other options either do not target communication or are secondary optimizations that do not resolve the fundamental interconnect limitation.
Exam trap
The trap here is assuming that increasing batch size or compressing gradients will fully resolve communication stalls, when the underlying high-latency interconnect is the true limiting factor.