PMLE Scaling Prototypes into ML Models Practice Question
You are running a distributed training job on Vertex AI using PyTorch and the DistributedDataParallel (DDP) strategy across 4 nodes, each with 8 GPUs. You notice that the training loss is not decreasing as expected and the job occasionally hangs. You suspect a communication issue between nodes. Which of the following should you check first?
⚠ Common exam trap
The trap here is assuming that changing batch size or framework will fix communication hangs, when the first step should be verifying the distributed training environment variables and network connectivity.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Ensure that the MASTER_ADDR and MASTER_PORT environment variables are correctly set and that all nodes can communicate over the network.
Communication issues in multi-node PyTorch DDP training often stem from incorrect MASTER_ADDR or MASTER_PORT settings, or network connectivity problems. These variables are critical for establishing the rendezvous point for all processes. Checking them first is a logical diagnostic step before considering other changes, as they are common causes of hangs and synchronization failures.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Switch to using the Horovod framework for distributed training.
Why it's wrong here
Switching frameworks is a major change and does not directly address the suspected communication issue. Horovod has its own configuration and may introduce new problems. It's better to first diagnose the current setup, as the issue is likely a misconfiguration or network problem that could affect any framework. This option is a costly workaround rather than a diagnostic step.
- ✗
Reduce the number of nodes to 2 to decrease communication overhead.
Why it's wrong here
Reducing nodes may lessen communication overhead but does not fix the underlying issue if it's a configuration or network problem. The job might still hang with fewer nodes. This is a temporary workaround that does not address the root cause and may reduce training efficiency. It's better to diagnose the communication setup first.
- ✗
Increase the batch size per GPU to improve gradient synchronization.
Why it's wrong here
Increasing batch size does not address communication issues and may worsen hanging if the network cannot handle the increased data transfer. It also does not fix misconfigured environment variables. While larger batches can improve throughput, they are not a solution for communication hangs or synchronization failures in distributed training.
- ✓
Ensure that the MASTER_ADDR and MASTER_PORT environment variables are correctly set and that all nodes can communicate over the network.
Why this is correct
In distributed training with PyTorch DDP, the MASTER_ADDR and MASTER_PORT are used for initializing the process group and coordinating communication. If these are misconfigured or if there are network issues preventing nodes from reaching each other, the job may hang or fail to synchronize gradients. Checking these first is essential for diagnosing communication problems in multi-node setups.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.