PMLE Scaling Prototypes into ML Models Practice Question
An ML engineer is training a TensorFlow model on Vertex AI using a custom training job with a single worker and multiple GPUs. The training script uses tf.distribute.MirroredStrategy. After a few epochs, the job fails with a NCCL timeout error. The engineer confirms the GPUs are healthy and the batch size is reasonable. What should they do to resolve the error?
⚠ Common exam trap
The trap here is assuming that NCCL timeout errors indicate a need for more workers or TF_CONFIG, when in a single-machine multi-GPU setup the issue is typically network interface selection or NCCL version compatibility.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Set the environment variable NCCL_SOCKET_IFNAME to the correct network interface and ensure the container image includes a compatible NCCL version.
NCCL timeout errors in a single-worker, multi-GPU Vertex AI training job usually stem from NCCL being unable to select the correct network interface or from incompatible NCCL versions. Explicitly setting NCCL_SOCKET_IFNAME to the primary interface and using a container image with a validated NCCL build ensures that inter-GPU communication initializes correctly. MirroredStrategy does not require TF_CONFIG, so adding workers or TF_CONFIG would not help.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the number of workers in the worker pool configuration to match the number of GPUs.
Why it's wrong here
MirroredStrategy is designed for synchronous training on multiple GPUs within a single machine. Adding more workers would switch to a multi-worker configuration, which requires MultiWorkerMirroredStrategy and a proper TF_CONFIG environment variable. Simply increasing the worker count on a single-machine job would not fix NCCL communication issues and could cause the job to fail differently.
- ✓
Set the environment variable NCCL_SOCKET_IFNAME to the correct network interface and ensure the container image includes a compatible NCCL version.
Why this is correct
NCCL timeouts on a single machine with multiple GPUs often occur when NCCL cannot determine the correct network interface or when there is a version mismatch between NCCL and the GPU driver or framework. Setting NCCL_SOCKET_IFNAME to the primary network interface (e.g., eth0) and using a pre-built container with a validated NCCL version ensures proper inter-GPU communication and resolves the timeout.
- ✗
Set the environment variable NCCL_DEBUG=INFO and reduce the per-GPU batch size to lower memory pressure.
Why it's wrong here
Enabling NCCL debug logging helps diagnose the cause but does not resolve it. Reducing batch size may alleviate out-of-memory errors, but the job is failing with a NCCL timeout, which typically indicates a communication or configuration problem rather than memory exhaustion. This action does not address the root cause of the timeout.
- ✗
Ensure the training job is configured with a single replica and that the container sets the TF_CONFIG environment variable appropriately for a single-worker, multi-GPU setup.
Why it's wrong here
For a single-worker, multi-GPU job using MirroredStrategy, TF_CONFIG is not required and is typically not set. In fact, setting TF_CONFIG with a worker list can confuse TensorFlow into expecting a multi-worker setup. The NCCL timeout is more likely caused by a mismatch in the NCCL version or network interface configuration, not by missing TF_CONFIG.
Visual reference
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.