PMLE Scaling Prototypes into ML Models Practice Question
An engineer is designing a distributed training job on Vertex AI for a TensorFlow model that uses the MultiWorkerMirroredStrategy. They need to ensure proper communication between workers. Which environment variable must be set correctly for each worker?
⚠ Common exam trap
The exam may confuse candidates with plausible but incorrect environment variable names like TF_CONFIG_JSON or TF_DISTRIBUTED_STRATEGY, but only TF_CONFIG is required.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
TF_CONFIG
In TensorFlow distributed training with MultiWorkerMirroredStrategy, the only required environment variable is `TF_CONFIG`. It provides the cluster topology and task identity, enabling gRPC communication between workers. The distribution strategy is defined in code, not via an environment variable. `TF_DISTRIBUTED_STRATEGY` is not a standard TensorFlow environment variable.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
CLUSTER_SPEC
Why it's wrong here
CLUSTER_SPEC is not a standard TensorFlow environment variable.
- ✗
TF_CPP_MIN_LOG_LEVEL
Why it's wrong here
TF_CPP_MIN_LOG_LEVEL controls logging verbosity; it is not required for distributed communication.
- ✗
TF_CONFIG_JSON
Why it's wrong here
TF_CONFIG_JSON is not a valid environment variable; the correct name is TF_CONFIG.
- ✗
TF_DISTRIBUTED_STRATEGY
Why it's wrong here
TF_DISTRIBUTED_STRATEGY is not an environment variable; the strategy is set in code.
- ✓
TF_CONFIG
Why this is correct
TF_CONFIG is the environment variable that carries the cluster specification and task details to each worker, letting MultiWorkerMirroredStrategy identify the chief and workers and establish collective communication. Without it correctly set per worker, distributed training on Vertex AI cannot coordinate.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
Same concept, more angles
1 more way this is tested on PMLE
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. A machine learning engineer is training a TensorFlow model on Vertex AI using distributed training with the MultiWorkerMirroredStrategy. The training job uses 4 workers with 4 GPUs each. The engineer notices that the training is not scaling linearly. What is the most likely cause?
medium- A.The model architecture is too simple to benefit from distribution
- B.The workers are not using the same version of TensorFlow
- ✓ C.Communication overhead due to gradient synchronization
- D.The GPUs are not configured correctly
Why C: With MultiWorkerMirroredStrategy, each worker computes gradients independently on its local batch, then all-reduces gradients across workers via collective communication (e.g., NCCL or gRPC). As the number of workers increases, the communication overhead for gradient synchronization grows, often dominating the per-step time and preventing linear scaling. This is the most common bottleneck in distributed TensorFlow training, especially with many workers or small batch sizes per worker.
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.