Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

A machine learning engineer is training a TensorFlow model on Vertex AI using distributed training with the MultiWorkerMirroredStrategy. The training job uses 4 workers with 4 GPUs each. The engineer notices that the training is not scaling linearly. What is the most likely cause?

⚠ Common exam trap

Test-takers frequently assume more workers always means linear speedup, ignoring the fixed overhead of gradient synchronization that becomes the dominant factor in distributed training.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Communication overhead due to gradient synchronization

With MultiWorkerMirroredStrategy, each worker computes gradients independently on its local batch, then all-reduces gradients across workers via collective communication (e.g., NCCL or gRPC). As the number of workers increases, the communication overhead for gradient synchronization grows, often dominating the per-step time and preventing linear scaling. This is the most common bottleneck in distributed TensorFlow training, especially with many workers or small batch sizes per worker.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The model architecture is too simple to benefit from distribution

    Why it's wrong here

    A simple model still scales if communication overhead is low; architecture complexity does not determine MultiWorkerMirroredStrategy efficiency. This option would fit a scenario where the model has too few parameters to occupy the GPUs, not one where four workers with four GPUs each underperform.

  • ✗

    The workers are not using the same version of TensorFlow

    Why it's wrong here

    Version mismatches cause job startup failures or collective-operation errors, not gradual scaling loss. Uniform TensorFlow versions are a prerequisite for launching MultiWorkerMirroredStrategy at all, so this would be the answer when workers refuse to form a cluster.

  • ✓

    Communication overhead due to gradient synchronization

    Why this is correct

    MultiWorkerMirroredStrategy performs all-reduce gradient synchronisation across workers each step, so inter-worker communication overhead grows with worker count and prevents linear scaling. With 4 workers and 16 GPUs, network latency and bandwidth for exchanging gradients, not compute, become the bottleneck limiting throughput gains.

  • ✗

    The GPUs are not configured correctly

    Why it's wrong here

    Misconfigured GPUs would typically cause job failure or errors, not merely sublinear scaling. GPU configuration matters when devices are absent from the visible device list or TensorFlow cannot allocate them, which is a setup fault rather than a scaling bottleneck.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.