Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

An ML engineer is using Vertex AI distributed training for a TensorFlow model that uses the MirroredStrategy. They notice that the training throughput drops significantly when moving from a single GPU to multiple GPUs on the same machine. What is the most likely cause?

⚠ Common exam trap

The trap here is assuming any multi-GPU slowdown must be a configuration error (TF_CONFIG, framework version) rather than recognizing that synchronization overhead in data-parallel training is amortized by batch size — a classic distributed-training performance pitfall.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The batch size is too small, causing each GPU to complete its forward pass quickly, but the sync wait dominates.

In MirroredStrategy, all GPUs must finish their forward/backward pass before the all-reduce gradient synchronization can occur, so throughput is bounded by the slowest replica plus the sync overhead. With a small batch size, each GPU's compute time is very short, so the fixed cost of the NCCL all-reduce (and the sync wait) dominates the step time, making multi-GPU training slower than single-GPU. Increasing the per-replica batch size amortizes the synchronization cost over more compute, restoring scaling efficiency.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The GPUs are not properly configured in TF_CONFIG.

    Why it's wrong here

    TF_CONFIG configures multi-worker, multi-machine clusters; MirroredStrategy handles single-machine multi-GPU synchronisation internally and ignores it, so misconfiguration cannot explain the drop. It is tempting because TF_CONFIG is central to distributed training, and would be correct for a MultiWorkerMirroredStrategy setup across separate hosts.

  • ✓

    The batch size is too small, causing each GPU to complete its forward pass quickly, but the sync wait dominates.

    Why this is correct

    MirroredStrategy performs an all-reduce gradient sync across GPUs after every step. With a small batch size, each GPU's forward and backward pass finishes quickly, so communication overhead dominates and throughput falls rather than scaling with added GPUs.

  • ✗

    The learning rate is too high, causing instability.

    Why it's wrong here

    Throughput loss on scaling GPUs stems from inter-GPU gradient synchronisation overhead, not learning-rate magnitude; an excessive rate causes divergence or oscillation, not reduced throughput. Learning-rate tuning is the correct remedy when loss curves destabilise or fail to converge during training.

  • ✗

    The model uses TensorFlow 1.x instead of 2.x.

    Why it's wrong here

    TensorFlow 1.x lacks the modern MirroredStrategy API and eager execution, but the stem states MirroredStrategy is already in use, so version cannot be the cause. It is tempting because framework upgrades often affect performance, and would be correct if the code relied on deprecated 1.x graph-mode training.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.