Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

You are training a TensorFlow model on Vertex AI using a custom container. The training job uses a single node with 4 GPUs and a global batch size of 1024. You notice that the training is slower than expected and GPU utilization is low. You suspect the input pipeline is the bottleneck. Which of the following should you do to improve training throughput?

⚠ Common exam trap

The trap here is assuming that adding more hardware or changing batch size will solve low GPU utilization, when the real issue is data starvation from an inefficient input pipeline.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use tf.data to prefetch data and parallelize data extraction and transformation with num_parallel_calls and prefetch.

Low GPU utilization in multi-GPU training often indicates the input pipeline cannot supply data fast enough. Using tf.data with prefetch and parallel calls overlaps data preprocessing with model training, keeping the GPUs busy. This is a standard optimization for TensorFlow input pipelines and directly targets the bottleneck without changing hardware or batch size.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Reduce the global batch size to 512 to decrease memory pressure on the GPUs.

    Why it's wrong here

    Reducing batch size might improve memory usage but does not address the input pipeline bottleneck. In fact, smaller batches may increase the frequency of data loading, potentially worsening the bottleneck. The issue is not memory pressure but data starvation, so this change is unlikely to improve throughput and may even reduce training efficiency.

  • ✓

    Use tf.data to prefetch data and parallelize data extraction and transformation with num_parallel_calls and prefetch.

    Why this is correct

    Optimizing the input pipeline with tf.data by using prefetch and parallel calls allows data to be prepared while the GPU is training. This overlaps I/O and preprocessing with computation, increasing GPU utilization and throughput. It directly addresses the bottleneck by ensuring the GPU is not waiting for data, which is a common cause of low utilization in multi-GPU training.

  • ✗

    Increase the number of GPUs to 8 and double the global batch size to 2048.

    Why it's wrong here

    Adding more GPUs and increasing batch size may improve throughput if the input pipeline can keep up, but if the input pipeline is already the bottleneck, it will only exacerbate the problem. More GPUs will demand more data, and if the pipeline cannot supply it, utilization will remain low. This approach does not address the root cause and may increase cost without benefit.

  • ✗

    Switch to using TPUs instead of GPUs for the training job.

    Why it's wrong here

    Switching to TPUs changes the hardware but does not fix the input pipeline bottleneck. If the data pipeline cannot feed the current GPUs, it likely cannot feed TPUs either. TPUs require efficient input pipelines as well, and the bottleneck would persist. This is a costly change that does not address the underlying issue.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.