Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

An ML engineer is training a very large PyTorch model on Vertex AI using a TPU v3 pod. The training is slower than expected, and the TPU utilization is low. What is the most likely cause?

⚠ Common exam trap

Google often tests the misconception that low utilization is caused by model architecture or hyperparameter issues, when in reality the most common bottleneck in distributed TPU training is the data pipeline, not the compute or memory limits.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The data pipeline is a bottleneck; the TPU is waiting for data.

The most likely cause of low TPU utilization is a data pipeline bottleneck, where the TPU spends a significant amount of time idle waiting for the next batch of data to be loaded and preprocessed. TPU v3 pods are designed for high-throughput matrix operations and can process data far faster than a typical CPU-based data loader can supply it, especially if the data pipeline uses inefficient I/O, lacks prefetching, or has insufficient workers. This mismatch starves the TPU, leading to low utilization and slower training.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    The data pipeline is a bottleneck; the TPU is waiting for data.

    Why this is correct

    TPUs stall when input tensors arrive slower than the accelerator can consume them, so low utilisation with slow training points to host-side starvation. Feeding data via tf.data with parallel extraction and prefetching keeps the TPU v3 pod saturated rather than idle waiting on I/O.

  • ✗

    The learning rate schedule is too aggressive.

    Why it's wrong here

    An aggressive learning rate schedule affects convergence quality and can cause divergence, not low TPU utilisation. Tuning the schedule is the right move when loss is unstable or training fails to converge, not when the accelerator sits idle.

  • ✗

    The model is using a single TensorFlow operation not supported by TPU.

    Why it's wrong here

    A single unsupported TensorFlow operation would typically fail compilation or fall back, not steadily underuse a TPU pod. Checking operation support matters when porting TensorFlow models to TPUs, whereas this scenario concerns PyTorch data feeding starving the device.

  • ✗

    The batch size is too large for the TPU memory.

    Why it's wrong here

    An oversized batch would raise memory pressure or trigger out-of-memory errors, not leave TPU utilisation low while training runs slowly. Large batches are chosen deliberately to improve throughput when TPU memory and pipeline feeding already permit it.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

4 more ways this is tested on PMLE

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. A machine learning engineer is preparing to train a Transformer-based model using TensorFlow on a single TPU v3-8 pod slice. The training script uses tf.distribute.TPUStrategy. Which environment variable must be set in Vertex AI to enable TPU training with the appropriate topology?

medium
  • A.TPU_NAME
  • B.XRT_TPU_CONFIG
  • ✓ C.TPU_CONFIG
  • D.TF_CONFIG

Why C: Vertex AI automatically sets the TPU_CONFIG environment variable to communicate the TPU worker IP address and port to the training container. TF_CONFIG is used for distributed training with CPUs/GPUs, but TPU_CONFIG is the correct one for TPU training.

Variation 2. A team is training a custom TensorFlow model on Vertex AI using a pre-built container. They need to use a TPU pod slice (v3-32). What THREE actions are required to set up the training job correctly?

hard
  • ✓ A.Configure TF_CONFIG for distributed training
  • B.Set the training worker pool to use only one worker
  • ✓ C.Specify the accelerator type as TPU_V3 and topology as '2x2x4'
  • D.Use a custom container with TensorFlow 2.12 and TPU support
  • E.Set the machine type to a high-memory VM with NVIDIA A100 GPUs

Why A: Option A is correct because a TPU pod slice (v3-32) is a distributed training environment, and TensorFlow requires the TF_CONFIG environment variable to define the cluster specification (worker, master, and TPU worker roles) so each worker process knows its task type and index. Option C is correct because Vertex AI custom training jobs specify TPU accelerators via acceleratorType: TPU_V3 and a topology string; for a v3-32 pod slice the topology is '2x2x4' (which equals 32 TPU v3 cores), and this must be set in the worker pool's machine spec. Option B is wrong because a TPU pod slice inherently requires multiple workers (v3-32 has 32 cores across 4 hosts), so restricting the worker pool to one worker would not provision the pod slice. Option D is wrong because the scenario explicitly uses a pre-built container, so a custom container is unnecessary and contradicts the requirement. Option E is wrong because NVIDIA A100 GPUs are GPU accelerators, not TPUs, and would not satisfy the TPU v3-32 pod slice requirement.

Variation 3. You are running a Vertex AI custom training job with pre-built TensorFlow container. You want to use TPU v3 pods for faster training. Which configuration is required?

medium
  • ✓ A.Specify machine type as tpu-v3-8 in the worker pool spec and use tf.distribute.TPUStrategy
  • B.Set the worker pool machine type to n1-standard-8 and add --accelerator type=tpu
  • C.Use a custom container with TPU libraries and set --tpu-topology=v3-8
  • D.Enable TPU by setting the environment variable TPU_NAME

Why A: To use TPU v3 pods in a Vertex AI custom training job with a pre-built TensorFlow container, you specify the TPU machine type (e.g., tpu-v3-8) in the worker pool spec and use tf.distribute.TPUStrategy in your training code. The pre-built container already includes the TPU libraries, so no custom container is needed, and the distribution strategy is required for the code to address the TPU cores correctly.

Variation 4. An ML engineer is using Vertex AI for distributed training of a PyTorch model across multiple nodes. The training job must use TPUs for high throughput. The engineer sets up the job configuration. Which THREE components are required for the training to work correctly? (Select 3)

medium
  • ✓ A.A startup script to configure the TPU pod (e.g., `xla_lib.sh`)
  • B.A MultiWorkerMirroredStrategy configuration
  • ✓ C.A Docker image that includes PyTorch and the TPU library (torch-xla)
  • D.A TF_CONFIG environment variable set for each worker
  • ✓ E.A CustomJob with a TPU accelerator type (e.g., v3-32)

Why A: A is correct because TPU pods require a startup script (e.g., `xla_lib.sh`) to initialize the XLA runtime, configure the TPU mesh, and set environment variables like `XRT_TPU_CONFIG`. Without this script, the TPU devices will not be discoverable by the PyTorch/XLA process, causing the training to fail with device-not-found errors.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.