Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

An ML engineer is using Vertex AI for distributed training of a PyTorch model across multiple nodes. The training job must use TPUs for high throughput. The engineer sets up the job configuration. Which THREE components are required for the training to work correctly? (Select 3)

⚠ Common exam trap

Google Cloud often tests the distinction between TensorFlow and PyTorch distributed training configurations, and the trap here is assuming that `TF_CONFIG` or `MultiWorkerMirroredStrategy` are universal for all frameworks, when in fact PyTorch uses its own environment variables and the `torch-xla` library for TPU training.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

A startup script to configure the TPU pod (e.g., `xla_lib.sh`)

A is correct because TPU pods require a startup script (e.g., `xla_lib.sh`) to initialize the XLA runtime, configure the TPU mesh, and set environment variables like `XRT_TPU_CONFIG`. Without this script, the TPU devices will not be discoverable by the PyTorch/XLA process, causing the training to fail with device-not-found errors.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    A startup script to configure the TPU pod (e.g., `xla_lib.sh`)

    Why this is correct

    Startup scripts are often needed to initialize TPU devices.

  • ✗

    A MultiWorkerMirroredStrategy configuration

    Why it's wrong here

    This is for TensorFlow data parallelism, not PyTorch TPU.

  • ✓

    A Docker image that includes PyTorch and the TPU library (torch-xla)

    Why this is correct

    The container must have the necessary dependencies for TPU training.

  • ✗

    A TF_CONFIG environment variable set for each worker

    Why it's wrong here

    TF_CONFIG is for TensorFlow distributed training, not PyTorch.

  • ✓

    A CustomJob with a TPU accelerator type (e.g., v3-32)

    Why this is correct

    TPU training requires specifying the TPU type in the job.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.