Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

A data scientist is training a very large neural network using Vertex AI with multiple GPUs across multiple nodes. The model does not fit on a single GPU, so they need to use both data parallelism and model parallelism (pipeline parallelism). Which THREE components or configurations are required to set up distributed training with Vertex AI?

⚠ Common exam trap

This question tests the misconception that Vertex AI automatically handles model parallelism (e.g., via AutoML or Vizier), when in reality the user must manually implement it in the training script using frameworks like PyTorch or TensorFlow.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Implementing pipeline parallelism manually in the training script using torch.distributed.pipeline.sync.Pipe

Pipeline parallelism requires explicit implementation in the training script, such as using `torch.distributed.pipeline.sync.Pipe` in PyTorch, to split the model layers across multiple GPUs. This is necessary when the model does not fit on a single GPU, and Vertex AI does not automatically handle model parallelism—it must be coded by the user.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Using Vertex AI Vizier to optimize the model parallelism strategy

    Why it's wrong here

    Vizier is for hyperparameter tuning, not for configuring parallelism.

  • ✗

    Enabling Vertex AI AutoML to automatically distribute the model

    Why it's wrong here

    AutoML is for automated model building, not for custom distributed training.

  • ✓

    Implementing pipeline parallelism manually in the training script using torch.distributed.pipeline.sync.Pipe

    Why this is correct

    Manual implementation of pipeline parallelism is required as Vertex AI does not provide built-in model parallelism.

  • ✓

    A custom container with the distributed framework (e.g., PyTorch DDP) installed

    Why this is correct

    Custom container needed to include the framework code and dependencies.

  • ✓

    Setting the --worker-machine-count flag when submitting the job

    Why this is correct

    Specifies the number of worker nodes for distributed training.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.