PMLE Spot VMs Practice Question
An ML team is using Vertex AI to train a deep learning model on a large dataset. To reduce costs, they want to use preemptible VMs for training jobs. However, training must complete within a bounded time. Which strategy should they use?
⚠ Common exam trap
A common misconception is that preemptible VMs are not supported in Vertex AI Training, but they are fully supported as spot VMs. The key to bounded-time completion is checkpointing to Cloud Storage for resumability.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Vertex AI Training with spot VMs and ensure the training code saves checkpoints periodically to Cloud Storage.
Vertex AI Training supports spot VMs (preemptible instances) for cost savings, and periodic checkpointing to Cloud Storage ensures that training can resume from the last saved state if a VM is preempted, allowing the job to complete within a bounded time despite interruptions.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use Cloud TPU instead of GPU; TPUs are not preemptible.
Why it's wrong here
Cloud TPUs are also offered in preemptible form, so this claim is false and does not address bounded completion. It is tempting because TPUs accelerate large deep learning workloads, but the scenario requires a checkpointing and restart strategy on preemptible resources, not a different accelerator type.
- ✗
Use Vertex AI Training without spot VMs, because preemptible VMs are not supported for training.
Why it's wrong here
Preemptible VMs are supported for Vertex AI training, so this premise is false and forgoes the cost saving entirely. It is tempting when guaranteed completion matters, but the correct approach combines preemptible VMs with checkpointing and automatic restart, satisfying both cost and bounded-time requirements.
- ✓
Use Vertex AI Training with spot VMs and ensure the training code saves checkpoints periodically to Cloud Storage.
Why this is correct
Spot VMs suit this scenario because Vertex AI automatically restarts preempted training jobs, satisfying the bounded-time constraint. Periodic checkpointing to Cloud Storage preserves progress across interruptions, so restarts resume from the last saved state rather than beginning again. This combination absorbs preemption while keeping costs low.
- ✗
Use a single powerful non-preemptible VM to avoid interruptions.
Why it's wrong here
A single non-preemptible VM removes interruption risk but abandons the requested cost reduction and cannot scale across the large dataset. It is tempting when a hard deadline forbids restarts, yet the correct strategy uses preemptible VMs with checkpointing and managed restart, preserving savings while bounding completion time.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.