PMLE Scaling Prototypes into ML Models Practice Question
You are fine-tuning a BERT model from Hugging Face Transformers on Vertex AI. You want to minimise cost for a short experiment. Which compute configuration should you use?
⚠ Common exam trap
PMLE often tests the misconception that more powerful hardware (TPUs or multiple high-end GPUs) is always better, but the key is matching compute to workload scale and cost constraints; candidates may overlook spot VMs as a cost-saving option for short, fault-tolerant jobs.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
A custom training job with a single NVIDIA T4 GPU using spot VMs
A single NVIDIA T4 GPU with spot VMs is the most cost-effective choice for a short BERT fine-tuning experiment on Vertex AI. T4 GPUs are inexpensive and well-suited for moderate training workloads, and spot VMs offer up to 60-70% discount over regular VMs. Since the experiment is short, the risk of preemption is acceptable, and the cost savings are significant. This configuration balances performance and cost effectively.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
A custom training job with a single NVIDIA T4 GPU using spot VMs
Why this is correct
A single T4 GPU on spot VMs gives the lowest cost for a short fine-tuning experiment. Spot capacity suits interruptible, brief jobs, and T4 provides sufficient memory and compute for BERT-scale training without paying for premium accelerators.
- ✗
A custom training job with a TPU v3-8 pod
Why it's wrong here
A TPU v3-8 pod is powerful but priced for sustained workloads and requires TensorFlow/JAX-compatible training code; Hugging Face BERT fine-tuning on PyTorch does not map cleanly. It is the correct choice for large-scale, TPU-optimised training runs, not a short cost-sensitive experiment.
- ✗
A custom training job with 8 NVIDIA V100 GPUs using regular VMs
Why it's wrong here
Eight V100 GPUs on regular VMs incur substantial hourly charges that a short experiment cannot amortise. This configuration is correct for large-scale distributed training where wall-clock time dominates cost, but for a brief fine-tuning run the accelerator spend outweighs any time saving.
- ✗
A standard n1-highmem-8 machine with no accelerator
Why it's wrong here
An n1-highmem-8 without accelerators cannot realistically fine-tune BERT within a short window; CPU-only training stretches to days, and Vertex AI bills for the whole duration. It is the right pick for lightweight inference or data preprocessing, not transformer fine-tuning.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.