Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

You are fine-tuning a BERT model from Hugging Face Transformers on Vertex AI. You want to minimise cost for a short experiment. Which compute configuration should you use?

⚠ Common exam trap

PMLE often tests the misconception that more powerful hardware (TPUs or multiple high-end GPUs) is always better, but the key is matching compute to workload scale and cost constraints; candidates may overlook spot VMs as a cost-saving option for short, fault-tolerant jobs.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

A custom training job with a single NVIDIA T4 GPU using spot VMs

A single NVIDIA T4 GPU with spot VMs is the most cost-effective choice for a short BERT fine-tuning experiment on Vertex AI. T4 GPUs are inexpensive and well-suited for moderate training workloads, and spot VMs offer up to 60-70% discount over regular VMs. Since the experiment is short, the risk of preemption is acceptable, and the cost savings are significant. This configuration balances performance and cost effectively.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    A custom training job with a single NVIDIA T4 GPU using spot VMs

    Why this is correct

    A single T4 GPU on spot VMs gives the lowest cost for a short fine-tuning experiment. Spot capacity suits interruptible, brief jobs, and T4 provides sufficient memory and compute for BERT-scale training without paying for premium accelerators.

  • ✗

    A custom training job with a TPU v3-8 pod

    Why it's wrong here

    A TPU v3-8 pod is powerful but priced for sustained workloads and requires TensorFlow/JAX-compatible training code; Hugging Face BERT fine-tuning on PyTorch does not map cleanly. It is the correct choice for large-scale, TPU-optimised training runs, not a short cost-sensitive experiment.

  • ✗

    A custom training job with 8 NVIDIA V100 GPUs using regular VMs

    Why it's wrong here

    Eight V100 GPUs on regular VMs incur substantial hourly charges that a short experiment cannot amortise. This configuration is correct for large-scale distributed training where wall-clock time dominates cost, but for a brief fine-tuning run the accelerator spend outweighs any time saving.

  • ✗

    A standard n1-highmem-8 machine with no accelerator

    Why it's wrong here

    An n1-highmem-8 without accelerators cannot realistically fine-tune BERT within a short window; CPU-only training stretches to days, and Vertex AI bills for the whole duration. It is the right pick for lightweight inference or data preprocessing, not transformer fine-tuning.

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.