Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

You are fine-tuning a pre-trained BERT model from Hugging Face on a custom text classification dataset using Vertex AI Training. You want to speed up training by using mixed precision. What should you do?

⚠ Common exam trap

It's easy for candidates to confuse framework-level training configuration (fp16=True in TrainingArguments) with infrastructure-level tuning (Vertex AI hyperparameter tuning) or model surgery (half-precision layers), when mixed precision is purely a Trainer API flag.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Set fp16=True in the TrainingArguments

Hugging Face's Trainer API exposes mixed precision through the TrainingArguments parameter fp16=True (or bf16=True for bfloat16). Setting this flag enables automatic mixed precision (AMP) via PyTorch's torch.cuda.amp, which casts eligible operations to FP16 while keeping master weights in FP32 for numerical stability. This is the standard, supported way to accelerate fine-tuning on GPU without rewriting the model.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Modify the model to use half-precision layers

    Why it's wrong here

    Casting the whole model to half precision (fp16) risks numerical instability and gradient underflow during fine-tuning, unlike mixed precision which keeps master weights in fp32 and uses fp16 only for forward and backward passes. It is tempting as the obvious precision reduction, and would suit inference where stability matters less.

  • ✗

    Use a custom container with TensorFlow instead of PyTorch

    Why it's wrong here

    Swapping PyTorch for TensorFlow changes the framework, not the numerical precision of the training loop; mixed precision must still be explicitly enabled via the framework's AMP API. It is tempting because TensorFlow supports mixed precision, and would be correct when framework compatibility, not precision, is the actual constraint.

  • ✗

    Enable mixed precision via Vertex AI hyperparameter tuning

    Why it's wrong here

    Hyperparameter tuning searches for optimal hyperparameter values; it does not enable mixed precision, which is a training-precision setting configured in the training code or container. It is tempting because it also speeds up training, and would be correct when searching learning rates or batch sizes rather than changing numerical precision.

  • ✓

    Set fp16=True in the TrainingArguments

    Why this is correct

    Setting fp16=True in TrainingArguments enables NVIDIA AMP mixed precision, using float16 for forward and backward passes while keeping master weights in float32. This reduces memory and speeds training on compatible GPUs without manual loss scaling.

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.