PMLE Serving and Scaling Models Practice Question
You are deploying a PyTorch model on Vertex AI using a custom container with NVIDIA Triton Inference Server. The model is a large transformer that requires GPU. You want to optimize GPU utilization and reduce memory footprint. Which technique should you apply?
⚠ Common exam trap
Google often tests the distinction between throughput optimization techniques (like dynamic batching) and memory footprint reduction techniques (like quantization), leading candidates to mistakenly choose dynamic batching when the question specifically asks about reducing memory footprint.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Apply model quantization using TensorRT.
Model quantization using TensorRT reduces the precision of model weights (e.g., from FP32 to FP16 or INT8), which directly decreases GPU memory usage and can improve throughput by enabling faster arithmetic operations on compatible NVIDIA GPUs. This technique is specifically designed to optimize GPU utilization and memory footprint for large transformer models deployed with Triton Inference Server.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable dynamic batching in Triton.
Why it's wrong here
Dynamic batching groups concurrent inference requests into a single GPU execution, raising throughput, but it does not reduce the model's weight or activation memory. It is the correct choice when latency-tolerant traffic underutilises the GPU, not when footprint reduction is the stated objective.
- ✗
Use CPU-only instances to avoid GPU memory issues.
Why it's wrong here
Moving to CPU-only instances removes GPU memory pressure but abandons the GPU acceleration the large transformer requires, collapsing inference latency. CPU deployment suits small models or offline batch scoring, not a GPU-dependent transformer where utilisation and footprint must both be optimised.
- ✗
Increase the number of GPU replicas.
Why it's wrong here
Adding GPU replicas multiplies the memory consumed per model instance; it raises aggregate throughput, not per-GPU efficiency, and does nothing to shrink the transformer's footprint. Horizontal scaling is right when traffic saturates existing replicas, not when the goal is reducing memory per device.
- ✓
Apply model quantization using TensorRT.
Why this is correct
TensorRT quantisation converts FP32 weights to lower precision such as FP16 or INT8, shrinking memory footprint and boosting throughput on NVIDIA GPUs. Triton serves the optimised engine, so GPU utilisation improves while latency drops, satisfying the memory and utilisation constraints.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
Same concept, more angles
1 more way this is tested on PMLE
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. You are optimizing a model for deployment on Vertex AI using NVIDIA Triton Inference Server. Which TWO actions can you take to improve inference performance?
medium- A.Increase the number of model replicas to the maximum.
- ✓ B.Use TensorRT to quantize the model to FP16 or INT8.
- C.Disable model caching to reduce memory usage.
- ✓ D.Enable dynamic batching in Triton to aggregate requests.
- E.Use a larger machine type with more vCPUs.
Why B: Option B is correct because TensorRT can quantize a model to FP16 or INT8, reducing precision and memory bandwidth requirements while leveraging NVIDIA GPU tensor cores, which directly lowers inference latency and increases throughput on Triton. Option D is correct because Triton's dynamic batching aggregates multiple incoming inference requests into a single batch at runtime, improving GPU utilization and throughput without requiring client-side batching. Option A is not appropriate because simply maxing out replicas can waste resources and does not guarantee better performance if the model is not compute-bound or if the machine lacks capacity. Option C is wrong because disabling model caching forces reloads and increases latency rather than improving inference performance. Option E is not ideal because adding vCPUs does not help GPU-bound inference and may not improve Triton throughput.
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.