PMLE Serving and Scaling Models Practice Question
You are optimizing a model for deployment on Vertex AI using NVIDIA Triton Inference Server. Which TWO actions can you take to improve inference performance?
⚠ Common exam trap
Google often tests the misconception that simply adding more replicas or CPU resources will linearly improve inference performance, ignoring the GPU-bound nature of model serving and the importance of batching and precision optimization.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use TensorRT to quantize the model to FP16 or INT8.
Option B is correct because TensorRT can quantize a model to FP16 or INT8, reducing precision and memory bandwidth requirements while leveraging NVIDIA GPU tensor cores, which directly lowers inference latency and increases throughput on Triton. Option D is correct because Triton's dynamic batching aggregates multiple incoming inference requests into a single batch at runtime, improving GPU utilization and throughput without requiring client-side batching. Option A is not appropriate because simply maxing out replicas can waste resources and does not guarantee better performance if the model is not compute-bound or if the machine lacks capacity. Option C is wrong because disabling model caching forces reloads and increases latency rather than improving inference performance. Option E is not ideal because adding vCPUs does not help GPU-bound inference and may not improve Triton throughput.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the number of model replicas to the maximum.
Why it's wrong here
Maximising replicas without matching GPU capacity causes queuing and resource contention, so latency rises instead of throughput. It is tempting because horizontal scaling usually helps, and it would be correct when traffic exceeds one replica's capacity and sufficient GPU quota exists.
- ✓
Use TensorRT to quantize the model to FP16 or INT8.
Why this is correct
TensorRT applies FP16 or INT8 quantisation, reducing weight precision and memory bandwidth while exploiting GPU tensor cores for faster matrix maths. This directly satisfies the stem's inference-performance constraint on Triton, where lower-precision kernels cut latency and raise throughput versus FP32 execution.
- ✗
Disable model caching to reduce memory usage.
Why it's wrong here
Disabling model caching forces Triton to reload weights from storage on each request, adding latency rather than improving throughput. It is tempting as a memory-saving measure, and would be correct when several large models share limited GPU memory and load time is not critical.
- ✓
Enable dynamic batching in Triton to aggregate requests.
Why this is correct
Dynamic batching aggregates multiple inference requests into a single batch on the server side, raising GPU utilisation and throughput without client changes. This directly addresses the inference performance constraint in the stem, since Triton can combine requests arriving within a configurable window, reducing per-request overhead.
- ✗
Use a larger machine type with more vCPUs.
Why it's wrong here
Adding vCPUs does not accelerate GPU-bound Triton inference; the bottleneck is GPU compute, not host CPU. It is tempting because scaling machine size helps CPU-bound workloads, and it would be the right choice when preprocessing or tokenisation on the host CPU saturates before the GPU does.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.