PMLE Serving and Scaling Models Practice Question
You are deploying a PyTorch model for online predictions on Vertex AI. The model expects input tensors and performs GPU-accelerated inference. You want to minimize prediction latency and maximize throughput. Which approach should you use?
⚠ Common exam trap
A common pitfall is assuming that model optimization alone (e.g., quantization) is sufficient for low-latency serving, when in fact the inference server's request handling and batching capabilities are critical for minimizing latency and maximizing throughput in online predictions on Vertex AI.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Deploy using a prebuilt PyTorch serving container with NVIDIA Triton Inference Server.
NVIDIA Triton Inference Server provides advanced features like dynamic batching, concurrent model execution, and GPU scheduling that maximize throughput and minimize latency for GPU-accelerated inference. Vertex AI's prebuilt PyTorch serving container with Triton is specifically designed to handle online prediction workloads efficiently, outperforming a plain custom container without an inference server.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Package the model in a custom container without any inference server.
Why it's wrong here
Omitting an inference server means you must implement batching, GPU memory management and HTTP handling yourself, typically increasing latency and limiting throughput. A custom container without a server suits bespoke runtimes, but Vertex AI's prebuilt PyTorch server provides optimised GPU serving out of the box.
- ✓
Deploy using a prebuilt PyTorch serving container with NVIDIA Triton Inference Server.
Why this is correct
NVIDIA Triton Inference Server provides dynamic batching and concurrent model execution, directly maximising GPU utilisation for PyTorch tensors. This satisfies the stem's twin constraints of minimising prediction latency and maximising throughput, whereas single-request serving containers leave GPU capacity idle between calls.
- ✗
Use Vertex AI Model Optimization to quantize the model to FP16 and deploy using the optimized model.
Why it's wrong here
FP16 quantisation reduces precision and memory footprint but does not itself guarantee lower latency or higher throughput for tensor inputs; it is intended for shrinking models to fit smaller accelerators or cut cost. The scenario demands GPU-accelerated online inference with minimal latency, which requires selecting an appropriate GPU-backed endpoint instead.
- ✗
Use batch prediction instead of online prediction to reduce latency.
Why it's wrong here
Batch prediction processes large jobs asynchronously and returns results after completion, so it cannot serve real-time online requests; latency increases rather than decreases. It is the right choice for offline scoring of large datasets where throughput matters and immediate responses are unnecessary, not for interactive online predictions.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
Same concept, more angles
1 more way this is tested on PMLE
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. A company uses Vertex AI for online predictions with a large ensemble model that requires GPU acceleration. They want to reduce inference latency by batching multiple requests into a single GPU inference call. What should they configure?
hard- A.Use Vertex AI Model Optimization for automatic compilation.
- ✓ B.Deploy the model with NVIDIA Triton Inference Server configured for dynamic batching.
- C.Increase the number of GPU replicas to handle higher concurrency.
- D.Enable model quantization using TensorRT.
Why B: NVIDIA Triton Inference Server supports dynamic batching, which automatically groups multiple inference requests into a single GPU call. This reduces overhead and improves GPU utilization, directly addressing the need to lower latency for online predictions with a large ensemble model on Vertex AI.
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.