Courseiva
Serving and Scaling Models →mediumMultiple Choice

PMLE Serving and Scaling Models Practice Question

You are deploying a PyTorch model for online predictions on Vertex AI. The model expects input tensors and performs GPU-accelerated inference. You want to minimize prediction latency and maximize throughput. Which approach should you use?

⚠ Common exam trap

A common pitfall is assuming that model optimization alone (e.g., quantization) is sufficient for low-latency serving, when in fact the inference server's request handling and batching capabilities are critical for minimizing latency and maximizing throughput in online predictions on Vertex AI.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Deploy using a prebuilt PyTorch serving container with NVIDIA Triton Inference Server.

NVIDIA Triton Inference Server provides advanced features like dynamic batching, concurrent model execution, and GPU scheduling that maximize throughput and minimize latency for GPU-accelerated inference. Vertex AI's prebuilt PyTorch serving container with Triton is specifically designed to handle online prediction workloads efficiently, outperforming a plain custom container without an inference server.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Package the model in a custom container without any inference server.

    Why it's wrong here

    Omitting an inference server means you must implement batching, GPU memory management and HTTP handling yourself, typically increasing latency and limiting throughput. A custom container without a server suits bespoke runtimes, but Vertex AI's prebuilt PyTorch server provides optimised GPU serving out of the box.

  • ✓

    Deploy using a prebuilt PyTorch serving container with NVIDIA Triton Inference Server.

    Why this is correct

    NVIDIA Triton Inference Server provides dynamic batching and concurrent model execution, directly maximising GPU utilisation for PyTorch tensors. This satisfies the stem's twin constraints of minimising prediction latency and maximising throughput, whereas single-request serving containers leave GPU capacity idle between calls.

  • ✗

    Use Vertex AI Model Optimization to quantize the model to FP16 and deploy using the optimized model.

    Why it's wrong here

    FP16 quantisation reduces precision and memory footprint but does not itself guarantee lower latency or higher throughput for tensor inputs; it is intended for shrinking models to fit smaller accelerators or cut cost. The scenario demands GPU-accelerated online inference with minimal latency, which requires selecting an appropriate GPU-backed endpoint instead.

  • ✗

    Use batch prediction instead of online prediction to reduce latency.

    Why it's wrong here

    Batch prediction processes large jobs asynchronously and returns results after completion, so it cannot serve real-time online requests; latency increases rather than decreases. It is the right choice for offline scoring of large datasets where throughput matters and immediate responses are unnecessary, not for interactive online predictions.

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

1 more way this is tested on PMLE

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. A company uses Vertex AI for online predictions with a large ensemble model that requires GPU acceleration. They want to reduce inference latency by batching multiple requests into a single GPU inference call. What should they configure?

hard
  • A.Use Vertex AI Model Optimization for automatic compilation.
  • ✓ B.Deploy the model with NVIDIA Triton Inference Server configured for dynamic batching.
  • C.Increase the number of GPU replicas to handle higher concurrency.
  • D.Enable model quantization using TensorRT.

Why B: NVIDIA Triton Inference Server supports dynamic batching, which automatically groups multiple inference requests into a single GPU call. This reduces overhead and improves GPU utilization, directly addressing the need to lower latency for online predictions with a large ensemble model on Vertex AI.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.