Courseiva
Serving and Scaling Models →mediumMultiple Choice

PMLE Serving and Scaling Models Practice Question

A team is deploying a large PyTorch model for online inference. They want to use NVIDIA Triton Inference Server to optimize serving performance. How can they integrate Triton with Vertex AI?

⚠ Common exam trap

PMLE often tests the misconception that managed services like Vertex AI automatically apply third-party optimizers (Triton, TensorRT) to any framework — in reality, integration requires the candidate to package the optimizer inside a custom container.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Package the model with Triton in a custom container and deploy it to Vertex AI

Vertex AI supports custom containers for prediction, so the standard integration pattern is to build a container that bundles the model with NVIDIA Triton Inference Server and deploy it as a Vertex AI Model with a custom prediction container. This lets Triton handle dynamic batching, model ensembles, and multi-framework serving while Vertex AI manages endpoints, autoscaling, and monitoring.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Package the model with Triton in a custom container and deploy it to Vertex AI

    Why this is correct

    Vertex AI lets you supply a custom prediction container, so bundling the PyTorch model with Triton and its model repository into that image gives Triton full control of inference, enabling its batching and optimisation features on Vertex AI's managed endpoint infrastructure.

  • ✗

    Vertex AI automatically uses Triton for all PyTorch models

    Why it's wrong here

    Vertex AI does not automatically select Triton for PyTorch models; it defaults to its own prebuilt PyTorch containers unless you supply a custom container. Automatic Triton selection would be correct only if Vertex AI offered it as a managed default, which it does not.

  • ✗

    Deploy the model to GKE and use Vertex AI as a frontend

    Why it's wrong here

    Running Triton on GKE bypasses Vertex AI's managed prediction endpoint, so the model is not deployed to Vertex AI at all. This approach suits teams wanting full control over serving infrastructure, but it fails the requirement to integrate Triton with Vertex AI's prediction service.

  • ✗

    Use a prebuilt Vertex AI PyTorch container that includes Triton

    Why it's wrong here

    Vertex AI's prebuilt PyTorch containers do not bundle Triton Inference Server; they include PyTorch and its serving stack only. A prebuilt Triton container would be correct if Vertex AI published one, but no such prebuilt image exists for this purpose.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.