PMLE Serving and Scaling Models Practice Question
A team is deploying a large PyTorch model for online inference. They want to use NVIDIA Triton Inference Server to optimize serving performance. How can they integrate Triton with Vertex AI?
⚠ Common exam trap
PMLE often tests the misconception that managed services like Vertex AI automatically apply third-party optimizers (Triton, TensorRT) to any framework — in reality, integration requires the candidate to package the optimizer inside a custom container.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Package the model with Triton in a custom container and deploy it to Vertex AI
Vertex AI supports custom containers for prediction, so the standard integration pattern is to build a container that bundles the model with NVIDIA Triton Inference Server and deploy it as a Vertex AI Model with a custom prediction container. This lets Triton handle dynamic batching, model ensembles, and multi-framework serving while Vertex AI manages endpoints, autoscaling, and monitoring.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Package the model with Triton in a custom container and deploy it to Vertex AI
Why this is correct
Vertex AI lets you supply a custom prediction container, so bundling the PyTorch model with Triton and its model repository into that image gives Triton full control of inference, enabling its batching and optimisation features on Vertex AI's managed endpoint infrastructure.
- ✗
Vertex AI automatically uses Triton for all PyTorch models
Why it's wrong here
Vertex AI does not automatically select Triton for PyTorch models; it defaults to its own prebuilt PyTorch containers unless you supply a custom container. Automatic Triton selection would be correct only if Vertex AI offered it as a managed default, which it does not.
- ✗
Deploy the model to GKE and use Vertex AI as a frontend
Why it's wrong here
Running Triton on GKE bypasses Vertex AI's managed prediction endpoint, so the model is not deployed to Vertex AI at all. This approach suits teams wanting full control over serving infrastructure, but it fails the requirement to integrate Triton with Vertex AI's prediction service.
- ✗
Use a prebuilt Vertex AI PyTorch container that includes Triton
Why it's wrong here
Vertex AI's prebuilt PyTorch containers do not bundle Triton Inference Server; they include PyTorch and its serving stack only. A prebuilt Triton container would be correct if Vertex AI published one, but no such prebuilt image exists for this purpose.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.