Courseiva

PMLE Serving and Scaling Models Practice Question

You are deploying a PyTorch model on Vertex AI and want to use NVIDIA Triton Inference Server for optimal performance. You have built a custom container with Triton. Which serving configuration should you use?

⚠ Common exam trap

PMLE often tests the misconception that prebuilt Vertex AI containers can be toggled to use Triton via environment variables, when in fact a custom container is required.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Upload your Triton container to Container Registry and specify it as the prediction container in Vertex AI Model.

To deploy a custom Triton Inference Server container on Vertex AI, you upload your container to Artifact Registry (or Container Registry) and specify it as the prediction container when creating the Vertex AI Model resource. Vertex AI supports custom containers for prediction, allowing you to run Triton with your model artifacts. This is the standard approach for custom serving frameworks.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Deploy the model on GKE with Triton and expose via Istio.

    Why it's wrong here

    GKE with Istio bypasses Vertex AI's managed prediction endpoint, so the custom Triton container cannot be registered as a Vertex AI Model with a prediction route. It is tempting when Kubernetes-native autoscaling is desired, but the stem requires Vertex AI serving, which accepts custom containers directly.

  • ✗

    Use the prebuilt Vertex AI PyTorch prediction container and set environment variables to enable Triton.

    Why it's wrong here

    The prebuilt PyTorch container runs TorchServe, not Triton, and its environment variables cannot substitute a custom Triton build. It is tempting because prebuilt containers reduce setup effort, but they are correct only when serving standard PyTorch models without a bespoke inference server.

  • ✗

    Use Vertex AI Model Optimization to automatically convert the model to TensorRT and deploy with built-in server.

    Why it's wrong here

    Model Optimization converts models to TensorRT for the built-in server; it does not deploy a custom Triton container, so the built Triton image is unused. It is tempting for GPU latency gains, but it applies when using Vertex AI's managed serving stack rather than a user-supplied Triton container.

  • ✓

    Upload your Triton container to Container Registry and specify it as the prediction container in Vertex AI Model.

    Why this is correct

    Vertex AI accepts a custom container as the prediction container, so pushing the Triton image to Container Registry and referencing it lets Vertex AI route prediction traffic through Triton's dynamic batching and concurrent model execution, satisfying the requirement to serve PyTorch via Triton.

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

2 more ways this is tested on PMLE

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. You need to deploy a PyTorch model for online inference on Vertex AI but the model was trained using custom ops that are not natively supported. You want to use NVIDIA Triton Inference Server for optimisation. How should you proceed?

medium
  • A.Convert the model to TFLite and deploy on an edge device.
  • ✓ B.Build a custom container with NVIDIA Triton Inference Server and deploy it to Vertex AI.
  • C.Export the model to ONNX and deploy using Vertex AI's built-in TensorFlow serving.
  • D.Use Vertex AI Model Optimisation to automatically quantise the model.

Why B: Vertex AI supports custom containers for prediction, so you can package NVIDIA Triton Inference Server with your PyTorch model and any custom ops/libraries it needs, then deploy that container as a Vertex AI Model. Triton natively supports PyTorch (via TorchScript), ONNX, TensorRT, and custom backends, making it the right choice when the model relies on non-standard operators that Vertex AI's pre-built PyTorch/TensorFlow containers cannot execute.

Variation 2. A team is deploying a large PyTorch model for online inference. They want to use NVIDIA Triton Inference Server to optimize serving performance. How can they integrate Triton with Vertex AI?

medium
  • ✓ A.Package the model with Triton in a custom container and deploy it to Vertex AI
  • B.Vertex AI automatically uses Triton for all PyTorch models
  • C.Deploy the model to GKE and use Vertex AI as a frontend
  • D.Use a prebuilt Vertex AI PyTorch container that includes Triton

Why A: Vertex AI supports custom containers for prediction, so the standard integration pattern is to build a container that bundles the model with NVIDIA Triton Inference Server and deploy it as a Vertex AI Model with a custom prediction container. This lets Triton handle dynamic batching, model ensembles, and multi-framework serving while Vertex AI manages endpoints, autoscaling, and monitoring.

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.