Courseiva

PMLE · domain

Serving and Scaling Models

This domain covers deploying, optimizing, and operating models on Vertex AI: endpoints, custom containers, batch vs online prediction, autoscaling, GPUs/TPUs, and Vector Search indexes. Questions are scenario-based, asking you to choose the configuration, index type, or rollout strategy that meets latency, throughput, cost, and freshness constraints.

102 questions23 easy53 medium26 hard

Focused practice

Practice Serving and Scaling Models questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Serving and Scaling Models

Be able to deploy a model to a Vertex AI Endpoint, split traffic for canary rollout, pick the right Vector Search index for latency versus recall, and cut cold starts. The single most important thing: match the serving configuration to the stated latency, throughput, and freshness requirement.

Configuring Vertex AI Endpoints with traffic splits, autoscaling, and deployed model settings

Choosing Vector Search index types (tree-AH, brute force) and StreamUpdate for freshness

Reducing cold start with smaller images, model warming, and min replica counts

Serving custom containers via Artifact Registry and Vertex AI Prediction with GPUs

Watch out for

Common Serving and Scaling Models exam traps

  • ▸Assuming streaming index updates are free; they trade latency and cost for freshness versus batch rebuilds.
  • ▸Confusing traffic split (percentage routing) with model version deployment; both are needed for canary rollout.
  • ▸Ignoring min replica count and image size when diagnosing cold start latency on custom containers.

Question index

All Serving and Scaling Models questions (102)

Click any question to see the full explanation, or start a practice session above.

1

You are deploying a PyTorch model for online predictions on Vertex AI. The model expects input tensors and performs GPU-accelerated inference. You want to minimize prediction latency and maximize throughput. Which approach should you use?

Medium
2

A company is deploying a model on Vertex AI for online predictions with strict latency SLOs. The model requires GPU acceleration. Which TWO configurations should they consider to meet the SLOs while optimizing cost?

Medium
3

Which of the following is a benefit of using Vertex AI Endpoints with autoscaling and scale-to-zero?

Easy
4

A company runs a high-throughput inference service on a Vertex AI Endpoint backed by a custom container. During peak hours, the endpoint's CPU utilization rises to 85%, but the autoscaler does not add replicas until utilization exceeds 95%. The team wants the autoscaler to react earlier to keep latency low. They have already deployed the model and cannot change the model artifact. What should they do?

Medium
5

You are using Vertex AI batch prediction and your model requires preprocessing that involves joining two BigQuery tables. The preprocessing logic is complex and must be done before inference. How should you design the pipeline?

Medium
6

Which Vertex AI service is best suited for finding similar items in a large dataset based on embedding vectors, such as product recommendations or image similarity search?

Easy
7

A company needs to perform real-time similarity search on a dataset of 10 million embedding vectors. They expect low latency (under 10ms) and high throughput. Which index type should they use in Vertex AI Vector Search?

Hard
8

You are deploying a large deep learning model on Vertex AI endpoints. The model requires GPU acceleration and you want to minimize cold-start latency. Which TWO actions should you take? (Choose 2 correct answers)

Medium
9

You have trained a scikit-learn model and want to deploy it to Vertex AI for online predictions. You need to minimize the effort to create a custom container and ensure the model is served with the default pre-built container. What should you do?

Easy
10

A retail company has deployed a scikit-learn model to a Vertex AI endpoint. The model's predictions are used to personalize the homepage. During a flash sale, the endpoint experiences a sudden 10x traffic spike, and the autoscaling configuration is set to minReplicaCount=1, maxReplicaCount=3. The endpoint becomes unresponsive. You need to modify the deployment to handle similar spikes while keeping costs low during normal hours. What should you do?

Medium
11

You are deploying a model to a Vertex AI endpoint for online predictions. You need to ensure that the endpoint can handle traffic spikes and that predictions are served with low latency. Which TWO of the following configurations should you apply? (Choose two.)

Medium
12

You have a Vertex AI endpoint with two deployed models: a champion (v1) and a challenger (v2). You set the traffic split to 90% v1 and 10% v2. After a week, you observe that v2 has better business metrics. You want to shift all traffic to v2 gradually over 3 days to avoid any risk. What should you do?

Hard
13

You need to run batch predictions on 10 TB of text data stored in BigQuery using a custom container model hosted in Vertex AI. What is the most cost-effective and simple approach?

Medium
14

You need to perform batch predictions on 10 TB of data stored in BigQuery using Vertex AI. The model requires some preprocessing that cannot be expressed in SQL. What is the most scalable approach?

Medium
15

You are using Vertex AI Prediction with a custom container that requires a large model file (5 GB). Deployment takes 10 minutes to start. You want to reduce cold start latency. Which action would be MOST effective?

Hard
16

A team wants to run a Vertex AI pipeline that deploys a model, runs a smoke test against the endpoint, and automatically rolls back if the smoke test fails. They need the deployment step to be reversible and the endpoint to remain available during the update. Which approach should they use?

Medium
17

You have a Vertex AI endpoint that serves a model for real-time predictions. You want to update the model to a new version with zero downtime. Which approach should you take?

Easy
18

A financial services company deploys a fraud detection model on a Vertex AI Endpoint. The model must process each transaction in under 50 ms. The team notices that p99 latency spikes to 200 ms every few minutes. Logs show that the model container performs a cold start when new replicas are added, and the autoscaler frequently adds and removes replicas. The endpoint currently has minReplicaCount=1 and maxReplicaCount=10. What should they do to reduce the latency spikes while controlling cost?

Hard
19

You are deploying a PyTorch model on Vertex AI using a custom container with NVIDIA Triton Inference Server. The model is a large transformer that requires GPU. You want to optimize GPU utilization and reduce memory footprint. Which technique should you apply?

Hard
20

Your team is deploying a large recommendation model on Vertex AI endpoints using GPUs. You need to minimise latency while optimising cost. The model serves many similar requests from the same users within short time windows. Which additional service would best reduce latency and cost?

Hard
21

A company runs a Vertex AI endpoint that serves a model for real-time predictions. The endpoint uses a custom container that loads a 10 GB model into memory. During a traffic spike, the autoscaler adds new replicas, but each new replica takes 8 minutes to become ready because it must download the model from Cloud Storage. The team wants to reduce scale-up time. Which approach is most effective?

Hard
22

A company wants to cache predictions for identical requests to reduce latency and cost. They use Vertex AI Prediction with a custom container. Which GCP service should they use to implement prediction caching?

Medium
23

A retail company serves a product-ranking model on a Vertex AI endpoint. Traffic is highly predictable: a steady baseline all day with a sharp peak every evening. During the evening peak, prediction latency exceeds the SLO for several minutes before autoscaling stabilises. The team wants to reduce this scale-up lag without over-provisioning hardware for the entire day. Which configuration should they apply to the deployed model?

Medium
24

A company uses Vertex AI Matching Engine for real-time recommendations. They need to serve queries with low latency and support frequent updates. Which two configurations are appropriate? (Choose 2)

Medium
25

A data scientist wants to perform A/B testing between two model versions deployed on the same Vertex AI endpoint. They need to route 10% of traffic to the challenger model. Which approach should they use?

Hard
26

Which Vertex AI service is designed for building and managing approximate nearest neighbor (ANN) indexes for similarity search at scale?

Easy
27

A company needs to reduce inference latency for their online prediction service on Vertex AI. Which two actions would help? (Choose 2)

Medium
28

A company is migrating from an on-premises ML serving infrastructure to Vertex AI. They have multiple models that need to be served from the same endpoint with different traffic percentages. They also need to monitor prediction quality. Which THREE actions should they take? (Choose 3)

Hard
29

A financial services company uses a Vertex AI Endpoint to serve a credit risk model. The model must always be available, even during maintenance windows, and they need to control the exact distribution of traffic across two model versions for a gradual rollout. They also want to minimize cold-start latency. Which deployment configuration should they use?

Medium
30

An engineer deploys a model to a Vertex AI endpoint with minReplicas=1 and maxReplicas=3. The endpoint receives a sudden traffic spike, but it does not scale up beyond 1 replica. The CPU utilization target is 60%. What is the most likely cause?

Medium
31

A machine learning engineer needs to run batch predictions on 50 TB of data stored in BigQuery using a Vertex AI model. The model is a custom container. What is the most efficient way to set up the batch prediction job?

Easy
32

Your team is serving a large language model on a Vertex AI endpoint using a custom container. You need to reduce inference latency for long prompts while keeping the deployment cost reasonable. The model uses an autoregressive decoder. Which optimization should you implement?

Hard
33

You are designing a batch prediction pipeline using Vertex AI. The input data is 100 TB of images stored in Cloud Storage. The model is a custom TensorFlow model that expects TFRecord format. The pipeline must be cost-effective and run within a time window of 2 hours. Which THREE steps should you include?

Hard
34

A data science team needs to serve multiple versions of the same ML model on Vertex AI Endpoints for A/B testing. They want to gradually shift traffic from the current 'champion' model to a new 'challenger' model. Which feature should they use?

Medium
35

A retail company deploys a new recommendation model alongside the current champion on Vertex AI Endpoints. They want to gradually shift traffic to the challenger while monitoring business metrics (conversion rate). Which two steps are required? (Choose 2)

Hard
36

What is the primary purpose of Vertex AI Edge Manager?

Easy
37

Which Vertex AI feature allows you to reduce the size of a trained model to improve inference speed on edge devices without significant accuracy loss?

Easy
38

An organization is deploying a mission-critical model on Vertex AI Endpoints. They need to ensure high availability and meet a strict SLO of 99.9% uptime. Which THREE steps should they take? (Choose 3)

Hard
39

You have trained a scikit-learn model and saved it as a joblib file in Cloud Storage. You need to deploy this model to Vertex AI for online predictions with minimal effort. What should you do?

Easy
40

A company is deploying a new model version to an existing Vertex AI endpoint. They want to test the new version with 5% of traffic before fully rolling it out. What is the correct approach?

Medium
41

You deploy a new version of a model to a Vertex AI endpoint and want to gradually shift traffic from the old version to the new version over 24 hours. The endpoint currently serves 100% traffic to the old version. What should you do?

Easy
42

A company uses Vertex AI Vector Search for similarity search. They have a dataset of 10 million 512-dimensional vectors. Which index type should they choose for lowest latency at high recall?

Medium
43

An ML engineer needs to update a model deployed on a Vertex AI endpoint without downtime. They want to gradually shift traffic to the new version while monitoring for errors. What is the correct procedure?

Medium
44

You need to serve a large embedding model for similarity search with low latency. The model was trained to generate 256-dimensional embeddings. You plan to use Vertex AI Vector Search. Which index type should you choose to balance accuracy and performance for a dataset with 10 million vectors?

Medium
45

You have a Vertex AI endpoint serving a model that returns predictions in about 200 ms. During a marketing campaign, traffic increases tenfold for short bursts. You want the endpoint to handle the bursts without manual intervention and without over-provisioning for the entire day. What should you do?

Easy
46

You are deploying a model to a Vertex AI Endpoint and need to reduce inference latency for a real-time application. Which two actions should you take? (Choose two.)

Medium
47

You are A/B testing a new model version (challenger) against the current version (champion) on Vertex AI. You want to gradually shift traffic from champion to challenger while measuring business metrics. Which approach should you use?

Medium
48

Your team has deployed a model on Vertex AI endpoints and you are planning an A/B test to compare a new challenger model (v2) against the current champion (v1). The test should measure business metrics such as click-through rate. Which THREE steps should you take to set up the A/B test correctly? (Choose 3 correct answers)

Hard
49

A company needs to run batch predictions on 10 TB of data stored in Cloud Storage. The predictions should be written to BigQuery. Which approach should they use?

Medium
50

You need to deploy a PyTorch model for online inference on Vertex AI but the model was trained using custom ops that are not natively supported. You want to use NVIDIA Triton Inference Server for optimisation. How should you proceed?

Medium
51

An ML engineer needs to run batch predictions on 10 TB of data stored in BigQuery using a TensorFlow model. The predictions must be written to BigQuery. Which service should they use?

Medium
52

A company needs to serve a high-throughput prediction service with strict latency requirements. They want to minimize cold starts and ensure consistent performance. Which endpoint configuration is most appropriate?

Medium
53

An application serving predictions from a Vertex AI endpoint receives many identical requests within a short time window. The team notices redundant computation and wants to cache responses to reduce latency and cost. What is the recommended solution?

Medium
54

You are designing a batch prediction pipeline using Vertex AI. The input data is 50 TB in CSV format on GCS. The model requires feature engineering that involves complex transformations (e.g., datetime parsing, one-hot encoding). Which TWO services or steps should you include in your pipeline?

Hard
55

A company uses Vertex AI Vector Search (Matching Engine) for a product recommendation system. The product embeddings are updated hourly. Which index update method should they use to ensure low latency for new items?

Medium
56

A machine learning engineer wants to deploy a trained model to Vertex AI for online predictions. Which Vertex AI resource is required to serve the model and provide an endpoint URL?

Easy
57

A startup is deploying a scikit-learn model to Vertex AI for online predictions. They want to minimize the effort required to containerize the model and ensure it can handle HTTP requests. What should they do?

Easy
58

Your Vertex AI endpoint receives many identical prediction requests (same input features). You want to cache responses to reduce latency and cost. Which Google Cloud service should you use?

Easy
59

Your company runs a high-traffic web application that serves the same machine learning model prediction for many identical requests (e.g., product recommendations for the same user profile). You want to reduce latency and load on the prediction endpoint by caching responses. Which Google Cloud service should you use?

Easy
60

An ML platform team deploys the same custom container to two Vertex AI endpoints: one for interactive scoring and one for nightly bulk scoring. The interactive endpoint must return predictions in under 200 ms and receives small single-record requests. The bulk endpoint sends large batched requests and tolerates seconds of latency. Both endpoints currently use the same machine type and the same container image, and the interactive endpoint frequently misses its latency target. Which change best resolves the interactive latency problem?

Hard
61

You are deploying a model to a Vertex AI endpoint and need to minimize latency for online predictions. Which machine type should you choose?

Easy
62

A media company uses a Vertex AI endpoint to serve a video recommendation model. The model is updated weekly with new embeddings. They want to minimize downtime during model updates and ensure that the new model performs well before fully rolling it out. They also need to be able to revert quickly if issues arise. What should they do?

Medium
63

You need to run a batch prediction job on Vertex AI for a large dataset stored in BigQuery. The model expects CSV input. Which input format should you specify for the batch prediction job?

Easy
64

You have a champion model serving 100% traffic on a Vertex AI endpoint. You want to deploy a challenger model and gradually shift 10% of traffic to it for A/B testing. What is the correct approach?

Medium
65

A team is deploying a large PyTorch model for online inference. They want to use NVIDIA Triton Inference Server to optimize serving performance. How can they integrate Triton with Vertex AI?

Medium
66

You are deploying a custom PyTorch model to a Vertex AI Endpoint for real-time inference. The model artifact is stored in a Cloud Storage bucket. Your security team requires that the model be served from a container that runs as a non-root user and has no network access except to the Vertex AI prediction service. Which deployment configuration should you use?

Medium
67

A retail team must run nightly batch predictions over 20 TB of Parquet data stored in Cloud Storage using a custom PyTorch model registered in Vertex AI Model Registry. They want the job to finish within a fixed maintenance window and prefer not to manage the underlying compute. Which configuration should they use?

Medium
68

A team has deployed a model on Vertex AI and wants to cache frequent identical prediction requests to improve latency and reduce cost. Which Google Cloud service should they use?

Medium
69

A company wants to run batch predictions on millions of records stored in BigQuery. They need to preprocess the data (e.g., feature engineering) before feeding it to the model. Which approach is most scalable and cost-effective?

Medium
70

A machine learning team deploys a PyTorch model for online prediction on Vertex AI using a custom container. They notice that the first few requests after scaling up experience high latency. What is the most likely cause and how should they mitigate it?

Medium
71

A team is building a batch prediction pipeline that processes raw data from Cloud Storage, performs complex preprocessing, and then runs predictions using a large model. The preprocessing step is compute-intensive and the prediction step is I/O-bound. Which TWO Google Cloud services should they combine to optimize cost and performance? (Choose 2)

Hard
72

Which TWO of the following can be used as input sources for Vertex AI batch prediction jobs? (Choose 2)

Easy
73

You have deployed a model to a Vertex AI Endpoint and need to perform a canary release of a new model version to 10% of traffic. You want to monitor the new version's performance before gradually increasing its traffic share. What should you do?

Medium
74

You are deploying a scikit-learn model to a Vertex AI endpoint for real-time inference. Prediction requests arrive as JSON payloads containing a single instance per request, and the model's predict method expects a pandas DataFrame with named columns. You want to avoid writing a custom container. Which approach should you take?

Medium
75

You are optimizing a model for deployment on Vertex AI using NVIDIA Triton Inference Server. Which TWO actions can you take to improve inference performance?

Medium
76

Your team serves a model on a Vertex AI endpoint with autoscaling. During a flash sale, traffic jumps from 50 to 900 requests per second within one minute, and many requests time out with 429 responses before new replicas become ready. You want to absorb the burst with the least user-visible impact. What should you do?

Hard
77

You have a TensorFlow model that you want to deploy on edge devices for real-time inference. The model was trained in Vertex AI. You need to convert it to a format suitable for on-device inference. Which approach should you use?

Medium
78

You need to serve multiple models on a single Vertex AI endpoint to reduce costs. How can you achieve this?

Easy
79

A data scientist wants to deploy a trained TensorFlow model to Vertex AI for online predictions. They need to serve predictions with low latency and want to leverage GPU acceleration. Which machine type should they select when creating the Vertex AI endpoint?

Easy
80

Your team has deployed a model to a Vertex AI endpoint and wants to route a small percentage of live traffic to a new model version for evaluation. You need to split traffic at the endpoint level without changing the client application. What should you do?

Medium
81

You are deploying a PyTorch model on Vertex AI and want to use NVIDIA Triton Inference Server for optimal performance. You have built a custom container with Triton. Which serving configuration should you use?

Hard
82

You deployed a model to a Vertex AI endpoint with minReplicas=0 and maxReplicas=5. After sending prediction requests, you notice the endpoint takes about 30 seconds to respond initially, but subsequent requests are fast. What is the most likely cause?

Easy
83

You are deploying a model to a Vertex AI endpoint that will serve predictions for a mobile application. The application sends a single request per user action and expects a response within 100 ms. The model is small and CPU-bound. You want to minimize cost while meeting the latency requirement. Which endpoint configuration should you choose?

Medium
84

You are deploying a large language model on a Vertex AI endpoint. The model is loaded from a Cloud Storage bucket at container startup, which adds 3 minutes to each cold start. You want to reduce cold-start time and ensure predictable latency during scale-out. Which approach should you take?

Medium
85

A company uses Vertex AI Matching Engine for a product recommendation system. They need to update the index with new product embeddings every hour, but the index is used for online queries with low latency. Which index update strategy should they use?

Hard
86

A fintech company needs to deploy a TensorFlow model for real-time fraud detection with strict latency SLO (p99 < 100ms). They expect variable traffic with spikes. They also want to minimize cold-start latency. Which two configurations should they use? (Choose 2)

Hard
87

Your team has deployed a model on Vertex AI endpoints. You need to monitor the prediction latency to ensure it meets a 99th percentile SLO of 500ms. You want to set up an alert if the latency exceeds this threshold. Which metric should you use?

Medium
88

You are serving a model on a Vertex AI endpoint that requires a GPU. The model is used for interactive predictions with a strict latency SLO. You notice that during peak hours, some requests time out because the endpoint's autoscaler is slow to add GPU replicas. Which action should you take to meet the SLO?

Hard
89

A media company is serving a video recommendation model on a Vertex AI Endpoint. The model receives a mix of requests: some require only a few features, while others require many features from a feature store. The team wants to reduce average latency and cost without retraining the model. Which TWO strategies should they use? (Choose two.)

Medium
90

A fraud-detection model is deployed on a Vertex AI endpoint and must respond within 30 ms for 95% of requests. During testing, the team sees that p95 latency is dominated by feature retrieval from an external online store, not by model inference. They want to reduce latency without retraining the model. What should they do first?

Hard
91

You are using Vertex AI Matching Engine for similarity search. Your index has 10 million embeddings of 512 dimensions. The query latency requirement is under 10ms for 99th percentile. Which index type should you choose?

Medium
92

You are using Vertex AI Vector Search with an approximate nearest neighbor index. You need to update the index with new data every hour. The updates must be available for queries immediately. Which update method should you use?

Medium
93

You are using Vertex AI Vector Search for a product recommendation system. Your index is updated with new embeddings every hour. To minimize query latency while keeping the index fresh, what should you do?

Hard
94

You need to run batch predictions on a large dataset stored in BigQuery using a Vertex AI model. The dataset contains 10 million rows, and each prediction takes about 100ms. You want to minimize cost and execution time. What should you do?

Medium
95

A financial services firm serves a fraud-detection model on a Vertex AI endpoint that consumes features from a Vertex AI Feature Store online store. During a load test, prediction latency is acceptable, but the firm discovers that the model's feature values in production drift from the values used at training time because the training pipeline read from a BigQuery table with different transformation logic. The team wants the serving path to use the same feature definitions as training so online and offline values match. Which approach should they take?

Hard
96

Which API is recommended for high-throughput, low-latency online prediction requests to Vertex AI endpoints?

Easy
97

You are deploying a model to a Vertex AI Endpoint that will serve predictions to a global user base. You want to minimize latency for users in different regions while ensuring high availability. What should you do?

Medium
98

An organization wants to deploy a TensorFlow model on edge devices such as smartphones and IoT devices for offline inference. Which format should they export the model to?

Medium
99

You need to query a Vertex AI Vector Search index for nearest neighbours. The index is deployed on an endpoint. Which API method should you use to perform the query?

Medium
100

A Vertex AI Endpoint hosts a model that must serve predictions with a strict 99th percentile latency under 100 ms. The model is a large TensorFlow model that processes images. During load testing, you observe that p99 latency spikes to 300 ms when batch size exceeds 1. You need to meet the latency SLO while maintaining reasonable throughput. What should you do?

Hard
101

You have a custom model deployed on a Vertex AI endpoint that receives online prediction requests. The model expects input features in a specific order, but clients sometimes send features in a different order. You want to ensure that the endpoint consistently receives correctly ordered features without modifying every client. What should you do?

Medium
102

A healthcare analytics team needs to serve a model on Vertex AI to internal applications, but compliance requires that no prediction request or response payload ever be written to logs. They still want basic operational metrics such as request count and latency. What should they configure on the endpoint?

Easy

Frequently asked questions

What does the Serving and Scaling Models domain cover on the PMLE exam?
Be able to deploy a model to a Vertex AI Endpoint, split traffic for canary rollout, pick the right Vector Search index for latency versus recall, and cut cold starts. The single most important thing: match the serving configuration to the stated latency, throughput, and freshness requirement.
How many questions are in this domain?
This page lists all 102 Serving and Scaling Models questions in the PMLE question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Serving and Scaling Models questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
google-pmle GOOGLE-PMLE pmle serving scaling Practice Questions