PMLE · domain
Serving and Scaling Models
This domain covers deploying, optimizing, and operating models on Vertex AI: endpoints, custom containers, batch vs online prediction, autoscaling, GPUs/TPUs, and Vector Search indexes. Questions are scenario-based, asking you to choose the configuration, index type, or rollout strategy that meets latency, throughput, cost, and freshness constraints.
Focused practice
Practice Serving and Scaling Models questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Serving and Scaling Models
Be able to deploy a model to a Vertex AI Endpoint, split traffic for canary rollout, pick the right Vector Search index for latency versus recall, and cut cold starts. The single most important thing: match the serving configuration to the stated latency, throughput, and freshness requirement.
Configuring Vertex AI Endpoints with traffic splits, autoscaling, and deployed model settings
Choosing Vector Search index types (tree-AH, brute force) and StreamUpdate for freshness
Reducing cold start with smaller images, model warming, and min replica counts
Serving custom containers via Artifact Registry and Vertex AI Prediction with GPUs
Watch out for
Common Serving and Scaling Models exam traps
- ▸Assuming streaming index updates are free; they trade latency and cost for freshness versus batch rebuilds.
- ▸Confusing traffic split (percentage routing) with model version deployment; both are needed for canary rollout.
- ▸Ignoring min replica count and image size when diagnosing cold start latency on custom containers.
Question index
All Serving and Scaling Models questions (102)
Click any question to see the full explanation, or start a practice session above.
You are deploying a PyTorch model for online predictions on Vertex AI. The model expects input tensors and performs GPU-accelerated inference. You want to minimize prediction latency and maximize throughput. Which approach should you use?
Medium2A company is deploying a model on Vertex AI for online predictions with strict latency SLOs. The model requires GPU acceleration. Which TWO configurations should they consider to meet the SLOs while optimizing cost?
Medium3Which of the following is a benefit of using Vertex AI Endpoints with autoscaling and scale-to-zero?
Easy4A company runs a high-throughput inference service on a Vertex AI Endpoint backed by a custom container. During peak hours, the endpoint's CPU utilization rises to 85%, but the autoscaler does not add replicas until utilization exceeds 95%. The team wants the autoscaler to react earlier to keep latency low. They have already deployed the model and cannot change the model artifact. What should they do?
Medium5You are using Vertex AI batch prediction and your model requires preprocessing that involves joining two BigQuery tables. The preprocessing logic is complex and must be done before inference. How should you design the pipeline?
Medium6Which Vertex AI service is best suited for finding similar items in a large dataset based on embedding vectors, such as product recommendations or image similarity search?
Easy7A company needs to perform real-time similarity search on a dataset of 10 million embedding vectors. They expect low latency (under 10ms) and high throughput. Which index type should they use in Vertex AI Vector Search?
Hard8You are deploying a large deep learning model on Vertex AI endpoints. The model requires GPU acceleration and you want to minimize cold-start latency. Which TWO actions should you take? (Choose 2 correct answers)
Medium9You have trained a scikit-learn model and want to deploy it to Vertex AI for online predictions. You need to minimize the effort to create a custom container and ensure the model is served with the default pre-built container. What should you do?
Easy10A retail company has deployed a scikit-learn model to a Vertex AI endpoint. The model's predictions are used to personalize the homepage. During a flash sale, the endpoint experiences a sudden 10x traffic spike, and the autoscaling configuration is set to minReplicaCount=1, maxReplicaCount=3. The endpoint becomes unresponsive. You need to modify the deployment to handle similar spikes while keeping costs low during normal hours. What should you do?
Medium11You are deploying a model to a Vertex AI endpoint for online predictions. You need to ensure that the endpoint can handle traffic spikes and that predictions are served with low latency. Which TWO of the following configurations should you apply? (Choose two.)
Medium12You have a Vertex AI endpoint with two deployed models: a champion (v1) and a challenger (v2). You set the traffic split to 90% v1 and 10% v2. After a week, you observe that v2 has better business metrics. You want to shift all traffic to v2 gradually over 3 days to avoid any risk. What should you do?
Hard13You need to run batch predictions on 10 TB of text data stored in BigQuery using a custom container model hosted in Vertex AI. What is the most cost-effective and simple approach?
Medium14You need to perform batch predictions on 10 TB of data stored in BigQuery using Vertex AI. The model requires some preprocessing that cannot be expressed in SQL. What is the most scalable approach?
Medium15You are using Vertex AI Prediction with a custom container that requires a large model file (5 GB). Deployment takes 10 minutes to start. You want to reduce cold start latency. Which action would be MOST effective?
Hard16A team wants to run a Vertex AI pipeline that deploys a model, runs a smoke test against the endpoint, and automatically rolls back if the smoke test fails. They need the deployment step to be reversible and the endpoint to remain available during the update. Which approach should they use?
Medium17You have a Vertex AI endpoint that serves a model for real-time predictions. You want to update the model to a new version with zero downtime. Which approach should you take?
Easy18A financial services company deploys a fraud detection model on a Vertex AI Endpoint. The model must process each transaction in under 50 ms. The team notices that p99 latency spikes to 200 ms every few minutes. Logs show that the model container performs a cold start when new replicas are added, and the autoscaler frequently adds and removes replicas. The endpoint currently has minReplicaCount=1 and maxReplicaCount=10. What should they do to reduce the latency spikes while controlling cost?
Hard19You are deploying a PyTorch model on Vertex AI using a custom container with NVIDIA Triton Inference Server. The model is a large transformer that requires GPU. You want to optimize GPU utilization and reduce memory footprint. Which technique should you apply?
Hard20Your team is deploying a large recommendation model on Vertex AI endpoints using GPUs. You need to minimise latency while optimising cost. The model serves many similar requests from the same users within short time windows. Which additional service would best reduce latency and cost?
Hard21A company runs a Vertex AI endpoint that serves a model for real-time predictions. The endpoint uses a custom container that loads a 10 GB model into memory. During a traffic spike, the autoscaler adds new replicas, but each new replica takes 8 minutes to become ready because it must download the model from Cloud Storage. The team wants to reduce scale-up time. Which approach is most effective?
Hard22A company wants to cache predictions for identical requests to reduce latency and cost. They use Vertex AI Prediction with a custom container. Which GCP service should they use to implement prediction caching?
Medium23A retail company serves a product-ranking model on a Vertex AI endpoint. Traffic is highly predictable: a steady baseline all day with a sharp peak every evening. During the evening peak, prediction latency exceeds the SLO for several minutes before autoscaling stabilises. The team wants to reduce this scale-up lag without over-provisioning hardware for the entire day. Which configuration should they apply to the deployed model?
Medium24A company uses Vertex AI Matching Engine for real-time recommendations. They need to serve queries with low latency and support frequent updates. Which two configurations are appropriate? (Choose 2)
Medium25A data scientist wants to perform A/B testing between two model versions deployed on the same Vertex AI endpoint. They need to route 10% of traffic to the challenger model. Which approach should they use?
Hard26Which Vertex AI service is designed for building and managing approximate nearest neighbor (ANN) indexes for similarity search at scale?
Easy27A company needs to reduce inference latency for their online prediction service on Vertex AI. Which two actions would help? (Choose 2)
Medium28A company is migrating from an on-premises ML serving infrastructure to Vertex AI. They have multiple models that need to be served from the same endpoint with different traffic percentages. They also need to monitor prediction quality. Which THREE actions should they take? (Choose 3)
Hard29A financial services company uses a Vertex AI Endpoint to serve a credit risk model. The model must always be available, even during maintenance windows, and they need to control the exact distribution of traffic across two model versions for a gradual rollout. They also want to minimize cold-start latency. Which deployment configuration should they use?
Medium30An engineer deploys a model to a Vertex AI endpoint with minReplicas=1 and maxReplicas=3. The endpoint receives a sudden traffic spike, but it does not scale up beyond 1 replica. The CPU utilization target is 60%. What is the most likely cause?
Medium31A machine learning engineer needs to run batch predictions on 50 TB of data stored in BigQuery using a Vertex AI model. The model is a custom container. What is the most efficient way to set up the batch prediction job?
Easy32Your team is serving a large language model on a Vertex AI endpoint using a custom container. You need to reduce inference latency for long prompts while keeping the deployment cost reasonable. The model uses an autoregressive decoder. Which optimization should you implement?
Hard33You are designing a batch prediction pipeline using Vertex AI. The input data is 100 TB of images stored in Cloud Storage. The model is a custom TensorFlow model that expects TFRecord format. The pipeline must be cost-effective and run within a time window of 2 hours. Which THREE steps should you include?
Hard34A data science team needs to serve multiple versions of the same ML model on Vertex AI Endpoints for A/B testing. They want to gradually shift traffic from the current 'champion' model to a new 'challenger' model. Which feature should they use?
Medium35A retail company deploys a new recommendation model alongside the current champion on Vertex AI Endpoints. They want to gradually shift traffic to the challenger while monitoring business metrics (conversion rate). Which two steps are required? (Choose 2)
Hard36What is the primary purpose of Vertex AI Edge Manager?
Easy37Which Vertex AI feature allows you to reduce the size of a trained model to improve inference speed on edge devices without significant accuracy loss?
Easy38An organization is deploying a mission-critical model on Vertex AI Endpoints. They need to ensure high availability and meet a strict SLO of 99.9% uptime. Which THREE steps should they take? (Choose 3)
Hard39You have trained a scikit-learn model and saved it as a joblib file in Cloud Storage. You need to deploy this model to Vertex AI for online predictions with minimal effort. What should you do?
Easy40A company is deploying a new model version to an existing Vertex AI endpoint. They want to test the new version with 5% of traffic before fully rolling it out. What is the correct approach?
Medium41You deploy a new version of a model to a Vertex AI endpoint and want to gradually shift traffic from the old version to the new version over 24 hours. The endpoint currently serves 100% traffic to the old version. What should you do?
Easy42A company uses Vertex AI Vector Search for similarity search. They have a dataset of 10 million 512-dimensional vectors. Which index type should they choose for lowest latency at high recall?
Medium43An ML engineer needs to update a model deployed on a Vertex AI endpoint without downtime. They want to gradually shift traffic to the new version while monitoring for errors. What is the correct procedure?
Medium44You need to serve a large embedding model for similarity search with low latency. The model was trained to generate 256-dimensional embeddings. You plan to use Vertex AI Vector Search. Which index type should you choose to balance accuracy and performance for a dataset with 10 million vectors?
Medium45You have a Vertex AI endpoint serving a model that returns predictions in about 200 ms. During a marketing campaign, traffic increases tenfold for short bursts. You want the endpoint to handle the bursts without manual intervention and without over-provisioning for the entire day. What should you do?
Easy46You are deploying a model to a Vertex AI Endpoint and need to reduce inference latency for a real-time application. Which two actions should you take? (Choose two.)
Medium47You are A/B testing a new model version (challenger) against the current version (champion) on Vertex AI. You want to gradually shift traffic from champion to challenger while measuring business metrics. Which approach should you use?
Medium48Your team has deployed a model on Vertex AI endpoints and you are planning an A/B test to compare a new challenger model (v2) against the current champion (v1). The test should measure business metrics such as click-through rate. Which THREE steps should you take to set up the A/B test correctly? (Choose 3 correct answers)
Hard49A company needs to run batch predictions on 10 TB of data stored in Cloud Storage. The predictions should be written to BigQuery. Which approach should they use?
Medium50You need to deploy a PyTorch model for online inference on Vertex AI but the model was trained using custom ops that are not natively supported. You want to use NVIDIA Triton Inference Server for optimisation. How should you proceed?
Medium51An ML engineer needs to run batch predictions on 10 TB of data stored in BigQuery using a TensorFlow model. The predictions must be written to BigQuery. Which service should they use?
Medium52A company needs to serve a high-throughput prediction service with strict latency requirements. They want to minimize cold starts and ensure consistent performance. Which endpoint configuration is most appropriate?
Medium53An application serving predictions from a Vertex AI endpoint receives many identical requests within a short time window. The team notices redundant computation and wants to cache responses to reduce latency and cost. What is the recommended solution?
Medium54You are designing a batch prediction pipeline using Vertex AI. The input data is 50 TB in CSV format on GCS. The model requires feature engineering that involves complex transformations (e.g., datetime parsing, one-hot encoding). Which TWO services or steps should you include in your pipeline?
Hard55A company uses Vertex AI Vector Search (Matching Engine) for a product recommendation system. The product embeddings are updated hourly. Which index update method should they use to ensure low latency for new items?
Medium56A machine learning engineer wants to deploy a trained model to Vertex AI for online predictions. Which Vertex AI resource is required to serve the model and provide an endpoint URL?
Easy57A startup is deploying a scikit-learn model to Vertex AI for online predictions. They want to minimize the effort required to containerize the model and ensure it can handle HTTP requests. What should they do?
Easy58Your Vertex AI endpoint receives many identical prediction requests (same input features). You want to cache responses to reduce latency and cost. Which Google Cloud service should you use?
Easy59Your company runs a high-traffic web application that serves the same machine learning model prediction for many identical requests (e.g., product recommendations for the same user profile). You want to reduce latency and load on the prediction endpoint by caching responses. Which Google Cloud service should you use?
Easy60An ML platform team deploys the same custom container to two Vertex AI endpoints: one for interactive scoring and one for nightly bulk scoring. The interactive endpoint must return predictions in under 200 ms and receives small single-record requests. The bulk endpoint sends large batched requests and tolerates seconds of latency. Both endpoints currently use the same machine type and the same container image, and the interactive endpoint frequently misses its latency target. Which change best resolves the interactive latency problem?
Hard61You are deploying a model to a Vertex AI endpoint and need to minimize latency for online predictions. Which machine type should you choose?
Easy62A media company uses a Vertex AI endpoint to serve a video recommendation model. The model is updated weekly with new embeddings. They want to minimize downtime during model updates and ensure that the new model performs well before fully rolling it out. They also need to be able to revert quickly if issues arise. What should they do?
Medium63You need to run a batch prediction job on Vertex AI for a large dataset stored in BigQuery. The model expects CSV input. Which input format should you specify for the batch prediction job?
Easy64You have a champion model serving 100% traffic on a Vertex AI endpoint. You want to deploy a challenger model and gradually shift 10% of traffic to it for A/B testing. What is the correct approach?
Medium65A team is deploying a large PyTorch model for online inference. They want to use NVIDIA Triton Inference Server to optimize serving performance. How can they integrate Triton with Vertex AI?
Medium66You are deploying a custom PyTorch model to a Vertex AI Endpoint for real-time inference. The model artifact is stored in a Cloud Storage bucket. Your security team requires that the model be served from a container that runs as a non-root user and has no network access except to the Vertex AI prediction service. Which deployment configuration should you use?
Medium67A retail team must run nightly batch predictions over 20 TB of Parquet data stored in Cloud Storage using a custom PyTorch model registered in Vertex AI Model Registry. They want the job to finish within a fixed maintenance window and prefer not to manage the underlying compute. Which configuration should they use?
Medium68A team has deployed a model on Vertex AI and wants to cache frequent identical prediction requests to improve latency and reduce cost. Which Google Cloud service should they use?
Medium69A company wants to run batch predictions on millions of records stored in BigQuery. They need to preprocess the data (e.g., feature engineering) before feeding it to the model. Which approach is most scalable and cost-effective?
Medium70A machine learning team deploys a PyTorch model for online prediction on Vertex AI using a custom container. They notice that the first few requests after scaling up experience high latency. What is the most likely cause and how should they mitigate it?
Medium71A team is building a batch prediction pipeline that processes raw data from Cloud Storage, performs complex preprocessing, and then runs predictions using a large model. The preprocessing step is compute-intensive and the prediction step is I/O-bound. Which TWO Google Cloud services should they combine to optimize cost and performance? (Choose 2)
Hard72Which TWO of the following can be used as input sources for Vertex AI batch prediction jobs? (Choose 2)
Easy73You have deployed a model to a Vertex AI Endpoint and need to perform a canary release of a new model version to 10% of traffic. You want to monitor the new version's performance before gradually increasing its traffic share. What should you do?
Medium74You are deploying a scikit-learn model to a Vertex AI endpoint for real-time inference. Prediction requests arrive as JSON payloads containing a single instance per request, and the model's predict method expects a pandas DataFrame with named columns. You want to avoid writing a custom container. Which approach should you take?
Medium75You are optimizing a model for deployment on Vertex AI using NVIDIA Triton Inference Server. Which TWO actions can you take to improve inference performance?
Medium76Your team serves a model on a Vertex AI endpoint with autoscaling. During a flash sale, traffic jumps from 50 to 900 requests per second within one minute, and many requests time out with 429 responses before new replicas become ready. You want to absorb the burst with the least user-visible impact. What should you do?
Hard77You have a TensorFlow model that you want to deploy on edge devices for real-time inference. The model was trained in Vertex AI. You need to convert it to a format suitable for on-device inference. Which approach should you use?
Medium78You need to serve multiple models on a single Vertex AI endpoint to reduce costs. How can you achieve this?
Easy79A data scientist wants to deploy a trained TensorFlow model to Vertex AI for online predictions. They need to serve predictions with low latency and want to leverage GPU acceleration. Which machine type should they select when creating the Vertex AI endpoint?
Easy80Your team has deployed a model to a Vertex AI endpoint and wants to route a small percentage of live traffic to a new model version for evaluation. You need to split traffic at the endpoint level without changing the client application. What should you do?
Medium81You are deploying a PyTorch model on Vertex AI and want to use NVIDIA Triton Inference Server for optimal performance. You have built a custom container with Triton. Which serving configuration should you use?
Hard82You deployed a model to a Vertex AI endpoint with minReplicas=0 and maxReplicas=5. After sending prediction requests, you notice the endpoint takes about 30 seconds to respond initially, but subsequent requests are fast. What is the most likely cause?
Easy83You are deploying a model to a Vertex AI endpoint that will serve predictions for a mobile application. The application sends a single request per user action and expects a response within 100 ms. The model is small and CPU-bound. You want to minimize cost while meeting the latency requirement. Which endpoint configuration should you choose?
Medium84You are deploying a large language model on a Vertex AI endpoint. The model is loaded from a Cloud Storage bucket at container startup, which adds 3 minutes to each cold start. You want to reduce cold-start time and ensure predictable latency during scale-out. Which approach should you take?
Medium85A company uses Vertex AI Matching Engine for a product recommendation system. They need to update the index with new product embeddings every hour, but the index is used for online queries with low latency. Which index update strategy should they use?
Hard86A fintech company needs to deploy a TensorFlow model for real-time fraud detection with strict latency SLO (p99 < 100ms). They expect variable traffic with spikes. They also want to minimize cold-start latency. Which two configurations should they use? (Choose 2)
Hard87Your team has deployed a model on Vertex AI endpoints. You need to monitor the prediction latency to ensure it meets a 99th percentile SLO of 500ms. You want to set up an alert if the latency exceeds this threshold. Which metric should you use?
Medium88You are serving a model on a Vertex AI endpoint that requires a GPU. The model is used for interactive predictions with a strict latency SLO. You notice that during peak hours, some requests time out because the endpoint's autoscaler is slow to add GPU replicas. Which action should you take to meet the SLO?
Hard89A media company is serving a video recommendation model on a Vertex AI Endpoint. The model receives a mix of requests: some require only a few features, while others require many features from a feature store. The team wants to reduce average latency and cost without retraining the model. Which TWO strategies should they use? (Choose two.)
Medium90A fraud-detection model is deployed on a Vertex AI endpoint and must respond within 30 ms for 95% of requests. During testing, the team sees that p95 latency is dominated by feature retrieval from an external online store, not by model inference. They want to reduce latency without retraining the model. What should they do first?
Hard91You are using Vertex AI Matching Engine for similarity search. Your index has 10 million embeddings of 512 dimensions. The query latency requirement is under 10ms for 99th percentile. Which index type should you choose?
Medium92You are using Vertex AI Vector Search with an approximate nearest neighbor index. You need to update the index with new data every hour. The updates must be available for queries immediately. Which update method should you use?
Medium93You are using Vertex AI Vector Search for a product recommendation system. Your index is updated with new embeddings every hour. To minimize query latency while keeping the index fresh, what should you do?
Hard94You need to run batch predictions on a large dataset stored in BigQuery using a Vertex AI model. The dataset contains 10 million rows, and each prediction takes about 100ms. You want to minimize cost and execution time. What should you do?
Medium95A financial services firm serves a fraud-detection model on a Vertex AI endpoint that consumes features from a Vertex AI Feature Store online store. During a load test, prediction latency is acceptable, but the firm discovers that the model's feature values in production drift from the values used at training time because the training pipeline read from a BigQuery table with different transformation logic. The team wants the serving path to use the same feature definitions as training so online and offline values match. Which approach should they take?
Hard96Which API is recommended for high-throughput, low-latency online prediction requests to Vertex AI endpoints?
Easy97You are deploying a model to a Vertex AI Endpoint that will serve predictions to a global user base. You want to minimize latency for users in different regions while ensuring high availability. What should you do?
Medium98An organization wants to deploy a TensorFlow model on edge devices such as smartphones and IoT devices for offline inference. Which format should they export the model to?
Medium99You need to query a Vertex AI Vector Search index for nearest neighbours. The index is deployed on an endpoint. Which API method should you use to perform the query?
Medium100A Vertex AI Endpoint hosts a model that must serve predictions with a strict 99th percentile latency under 100 ms. The model is a large TensorFlow model that processes images. During load testing, you observe that p99 latency spikes to 300 ms when batch size exceeds 1. You need to meet the latency SLO while maintaining reasonable throughput. What should you do?
Hard101You have a custom model deployed on a Vertex AI endpoint that receives online prediction requests. The model expects input features in a specific order, but clients sometimes send features in a different order. You want to ensure that the endpoint consistently receives correctly ordered features without modifying every client. What should you do?
Medium102A healthcare analytics team needs to serve a model on Vertex AI to internal applications, but compliance requires that no prediction request or response payload ever be written to logs. They still want basic operational metrics such as request count and latency. What should they configure on the endpoint?
EasyOther domains
All PMLE exam domains
Frequently asked questions
- What does the Serving and Scaling Models domain cover on the PMLE exam?
- Be able to deploy a model to a Vertex AI Endpoint, split traffic for canary rollout, pick the right Vector Search index for latency versus recall, and cut cold starts. The single most important thing: match the serving configuration to the stated latency, throughput, and freshness requirement.
- How many questions are in this domain?
- This page lists all 102 Serving and Scaling Models questions in the PMLE question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Serving and Scaling Models questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.