Courseiva

CCNA Serving and Scaling Models Questions

75 of 102 questions · Page 1/2 · Serving and Scaling Models · Answers revealed

1
MCQmedium

You are deploying a PyTorch model for online predictions on Vertex AI. The model expects input tensors and performs GPU-accelerated inference. You want to minimize prediction latency and maximize throughput. Which approach should you use?

A.Package the model in a custom container without any inference server.
B.Deploy using a prebuilt PyTorch serving container with NVIDIA Triton Inference Server.
C.Use Vertex AI Model Optimization to quantize the model to FP16 and deploy using the optimized model.
D.Use batch prediction instead of online prediction to reduce latency.
AnswerB

NVIDIA Triton Inference Server provides dynamic batching and concurrent model execution, directly maximising GPU utilisation for PyTorch tensors. This satisfies the stem's twin constraints of minimising prediction latency and maximising throughput, whereas single-request serving containers leave GPU capacity idle between calls.

Why this answer

NVIDIA Triton Inference Server provides advanced features like dynamic batching, concurrent model execution, and GPU scheduling that maximize throughput and minimize latency for GPU-accelerated inference. Vertex AI's prebuilt PyTorch serving container with Triton is specifically designed to handle online prediction workloads efficiently, outperforming a plain custom container without an inference server.

Exam trap

A common pitfall is assuming that model optimization alone (e.g., quantization) is sufficient for low-latency serving, when in fact the inference server's request handling and batching capabilities are critical for minimizing latency and maximizing throughput in online predictions on Vertex AI.

How to eliminate wrong answers

Option A is wrong because a custom container without any inference server lacks request batching, model queuing, and GPU utilization optimizations, leading to higher latency and lower throughput under concurrent requests. Option C is wrong because Vertex AI Model Optimization for FP16 quantization reduces model size and can improve throughput, but it does not address the serving infrastructure needed for low-latency online predictions; the deployment still requires an inference server like Triton to handle request management and GPU scheduling. Option D is wrong because batch prediction is designed for high-throughput, offline processing of large datasets and typically has higher latency per request due to job queuing and resource provisioning, making it unsuitable for minimizing prediction latency in online scenarios.

2
Multi-Selectmedium

A company is deploying a model on Vertex AI for online predictions with strict latency SLOs. The model requires GPU acceleration. Which TWO configurations should they consider to meet the SLOs while optimizing cost?

Select 2 answers
A.Use n1-highmem-32 machine types without GPU
B.Set min_replica_count to handle base traffic and max_replica_count to handle spikes
C.Use GPU-enabled machine types such as n1-standard-4 with T4
D.Enable autoscaling with min_replica_count=0 and max_replica_count=10
E.Disable autoscaling and set a fixed number of replicas equal to peak load
AnswersB, C

Setting min_replica_count to cover baseline traffic keeps GPU-backed replicas warm, eliminating cold-start latency that would breach the SLO, while max_replica_count caps autoscaling during spikes so cost scales only with genuine demand. This directly satisfies the strict latency requirement without over-provisioning GPUs continuously.

Why this answer

Option B is correct because configuring min_replica_count to cover baseline traffic keeps enough warm replicas to serve steady requests within latency SLOs, while max_replica_count allows Vertex AI autoscaling to absorb traffic spikes without over-provisioning permanently. Option C is correct because the model requires GPU acceleration, and GPU-enabled machine types such as n1-standard-4 with an NVIDIA T4 provide the necessary hardware acceleration for online predictions. Option A is incorrect because n1-highmem-32 without a GPU cannot satisfy the GPU acceleration requirement.

Option D is incorrect because setting min_replica_count=0 can cause cold-start delays and scale-from-zero latency that violate strict latency SLOs. Option E is incorrect because disabling autoscaling and fixing replicas at peak load wastes cost and does not optimize for variable traffic.

Exam trap

PMLE often tests the trade-off between cost and latency, and candidates may incorrectly choose min_replica_count=0 to save cost, ignoring the latency impact.

3
MCQeasy

Which of the following is a benefit of using Vertex AI Endpoints with autoscaling and scale-to-zero?

A.It eliminates the need for a load balancer.
B.It reduces costs by scaling down to zero replicas when no requests are received.
C.It reduces model training time.
D.It automatically upgrades the model version.
AnswerB

Scale-to-zero removes all replicas when no requests arrive, so you pay nothing during idle periods while still serving traffic when demand returns. This satisfies the cost-reduction benefit, unlike always-on endpoints that bill for idle capacity.

Why this answer

Vertex AI Endpoints with autoscaling and scale-to-zero allow the number of serving replicas to dynamically adjust based on incoming traffic. When no requests are received, the endpoint can scale down to zero replicas, meaning you are not charged for idle compute resources. This directly reduces operational costs compared to maintaining a minimum number of always-on instances.

Exam trap

A common misconception is that autoscaling eliminates the need for a load balancer, but in Vertex AI Endpoints, the load balancer is a separate component that remains essential for request distribution even when scaling to zero.

How to eliminate wrong answers

Option A is wrong because Vertex AI Endpoints still require a load balancer (the built-in Google Cloud Load Balancer) to distribute incoming requests across replicas; autoscaling does not eliminate this need. Option C is wrong because model training time is a function of training infrastructure and algorithm, not of serving endpoint configuration like autoscaling. Option D is wrong because Vertex AI Endpoints do not automatically upgrade model versions; you must explicitly deploy a new model version or use a traffic split to route requests to a different version.

4
MCQmedium

A company runs a high-throughput inference service on a Vertex AI Endpoint backed by a custom container. During peak hours, the endpoint's CPU utilization rises to 85%, but the autoscaler does not add replicas until utilization exceeds 95%. The team wants the autoscaler to react earlier to keep latency low. They have already deployed the model and cannot change the model artifact. What should they do?

A.Update the endpoint's deployed model to set a lower autoscaling metric threshold, such as 70% CPU utilization.
B.Enable request-response logging on the endpoint and use Cloud Monitoring alerts to manually add replicas when CPU exceeds 70%.
C.Increase the endpoint's minReplicaCount to match the peak traffic level so that replicas are always available.
D.Redeploy the model with a larger machine type so that each replica handles more traffic and CPU never reaches the threshold.
AnswerA

Vertex AI Endpoints expose an autoscaling configuration on the DeployedModel, including the target metric and threshold. Lowering the CPU utilization target to 70% makes the autoscaler add replicas sooner, which reduces latency during ramp-up. This is a configuration change on the deployed model and does not require retraining or replacing the model artifact.

Why this answer

The autoscaler on a Vertex AI Endpoint uses a target metric and threshold defined in the deployed model's autoscaling configuration. Adjusting the CPU utilization target to a lower value makes the system add replicas before latency degrades. Other approaches either do not change the scaling trigger or introduce manual processes that are slower and less reliable.

Exam trap

The trap here is assuming that increasing the minimum replica count changes when the autoscaler scales, when in fact it only raises the floor of running replicas.

5
MCQmedium

You are using Vertex AI batch prediction and your model requires preprocessing that involves joining two BigQuery tables. The preprocessing logic is complex and must be done before inference. How should you design the pipeline?

A.Write a Cloud Composer workflow that runs the preprocessing and then triggers the batch prediction job.
B.Use Dataflow to read from both BigQuery tables, perform the join and preprocessing, write the results to GCS, then run Vertex AI batch prediction with GCS source.
C.Use Vertex AI batch prediction with a custom container that includes logic to read and join tables on the fly.
D.Use BigQuery to create a materialized view that joins the tables and directly use that as the batch prediction source.
AnswerB

Dataflow performs the complex two-table BigQuery join and preprocessing, writing results to Cloud Storage, which Vertex AI batch prediction then reads as its source. This satisfies the constraint that preprocessing must complete before inference, since batch prediction cannot join BigQuery tables itself.

Why this answer

Dataflow (Apache Beam) is designed for complex, stateful data processing like joining two BigQuery tables and performing custom preprocessing. It can read from BigQuery, execute the join logic, and write the preprocessed results to Cloud Storage (GCS). Vertex AI batch prediction then reads the preprocessed data from GCS, which is the recommended pattern for non-trivial transformations before inference, as it decouples preprocessing from prediction and avoids resource contention.

Exam trap

Google often tests the misconception that batch prediction can handle live data transformations within the prediction container, but the correct design is to preprocess data in a separate, scalable data processing service like Dataflow before feeding it to batch prediction.

How to eliminate wrong answers

Option A is wrong because Cloud Composer (Apache Airflow) is an orchestration tool, not a data processing engine; using it to run the preprocessing itself would be inefficient and error-prone, as it lacks native support for large-scale data joins and transformations. Option C is wrong because Vertex AI batch prediction with a custom container that reads and joins tables on the fly violates the principle of separation of concerns, leading to longer inference latency, higher memory usage, and potential timeouts during prediction, as batch prediction expects preprocessed input, not live database joins. Option D is wrong because BigQuery materialized views are precomputed, read-only snapshots that cannot be used directly as a batch prediction source; batch prediction requires input data in GCS (JSON/CSV) or BigQuery tables, but a materialized view is not a table and cannot be referenced as a source URI.

6
MCQeasy

Which Vertex AI service is best suited for finding similar items in a large dataset based on embedding vectors, such as product recommendations or image similarity search?

A.Vertex AI Prediction Endpoint
B.Vertex AI Model Monitoring
C.Vertex AI Feature Store
D.Vertex AI Matching Engine
AnswerD

Matching Engine performs approximate nearest-neighbour search over embedding vectors, returning semantically similar items at scale. It satisfies the stem's similarity-search requirement for recommendations and image search, unlike tabular or forecasting services that do not index vector embeddings.

Why this answer

Vertex AI Matching Engine is specifically designed for high-performance vector similarity search (also known as approximate nearest neighbor search) using embedding vectors. It scales to billions of vectors and is ideal for use cases like product recommendations and image similarity search, where you need to find the most similar items based on dense vector representations.

Exam trap

Candidates often confuse Vertex AI Prediction (model serving) with Vertex AI Matching Engine (vector similarity search). The key distinction is that Prediction serves model inference on input data, while Matching Engine retrieves similar items based on embedding vectors.

How to eliminate wrong answers

Option A is wrong because Vertex AI Prediction Endpoint serves model predictions via HTTP requests but does not provide built-in vector similarity search or indexing capabilities. Option B is wrong because Vertex AI Model Monitoring tracks prediction quality and data drift over time, not similarity search. Option C is wrong because Vertex AI Feature Store is a centralized repository for storing, serving, and sharing feature data, but it does not perform nearest neighbor search on embedding vectors.

7
MCQhard

A company needs to perform real-time similarity search on a dataset of 10 million embedding vectors. They expect low latency (under 10ms) and high throughput. Which index type should they use in Vertex AI Vector Search?

A.Brute-force index
B.Hash-based index
C.Tree-based index
D.Approximate nearest neighbor (ANN) index with ScaNN
AnswerD

ScaNN is Google's ANN algorithm optimised for high-dimensional vectors, using anisotropic vector quantisation to deliver sub-10ms latency and high throughput. It satisfies both stated constraints on 10 million embeddings, whereas exact or tree-based indexes cannot meet that latency at this scale.

Why this answer

For large datasets requiring low latency, an approximate nearest neighbor (ANN) index is appropriate. The Scann algorithm (ScaNN) is used by Vertex AI Vector Search for ANN.

8
Multi-Selectmedium

You are deploying a large deep learning model on Vertex AI endpoints. The model requires GPU acceleration and you want to minimize cold-start latency. Which TWO actions should you take? (Choose 2 correct answers)

Select 2 answers
A.Set minReplicaCount to 0 to allow scale-to-zero.
B.Use a custom container that loads the model during startup.
C.Increase maxReplicaCount to a high number.
D.Use batch prediction instead of online prediction.
E.Set minReplicaCount to 1 to always have at least one replica running.
AnswersB, E

Loading the model during container startup means weights are already in memory when the first request arrives, rather than fetched lazily per request. This directly reduces cold-start latency for the GPU-backed deep learning model described in the stem, complementing replica-warming measures.

Why this answer

Option B is correct because using a custom container that loads the model during startup lets you control and optimize the initialization process (for example, preloading weights and warming up the GPU) so the model is ready as soon as the container starts, which directly reduces cold-start latency. Option E is correct because setting minReplicaCount to 1 keeps at least one replica always provisioned and running on Vertex AI, so incoming requests hit an already-loaded model instead of triggering a new replica spin-up, eliminating the cold start entirely. Option A is wrong because setting minReplicaCount to 0 enables scale-to-zero, meaning replicas are torn down when idle and every new request incurs a full cold start, which is the opposite of the goal.

Option C is wrong because increasing maxReplicaCount only raises the ceiling for horizontal scaling under load and does nothing to reduce the latency of starting an individual replica. Option D is wrong because batch prediction is an asynchronous, job-based mode that does not serve real-time online requests and therefore is irrelevant to minimizing endpoint cold-start latency.

Exam trap

Google often tests the misconception that scale-to-zero (minReplicaCount=0) reduces latency, when in fact it increases cold-start latency; the correct approach is to keep at least one replica always warm (minReplicaCount=1) and pre-load the model during container startup.

9
MCQeasy

You have trained a scikit-learn model and want to deploy it to Vertex AI for online predictions. You need to minimize the effort to create a custom container and ensure the model is served with the default pre-built container. What should you do?

A.Package the model into a custom container with a Flask app and deploy it to Vertex AI.
B.Export the model as a PMML file and deploy using the pre-built XGBoost container.
C.Convert the model to TensorFlow SavedModel format and deploy using the pre-built TensorFlow container.
D.Save the model using joblib and upload it to Vertex AI Model Registry, then deploy using the pre-built scikit-learn container.
AnswerD

Vertex AI provides pre-built containers for scikit-learn that expect the model to be saved in a specific format, typically joblib or pickle. By saving the model with joblib and uploading it to Model Registry, you can deploy it directly using the pre-built container without building a custom image. This minimizes effort and leverages the managed serving stack.

Why this answer

The least-effort path is to use the pre-built scikit-learn container on Vertex AI. Saving the model with joblib and uploading it to Model Registry allows you to deploy without writing any serving code or building a container. The pre-built container handles loading the model and serving predictions.

Other options involve unnecessary conversion or custom packaging, which increase effort and risk.

Exam trap

The trap here is overcomplicating the deployment by assuming a custom container is needed, when Vertex AI provides a pre-built container for scikit-learn models.

10
MCQmedium

A retail company has deployed a scikit-learn model to a Vertex AI endpoint. The model's predictions are used to personalize the homepage. During a flash sale, the endpoint experiences a sudden 10x traffic spike, and the autoscaling configuration is set to minReplicaCount=1, maxReplicaCount=3. The endpoint becomes unresponsive. You need to modify the deployment to handle similar spikes while keeping costs low during normal hours. What should you do?

A.Increase maxReplicaCount to 10 and configure the autoscaling metric to CPU utilization with a target of 60%.
B.Deploy the model to a new endpoint with minReplicaCount=1 and maxReplicaCount=10, and use a traffic split to gradually shift traffic from the old endpoint.
C.Set minReplicaCount to 3 and maxReplicaCount to 10, and enable autoscaling based on CPU utilization with a target of 80%.
D.Keep minReplicaCount=1 but set maxReplicaCount to 10, and configure autoscaling based on a custom metric that tracks the number of incoming requests per second.
AnswerD

This approach allows the endpoint to scale out rapidly during traffic spikes by using a request-based metric that directly reflects load, while keeping the minimum replica count at 1 to save costs during idle periods. A custom metric such as requests per second is more responsive for sudden spikes than CPU utilization, which may lag. Increasing max replicas provides headroom.

Why this answer

The best solution is to increase the maximum replica count to handle spikes and use a custom metric that directly measures request load, such as requests per second, for autoscaling. This allows rapid scaling during flash sales while keeping the minimum replicas low to control costs. CPU-based autoscaling may not react quickly enough for sudden spikes, and raising the minimum replicas increases baseline cost unnecessarily.

Exam trap

The trap here is assuming that CPU utilization is always the best autoscaling metric, when in fact request-based metrics can be more responsive for sudden traffic spikes.

11
Multi-Selectmedium

You are deploying a model to a Vertex AI endpoint for online predictions. You need to ensure that the endpoint can handle traffic spikes and that predictions are served with low latency. Which TWO of the following configurations should you apply? (Choose two.)

Select 2 answers
A.Deploy the model to multiple endpoints and use a load balancer to distribute traffic.
B.Configure autoscaling based on CPU utilization with a target utilization that triggers scale-out early.
C.Use a larger machine type with more CPU and memory for each replica to increase per-replica throughput.
D.Enable request-response logging to capture detailed latency metrics for each prediction.
E.Set a minimum replica count greater than 1 to ensure that the endpoint has warm replicas ready to serve traffic.
AnswersB, E

Autoscaling based on CPU utilization with a low target utilization triggers scale-out earlier, adding replicas before the existing ones become saturated. This proactive scaling helps handle traffic spikes by provisioning additional capacity in advance. It balances cost and performance by scaling out only when needed but doing so early enough to maintain low latency. This is a standard practice for latency-sensitive endpoints.

Why this answer

To handle traffic spikes and ensure low latency, you should maintain a minimum replica count greater than 1 to have warm replicas ready, and configure autoscaling with a low target utilization to trigger scale-out early. These two settings together provide both baseline capacity and proactive scaling, reducing the risk of cold starts and overload during spikes. Other options like logging or larger machines do not directly address dynamic traffic handling.

Exam trap

The trap here is assuming that logging or larger machines are sufficient for handling spikes, when the key is to have warm replicas and proactive autoscaling.

12
MCQhard

You have a Vertex AI endpoint with two deployed models: a champion (v1) and a challenger (v2). You set the traffic split to 90% v1 and 10% v2. After a week, you observe that v2 has better business metrics. You want to shift all traffic to v2 gradually over 3 days to avoid any risk. What should you do?

A.Deploy v2 to a new endpoint and update your clients to use the new endpoint.
B.Use Vertex AI Experiments to compare v1 and v2, then redeploy v2 with 100% traffic.
C.Update the traffic split configuration on the endpoint multiple times over the 3 days to gradually increase v2's percentage.
D.Delete v1 from the endpoint so that all traffic automatically goes to v2.
AnswerC

Updating the endpoint's traffic split repeatedly lets you raise v2's percentage incrementally, satisfying the gradual three-day shift while keeping v1 serving the remainder. Vertex AI supports modifying the deployed model traffic split on a live endpoint without redeployment, so risk is bounded at each step.

Why this answer

Vertex AI endpoints support live traffic splitting between deployed models, allowing you to gradually shift traffic from v1 to v2 by updating the traffic split configuration multiple times over the 3-day period. This approach minimizes risk by enabling incremental rollouts and immediate rollback if issues arise, without requiring client-side changes or downtime.

Exam trap

The trap here is that candidates may assume deleting the old model or redeploying with 100% traffic is acceptable, but the question explicitly requires a gradual shift over 3 days to avoid risk, which only incremental traffic split updates can achieve.

How to eliminate wrong answers

Option A is wrong because deploying v2 to a new endpoint and updating clients introduces unnecessary complexity, potential downtime, and defeats the purpose of gradual traffic shifting; it also requires client-side changes, which is riskier and not aligned with the goal of avoiding risk. Option B is wrong because Vertex AI Experiments are used for offline model evaluation and comparison, not for live traffic management; redeploying v2 with 100% traffic would be an abrupt switch, not a gradual shift over 3 days. Option D is wrong because deleting v1 from the endpoint would immediately route 100% of traffic to v2, which is an abrupt change, not gradual, and violates the requirement to shift traffic gradually over 3 days to avoid risk.

13
MCQmedium

You need to run batch predictions on 10 TB of text data stored in BigQuery using a custom container model hosted in Vertex AI. What is the most cost-effective and simple approach?

A.Use Vertex AI batch prediction with BigQuery source and sink.
B.Use Cloud Run jobs to read from BigQuery and write results back.
C.Export BigQuery data to GCS, then run a Dataflow pipeline to call the model's online prediction endpoint for each row.
D.Use Cloud Dataproc to spin up a Spark cluster and run the model inference in parallel.
AnswerA

Vertex AI batch prediction natively reads from and writes to BigQuery, avoiding data export and re-import. This directly satisfies the cost-effective and simple constraint for 10 TB of text, since no intermediate Cloud Storage staging or custom extraction code is needed.

Why this answer

Vertex AI batch prediction natively supports BigQuery as both source and sink, allowing you to run predictions on 10 TB of text data without any data movement or intermediate storage. This is the most cost-effective and simple approach because it eliminates the need for exporting data, managing infrastructure, or calling online endpoints, and it leverages Vertex AI's optimized batch inference infrastructure that scales automatically.

Exam trap

The exam often tests the misconception that you must export data from BigQuery to GCS before running batch predictions, when in fact Vertex AI batch prediction can directly read from and write to BigQuery, making the export step unnecessary and cost-inefficient.

How to eliminate wrong answers

Option B is wrong because Cloud Run jobs have a maximum request timeout of 60 minutes and are not designed for processing 10 TB of data efficiently; they would require complex batching and retry logic, and would incur higher costs due to per-request pricing and lack of native BigQuery integration. Option C is wrong because exporting data to GCS and then using Dataflow to call the online prediction endpoint for each row introduces unnecessary data movement, storage costs, and network latency; online endpoints are designed for low-latency single requests, not high-throughput batch processing, and this approach would be both slower and more expensive. Option D is wrong because Cloud Dataproc requires you to manage a Spark cluster, handle autoscaling, and write custom inference code, which adds operational complexity and cost for a task that Vertex AI batch prediction can handle natively with no infrastructure management.

14
MCQmedium

You need to perform batch predictions on 10 TB of data stored in BigQuery using Vertex AI. The model requires some preprocessing that cannot be expressed in SQL. What is the most scalable approach?

A.Use a Cloud Function to preprocess each row and write to a new BigQuery table, then run batch prediction.
B.Use Dataflow to read from BigQuery, perform preprocessing, write results to GCS, then run Vertex AI batch prediction job with GCS source.
C.Use Vertex AI batch prediction with BigQuery source and include preprocessing logic in the model container.
D.Export BigQuery data to CSV, run a local Python script for preprocessing, then upload to GCS and start a batch prediction job.
AnswerB

Dataflow handles the arbitrary preprocessing that SQL cannot express, reading from BigQuery and writing prepared data to Cloud Storage. Vertex AI batch prediction then consumes that GCS source at scale, decoupling preprocessing from the prediction job.

Why this answer

Dataflow (Apache Beam) provides a fully managed, auto-scaling, serverless execution environment that can read from BigQuery, apply arbitrary Python/Java preprocessing logic (e.g., feature engineering, normalization) that cannot be expressed in SQL, and write the preprocessed results to Cloud Storage (GCS). Vertex AI batch prediction can then read from GCS as input, making this the most scalable approach for 10 TB of data without requiring custom model container changes or manual data movement.

Exam trap

A common misconception is that Cloud Functions can handle large-scale batch processing, but the trap here is that Cloud Functions are designed for event-driven, short-lived tasks, not for processing terabytes of data in a batch pipeline. Dataflow is the appropriate Google Cloud service for this scenario.

How to eliminate wrong answers

Option A is wrong because Cloud Functions have a 9-minute timeout and limited memory (up to 8 GB), making them unsuitable for processing 10 TB of data row-by-row; they would require an impractical number of invocations and lack built-in parallelization for large-scale batch workloads. Option C is wrong because Vertex AI batch prediction with a BigQuery source does not support preprocessing logic inside the model container — the container receives raw data and must handle all transformations itself, which couples preprocessing to the model and violates separation of concerns; also, BigQuery source does not allow custom preprocessing steps before inference. Option D is wrong because exporting 10 TB of data to CSV, running a local Python script (single machine, no distributed processing), then uploading to GCS is not scalable — it creates a bottleneck at the local script, requires significant network transfer, and does not leverage managed services for parallel processing.

15
MCQhard

You are using Vertex AI Prediction with a custom container that requires a large model file (5 GB). Deployment takes 10 minutes to start. You want to reduce cold start latency. Which action would be MOST effective?

A.Compress the model file and decompress on startup.
B.Use a machine type with local SSD to speed up model loading.
C.Switch to batch prediction to avoid online cold start.
D.Set minReplicas to 1 to keep at least one instance always running.
AnswerD

Keeping minReplicas at 1 maintains a warm instance, so the 5 GB model file is already loaded and the container initialised. This directly eliminates the 10-minute cold start, since new requests hit a running replica rather than triggering a fresh container pull and model load.

Why this answer

Setting minReplicas to 1 keeps at least one prediction instance warm, eliminating the cold start caused by loading the 5 GB model from scratch on each new deployment. Vertex AI Prediction scales replicas based on traffic; with minReplicas=1, the model stays loaded in memory and responds immediately. Compression, local SSD, and batch prediction do not address the fundamental issue of keeping a warm online endpoint.

Exam trap

PMLE often tests cold-start mitigation — candidates confuse storage-level optimizations (SSD, compression) with the actual fix, which is keeping a warm replica via minReplicas.

How to eliminate wrong answers

Option A is wrong because compressing and decompressing on startup still requires loading the full model into memory on every cold start, adding CPU overhead and not eliminating the latency. Option B is wrong because local SSD may speed disk reads but the dominant cold-start cost is model loading and container initialization, which minReplicas avoids entirely. Option C is wrong because batch prediction is asynchronous and not suitable for online low-latency serving; it sidesteps rather than solves the cold-start problem.

16
MCQmedium

A team wants to run a Vertex AI pipeline that deploys a model, runs a smoke test against the endpoint, and automatically rolls back if the smoke test fails. They need the deployment step to be reversible and the endpoint to remain available during the update. Which approach should they use?

A.Deploy the new model to a separate endpoint, run the smoke test there, and then update the application's endpoint URL after the test passes.
B.Use a Vertex AI pipeline with a custom training component that trains and deploys the model in a single step, and rely on pipeline retry policies for rollback.
C.Deploy the new model as a second deployment resource on the endpoint with 0% traffic, run the smoke test, then shift traffic and delete the old deployment if the test passes.
D.Undeploy the current model, deploy the new model on the same endpoint, and run the smoke test; if it fails, redeploy the previous model.
AnswerC

Creating a new deployment resource on the same endpoint keeps the existing model serving while the new one is validated. Traffic can be shifted only after the smoke test passes, and the old resource can be removed afterward. This pattern preserves availability and makes rollback trivial because the old deployment is untouched until the new one is proven.

Why this answer

Adding a new deployment resource to the existing endpoint lets the team validate the new model without disrupting current traffic. Traffic can be moved only after the smoke test succeeds, and removing the old resource completes the rollout; if the test fails, the old resource is still serving and no rollback action is needed.

Exam trap

The trap here is treating redeployment after failure as equivalent to a rollback, when a true rollback leaves the prior deployment serving continuously with no gap.

17
MCQeasy

You have a Vertex AI endpoint that serves a model for real-time predictions. You want to update the model to a new version with zero downtime. Which approach should you take?

A.Delete the endpoint and recreate it with the new model.
B.Deploy the new model version to the same endpoint and then set traffic to 100% for the new version.
C.Use Cloud Load Balancing to switch traffic between two endpoints.
D.Create a new endpoint and update the client application to point to the new endpoint.
AnswerB

Deploying the new version alongside the existing one on the same endpoint, then shifting traffic to 100%, uses Vertex AI's traffic splitting to avoid downtime. Requests route to the old version until the switch, satisfying the zero-downtime requirement.

Why this answer

Vertex AI endpoints support canary deployments by allowing you to deploy a new model version to the same endpoint and then gradually shift traffic to it using the `traffic_split` parameter. Setting traffic to 100% for the new version after deployment ensures zero downtime, as the endpoint remains active and serves requests from the old version until the switch is complete.

Exam trap

The trap here is that candidates assume a new endpoint or load balancer is required for zero-downtime updates, but Vertex AI endpoints natively support traffic splitting between model versions on the same endpoint, making external components unnecessary.

How to eliminate wrong answers

Option A is wrong because deleting and recreating the endpoint causes downtime during the deletion and creation process, and the endpoint URL changes, requiring client updates. Option C is wrong because Cloud Load Balancing is an external traffic management layer that adds unnecessary complexity and latency; Vertex AI endpoints natively support traffic splitting without needing an external load balancer. Option D is wrong because creating a new endpoint changes the endpoint URL, which requires updating client applications, leading to potential downtime or misrouting during the transition.

18
MCQhard

A financial services company deploys a fraud detection model on a Vertex AI Endpoint. The model must process each transaction in under 50 ms. The team notices that p99 latency spikes to 200 ms every few minutes. Logs show that the model container performs a cold start when new replicas are added, and the autoscaler frequently adds and removes replicas. The endpoint currently has minReplicaCount=1 and maxReplicaCount=10. What should they do to reduce the latency spikes while controlling cost?

A.Deploy the model to a Vertex AI Batch Prediction job and use online predictions only for high-value transactions.
B.Set minReplicaCount to a value that covers baseline traffic and increase the autoscaling cool-down period to avoid rapid scale-down.
C.Configure the endpoint to use a smaller machine type so that replicas start faster and cold starts are shorter.
D.Increase maxReplicaCount to 20 and lower the autoscaling target CPU utilization to 50%.
AnswerB

Keeping a baseline number of warm replicas prevents cold starts during normal traffic fluctuations. Extending the cool-down period reduces thrashing, so replicas are not removed immediately after a spike. This combination maintains low p99 latency while avoiding unnecessary replica churn, which is the main cause of the 200 ms spikes.

Why this answer

Cold starts and rapid scale-down cause p99 latency spikes. Maintaining a baseline of warm replicas ensures that new requests do not hit a cold container, and extending the cool-down period prevents replicas from being removed too quickly. This stabilizes the replica count and keeps latency low without over-provisioning for peak traffic.

Exam trap

The trap here is focusing on increasing maximum capacity or lowering the scaling target, when the real issue is cold starts and replica thrashing.

19
MCQhard

You are deploying a PyTorch model on Vertex AI using a custom container with NVIDIA Triton Inference Server. The model is a large transformer that requires GPU. You want to optimize GPU utilization and reduce memory footprint. Which technique should you apply?

A.Enable dynamic batching in Triton.
B.Use CPU-only instances to avoid GPU memory issues.
C.Increase the number of GPU replicas.
D.Apply model quantization using TensorRT.
AnswerD

TensorRT quantisation converts FP32 weights to lower precision such as FP16 or INT8, shrinking memory footprint and boosting throughput on NVIDIA GPUs. Triton serves the optimised engine, so GPU utilisation improves while latency drops, satisfying the memory and utilisation constraints.

Why this answer

Model quantization using TensorRT reduces the precision of model weights (e.g., from FP32 to FP16 or INT8), which directly decreases GPU memory usage and can improve throughput by enabling faster arithmetic operations on compatible NVIDIA GPUs. This technique is specifically designed to optimize GPU utilization and memory footprint for large transformer models deployed with Triton Inference Server.

Exam trap

Google often tests the distinction between throughput optimization techniques (like dynamic batching) and memory footprint reduction techniques (like quantization), leading candidates to mistakenly choose dynamic batching when the question specifically asks about reducing memory footprint.

How to eliminate wrong answers

Option A is wrong because dynamic batching improves throughput by grouping inference requests, but it does not reduce the memory footprint per model instance or optimize GPU utilization in terms of memory efficiency. Option B is wrong because CPU-only instances cannot run the large transformer model with acceptable latency or throughput, and the question explicitly requires GPU. Option C is wrong because increasing the number of GPU replicas scales horizontally, which increases total memory footprint and cost, rather than reducing memory footprint per replica or optimizing utilization of a single GPU.

20
MCQhard

Your team is deploying a large recommendation model on Vertex AI endpoints using GPUs. You need to minimise latency while optimising cost. The model serves many similar requests from the same users within short time windows. Which additional service would best reduce latency and cost?

A.Switch to CPU-only instances to reduce cost.
B.Increase maxReplicas to handle the load without caching.
C.Set up a Cloud CDN in front of the endpoint.
D.Use Cloud Memorystore to cache prediction results.
AnswerD

Cloud Memorystore caches repeated predictions for the same users within short windows, so identical requests bypass GPU inference entirely. This cuts both latency and GPU cost, directly addressing the repeated-request pattern while the endpoint stays GPU-backed.

Why this answer

Cloud Memorystore (Redis) in front of the Vertex AI endpoint lets you cache prediction results keyed by user/request signature, so repeated similar requests within short windows are served from cache instead of hitting the GPU-backed model. This reduces both latency (cache hit is sub-millisecond) and cost (fewer GPU inference calls).

Exam trap

PMLE often tests whether candidates reach for infrastructure scaling (more replicas, CDN) when the correct answer is application-level caching for repeated identical or near-identical requests.

How to eliminate wrong answers

Option A is wrong because switching to CPU-only instances would increase latency for a large recommendation model and likely reduce throughput, defeating the latency goal. Option B is wrong because increasing maxReplicas scales out GPU instances, which increases cost and does nothing to reduce per-request latency for repeated queries. Option C is wrong because Cloud CDN caches HTTP responses at edge locations but is not designed for dynamic, user-specific prediction payloads and does not integrate with Vertex AI prediction semantics.

21
MCQhard

A company runs a Vertex AI endpoint that serves a model for real-time predictions. The endpoint uses a custom container that loads a 10 GB model into memory. During a traffic spike, the autoscaler adds new replicas, but each new replica takes 8 minutes to become ready because it must download the model from Cloud Storage. The team wants to reduce scale-up time. Which approach is most effective?

A.Increase the minReplicaCount to a value that covers peak traffic so that new replicas are not needed.
B.Bake the model artifacts into the custom container image and push it to Artifact Registry, then deploy the endpoint with that image.
C.Store the model artifacts in a regional Cloud Storage bucket and enable parallel downloads in the custom container.
D.Enable Vertex AI Model Monitoring to detect when the model is slow to load and trigger a preemptive scale-up.
AnswerB

Baking the model into the container image eliminates the runtime download from Cloud Storage. The container registry can serve the image layers efficiently, and the model is available as soon as the container starts. This significantly reduces startup time, often from minutes to seconds, and is the recommended pattern for large models. It also avoids dependency on external storage during scale-up.

Why this answer

The slow scale-up is caused by downloading a 10 GB model from Cloud Storage at container startup. By embedding the model into the container image and storing it in Artifact Registry, the model is available locally when the container starts. This removes the network transfer from the critical path and dramatically reduces the time for a new replica to become ready, allowing the autoscaler to respond faster to traffic spikes.

Exam trap

The trap here is focusing on autoscaler settings or monitoring instead of addressing the root cause: the model is fetched at runtime rather than being part of the image.

22
MCQmedium

A company wants to cache predictions for identical requests to reduce latency and cost. They use Vertex AI Prediction with a custom container. Which GCP service should they use to implement prediction caching?

A.Cloud Bigtable
B.Cloud Memorystore for Redis
C.Cloud Storage
D.Cloud Firestore
AnswerB

Memorystore for Redis provides sub-millisecond key-value lookups, letting the custom container return cached predictions for identical requests before invoking the model. This satisfies the latency and cost reduction goal, since repeated inference is skipped entirely.

Why this answer

Cloud Memorystore for Redis is an in-memory data store with sub-millisecond latency, making it the ideal GCP service for caching prediction results keyed by request hash. It supports TTL-based expiration and high-throughput reads, which directly reduce latency and repeated model inference costs. This is the canonical GCP caching layer for Vertex AI prediction workloads.

Exam trap

PMLE often tests the confusion between caching (in-memory, sub-millisecond, Redis/Memorystore) and persistent storage (Bigtable, Firestore, Cloud Storage) — candidates must match the latency and access-pattern requirements to the correct service class.

How to eliminate wrong answers

Option A (Cloud Bigtable) is wrong because it is a wide-column NoSQL store optimized for large-scale analytical and time-series workloads with millisecond latency — not sub-millisecond in-memory caching. Option C (Cloud Storage) is wrong because it is object storage with high latency (tens to hundreds of milliseconds) and is unsuitable for low-latency cache lookups. Option D (Cloud Firestore) is wrong because it is a document database with single-digit millisecond latency at best and is designed for mobile/web app data, not high-throughput prediction caching.

23
MCQmedium

A retail company serves a product-ranking model on a Vertex AI endpoint. Traffic is highly predictable: a steady baseline all day with a sharp peak every evening. During the evening peak, prediction latency exceeds the SLO for several minutes before autoscaling stabilises. The team wants to reduce this scale-up lag without over-provisioning hardware for the entire day. Which configuration should they apply to the deployed model?

A.Switch the endpoint to a private endpoint and increase the machine type of each replica.
B.Enable request-response logging and reduce the model's input feature count.
C.Configure a dedicated autoscaling metric with a lower utilization target and define a scale-up schedule aligned to the evening peak.
D.Set a higher maxReplicaCount and rely on the default autoscaling metrics.
AnswerC

Vertex AI Model Deployment autoscaling supports both a target utilization metric and scheduling options. Lowering the target utilization makes the autoscaler add replicas earlier, while a schedule pre-warms capacity exactly when the predictable evening surge begins. Together they cut the reactive ramp-up delay without paying for peak capacity around the clock, which is precisely the SLO gap described.

Why this answer

The latency breach is a capacity-timing problem, not a capacity-ceiling problem, so the fix must make replicas available before demand arrives. Vertex AI Model Deployment autoscaling lets you choose the metric the autoscaler tracks and set a target utilization, and it also supports scheduled scaling. Lowering the target makes scale-out begin earlier, and a schedule aligned with the predictable evening peak pre-warms replicas, eliminating the reactive ramp-up lag while keeping off-peak cost low.

Exam trap

The trap here is assuming that a larger maximum replica count automatically reduces scale-up latency, when the delay actually comes from reactive scaling that only starts after load is already observed.

24
Multi-Selectmedium

A company uses Vertex AI Matching Engine for real-time recommendations. They need to serve queries with low latency and support frequent updates. Which two configurations are appropriate? (Choose 2)

Select 2 answers
A.Store the index in Cloud Storage and query via Python
B.Enable streaming updates for the index
C.Use a brute-force index for exact results
D.Deploy the index to a Vertex AI Matching Engine endpoint
E.Use batch updates only
AnswersB, D

Streaming updates allow new embeddings to be inserted or removed incrementally while the index remains queryable, avoiding the full rebuild and redeploy that batch updates force. This satisfies the frequent-update requirement without interrupting the low-latency online serving the stem specifies.

Why this answer

Option B is correct because Vertex AI Matching Engine supports streaming updates, which allow the index to be modified incrementally as new data arrives without requiring a full rebuild, directly satisfying the requirement for frequent updates in a real-time recommendation system. Option D is correct because deploying the index to a Vertex AI Matching Engine endpoint is the standard mechanism for serving low-latency, real-time nearest-neighbor queries at scale, which is exactly what the scenario demands. Option A is not appropriate because storing the index in Cloud Storage and querying it via Python does not provide the managed, low-latency serving infrastructure of Matching Engine.

Option C is not appropriate because a brute-force index performs exhaustive comparisons, which is far too slow for low-latency real-time serving at scale. Option E is not appropriate because batch-only updates cannot keep pace with frequent updates and would introduce staleness in a real-time recommendation system.

Exam trap

The trap here is that in Google's Vertex AI Matching Engine, candidates often confuse batch updates with streaming updates, assuming that batch updates can be made frequent enough to approximate real-time, but they fail to recognize that batch updates require full index rebuilds, which introduce significant latency and downtime for serving.

25
MCQhard

A data scientist wants to perform A/B testing between two model versions deployed on the same Vertex AI endpoint. They need to route 10% of traffic to the challenger model. Which approach should they use?

A.Use Vertex AI Experiments to compare models offline, then deploy the winner
B.Deploy the challenger model to a separate endpoint and use a load balancer to split traffic
C.Update the champion model with a new version and use model version aliases
D.Deploy both models to the same endpoint and set traffic_split to 90 for champion and 10 for challenger
AnswerD

A single Vertex AI endpoint can host multiple model versions as DeployModel resources, and the endpoint's traffic_split dictionary assigns integer percentage weights per deployed model ID. Setting 90/10 routes exactly one tenth of prediction requests to the challenger, giving the required A/B comparison without a second endpoint.

Why this answer

Vertex AI endpoints support traffic splitting directly, allowing you to route a percentage of requests to different model versions deployed on the same endpoint. By setting `traffic_split` to 90 for the champion and 10 for the challenger, the data scientist can perform online A/B testing without additional infrastructure. This is the simplest and most cost-effective approach, as it avoids managing separate endpoints or load balancers.

Exam trap

A common trap in Google exams is the misconception that separate endpoints or load balancers are required for A/B testing, when in fact Vertex AI endpoints provide built-in traffic splitting for this exact purpose.

How to eliminate wrong answers

Option A is wrong because Vertex AI Experiments is an offline evaluation tool for comparing model performance on historical data, not for live traffic splitting. Option B is wrong because deploying the challenger to a separate endpoint and using a load balancer adds unnecessary complexity and cost; Vertex AI endpoints natively support traffic splitting across model versions. Option C is wrong because updating the champion model with a new version and using model version aliases does not provide granular traffic splitting; aliases are for version management, not for routing a specific percentage of live traffic.

26
MCQeasy

Which Vertex AI service is designed for building and managing approximate nearest neighbor (ANN) indexes for similarity search at scale?

A.Vertex AI AutoML
B.Vertex AI Workbench
C.Vertex AI Prediction
D.Vertex AI Matching Engine (Vector Search)
AnswerD

Matching Engine (Vector Search) is the Vertex AI service purpose-built for creating and serving ANN indexes, delivering scalable similarity search across billions of embeddings. It directly satisfies the stem's requirement for approximate nearest neighbour indexing at scale, unlike general prediction or training services.

Why this answer

Vertex AI Matching Engine (Vector Search) is specifically designed for building and managing approximate nearest neighbor (ANN) indexes for similarity search at scale. It allows you to create indexes from embeddings and perform low-latency similarity queries. This service is optimized for large-scale vector search use cases.

Exam trap

PMLE often tests the distinction between Vertex AI services, and candidates may confuse Matching Engine with Prediction or AutoML, not realizing it's specifically for vector similarity search.

How to eliminate wrong answers

Option A is wrong because Vertex AI AutoML is for building custom machine learning models without coding, not for ANN indexes. Option B is wrong because Vertex AI Workbench is a notebook environment for data science, not a similarity search service. Option C is wrong because Vertex AI Prediction is for deploying models for online or batch predictions, not for managing ANN indexes.

27
Multi-Selectmedium

A company needs to reduce inference latency for their online prediction service on Vertex AI. Which two actions would help? (Choose 2)

Select 2 answers
A.Increase the maximum number of replicas
B.Deploy the model on a GPU-enabled machine
C.Enable model quantization via Vertex AI Model Optimization
D.Use a smaller machine type with less memory
E.Enable autoscaling with a lower target CPU utilization
AnswersB, C

GPUs accelerate the matrix operations underpinning neural network inference, cutting per-request compute time. This addresses the stem's latency constraint directly, since the bottleneck in online prediction is typically model computation rather than network transfer or storage access.

Why this answer

Option B is correct because deploying the model on a GPU-enabled machine type gives the inference workload hardware acceleration, which speeds up the matrix/tensor computations typical of ML models and thereby reduces per-request prediction latency. Option C is correct because enabling model quantization via Vertex AI Model Optimization reduces the numerical precision of weights and activations (for example to INT8), shrinking model size and memory bandwidth requirements and enabling faster inference, which lowers latency. Option A is not correct because increasing the maximum number of replicas improves throughput and availability under load, but it does not reduce the latency of an individual prediction.

Option D is not correct because using a smaller machine type with less memory would likely increase latency or cause out-of-memory failures rather than reduce it. Option E is not correct because autoscaling with a lower target CPU utilization only triggers scaling sooner to add capacity; it addresses throughput/load handling, not the intrinsic latency of a single inference.

Exam trap

Google often tests the distinction between scaling for throughput (replicas, autoscaling) versus reducing per-request latency (hardware acceleration, model optimization), leading candidates to confuse horizontal scaling with performance optimization.

28
Multi-Selecthard

A company is migrating from an on-premises ML serving infrastructure to Vertex AI. They have multiple models that need to be served from the same endpoint with different traffic percentages. They also need to monitor prediction quality. Which THREE actions should they take? (Choose 3)

Select 3 answers
A.Deploy multiple model versions on the same endpoint with traffic_split parameter.
B.Deploy each model as a separate endpoint and use Cloud Load Balancing.
C.Enable Vertex AI Model Monitoring to detect prediction drift.
D.Use Cloud Monitoring to create custom metrics based on business outcomes.
E.Export logs to BigQuery for manual analysis only.
AnswersA, C, D

Deploying multiple models to one endpoint with the `traffic_split` parameter directly satisfies the requirement for serving several models from a single endpoint at differing traffic percentages. Vertex AI assigns each deployed model a percentage of incoming prediction requests, enabling gradual rollouts or A/B testing without provisioning separate endpoints.

Why this answer

Option A is correct because Vertex AI endpoints natively support serving multiple DeployedModels behind a single endpoint, and the traffic_split parameter on the endpoint lets you route configurable percentages of prediction traffic to each model version. Option C is correct because Vertex AI Model Monitoring is the managed service that detects training-serving skew and prediction drift on deployed models, directly addressing the requirement to monitor prediction quality. Option D is correct because prediction quality often depends on business outcomes that Model Monitoring does not capture, so Cloud Monitoring custom metrics let you track those outcome-based signals alongside Vertex AI metrics.

Option B is not appropriate because separate endpoints with Cloud Load Balancing adds unnecessary complexity and does not provide Vertex AI's built-in traffic splitting or model monitoring integration. Option E is insufficient because exporting logs to BigQuery for manual analysis only is not an automated monitoring solution and does not satisfy the prediction-quality monitoring requirement.

Exam trap

The trap here is that candidates may think separate endpoints with a load balancer (Option B) are required for traffic distribution, overlooking Google Vertex AI's built-in traffic splitting on a single endpoint, which is simpler and more aligned with the platform's design.

29
MCQmedium

A financial services company uses a Vertex AI Endpoint to serve a credit risk model. The model must always be available, even during maintenance windows, and they need to control the exact distribution of traffic across two model versions for a gradual rollout. They also want to minimize cold-start latency. Which deployment configuration should they use?

A.Deploy a single model version with minReplicaCount=0 and maxReplicaCount=10, and use a custom prediction routine to handle both versions based on request headers.
B.Create two separate endpoints, one for each model version, and use a global load balancer to distribute traffic equally between them.
C.Deploy the model to a Vertex AI Batch Prediction job and schedule it to run every hour, then use a Cloud Function to serve predictions from the latest batch results.
D.Deploy both model versions to the same endpoint, set minReplicaCount to at least 1 for each, and use traffic splitting to route a percentage of requests to each version.
AnswerD

This configuration ensures high availability with dedicated replicas for each model version, allows precise traffic splitting for gradual rollout, and minimizes cold starts by keeping at least one replica warm. It directly addresses all requirements: availability, traffic control, and latency.

Why this answer

Deploying both versions to the same endpoint with dedicated replicas and traffic splitting satisfies availability, precise rollout control, and low latency. It leverages Vertex AI's native capabilities for version management and gradual traffic shifting without additional infrastructure.

Exam trap

The trap here is assuming that separate endpoints or batch processing can achieve the same level of control and low latency as a single endpoint with traffic splitting and minimum replicas.

30
MCQmedium

An engineer deploys a model to a Vertex AI endpoint with minReplicas=1 and maxReplicas=3. The endpoint receives a sudden traffic spike, but it does not scale up beyond 1 replica. The CPU utilization target is 60%. What is the most likely cause?

A.The model is not deployed correctly.
B.The endpoint is configured with the wrong machine type.
C.The CPU utilization is below the target threshold, so the autoscaler does not add replicas.
D.The endpoint is using GPU which cannot autoscale.
AnswerC

Vertex AI autoscaling adds replicas only when average CPU utilisation exceeds the 60% target; if observed utilisation stays below that threshold, the scaler holds at minReplicas=1 regardless of request volume. The spike therefore did not drive per-replica CPU high enough to trigger scale-out, making this the likely cause.

Why this answer

Vertex AI's autoscaler uses CPU utilization as a metric to decide when to add replicas. If the CPU utilization remains below the 60% target threshold, the autoscaler will not trigger scale-up, even during a traffic spike. The endpoint is configured with minReplicas=1 and maxReplicas=3, but without exceeding the target, it stays at the minimum.

Exam trap

The trap here is that candidates assume any traffic spike automatically triggers scaling, but Vertex AI's autoscaler only scales based on the configured metric (CPU utilization), not request volume directly.

How to eliminate wrong answers

Option A is wrong because the model being deployed correctly is unrelated to autoscaling behavior; a misdeployment would typically cause prediction failures or errors, not a failure to scale. Option B is wrong because the machine type affects performance and cost, but does not directly prevent the autoscaler from adding replicas when CPU utilization exceeds the target. Option D is wrong because GPU-enabled endpoints can autoscale; Vertex AI supports autoscaling for GPU instances, though GPU metrics may require custom configuration.

31
MCQeasy

A machine learning engineer needs to run batch predictions on 50 TB of data stored in BigQuery using a Vertex AI model. The model is a custom container. What is the most efficient way to set up the batch prediction job?

A.Create a Vertex AI batch prediction job with BigQuery source and BigQuery destination.
B.Use Dataflow to process the data and call the model via Vertex AI online prediction.
C.Export BigQuery data to CSV in GCS, then create a batch prediction job with GCS source.
D.Create a Cloud Function to iterate over BigQuery rows and call the endpoint.
AnswerA

Native BigQuery source and destination lets Vertex AI read and write directly without exporting 50 TB to Cloud Storage, avoiding costly intermediate copies. This satisfies the efficiency constraint by keeping data in place and streaming results back to BigQuery.

Why this answer

Vertex AI batch prediction natively supports BigQuery as both input source and output destination, allowing the service to read the 50 TB directly from BigQuery and write predictions back without exporting data. This avoids data movement, leverages BigQuery's scalability, and is the most efficient, fully managed approach for large-scale batch inference with a custom container.

Exam trap

PMLE often tests whether candidates default to exporting data to GCS out of habit — the trap is missing that Vertex AI batch prediction supports BigQuery natively, making export steps unnecessary and inefficient.

How to eliminate wrong answers

Option B is wrong because using Dataflow to call online prediction introduces per-row network calls, latency, and cost overhead, and online prediction endpoints are not designed for 50 TB batch workloads. Option C is wrong because exporting 50 TB to CSV in GCS adds significant time, storage cost, and an unnecessary data movement step when BigQuery source is natively supported. Option D is wrong because a Cloud Function iterating over BigQuery rows is not scalable, has execution time and memory limits, and would be prohibitively slow and expensive for 50 TB.

32
MCQhard

Your team is serving a large language model on a Vertex AI endpoint using a custom container. You need to reduce inference latency for long prompts while keeping the deployment cost reasonable. The model uses an autoregressive decoder. Which optimization should you implement?

A.Use a smaller quantized version of the model and serve it with a batch size of one.
B.Implement continuous batching and paged attention in the serving container.
C.Increase maxReplicaCount and rely on autoscaling to handle long prompts.
D.Enable response streaming and return tokens as they are generated.
AnswerB

Continuous batching allows the server to add new requests to an in-flight batch at each decoding step, improving GPU utilization and throughput. Paged attention manages the key-value cache in non-contiguous blocks, reducing memory fragmentation and allowing more concurrent sequences. Together they lower per-token latency and increase the number of requests served per GPU, which directly addresses long-prompt inference cost and latency.

Why this answer

For autoregressive LLM serving, the key bottlenecks are GPU underutilization during decoding and memory fragmentation in the key-value cache. Continuous batching keeps the GPU busy by admitting new requests into the current batch, while paged attention stores the KV cache in fixed-size blocks to avoid fragmentation and support more concurrent sequences. These techniques reduce per-token latency and improve throughput, making them the right optimization for long prompts on a cost-conscious deployment.

Exam trap

The trap here is confusing perceived latency improvements from streaming with actual compute reductions, when the real gains come from batching and KV-cache memory management.

33
Multi-Selecthard

You are designing a batch prediction pipeline using Vertex AI. The input data is 100 TB of images stored in Cloud Storage. The model is a custom TensorFlow model that expects TFRecord format. The pipeline must be cost-effective and run within a time window of 2 hours. Which THREE steps should you include?

Select 3 answers
A.Store batch prediction results in BigQuery.
B.Create a Vertex AI batch prediction job with input from GCS (TFRecord files).
C.Use Dataflow to read images and write TFRecord files to GCS.
D.Store batch prediction results in GCS.
E.Use Cloud Functions to convert images to TFRecord.
AnswersB, C, D

A batch prediction job reading TFRecord files directly from Cloud Storage matches the model's expected input format, avoiding costly conversion of 100 TB. Vertex AI distributes the job across managed workers, meeting the two-hour window cost-effectively.

Why this answer

Option B is correct because Vertex AI batch prediction natively accepts TFRecord input files from Cloud Storage, which matches the custom TensorFlow model's expected format and lets the managed service handle the large-scale inference job within the 2-hour window. Option C is correct because Dataflow is a scalable, cost-effective managed service for reading 100 TB of images from GCS and transforming them into TFRecord files, a task Cloud Functions cannot handle at this volume. Option D is correct because Vertex AI batch prediction writes its output to Cloud Storage, which is the appropriate and cost-effective destination for large result sets from 100 TB of input.

Option A is not appropriate because BigQuery is not the default or cost-effective sink for massive batch prediction outputs and would add unnecessary loading overhead. Option E is not appropriate because Cloud Functions has execution time, memory, and concurrency limits that make it unsuitable for converting 100 TB of images to TFRecord.

Exam trap

Google often tests the misconception that Cloud Functions can handle large-scale data processing tasks, but the trap here is that Cloud Functions have strict timeout and memory limits, making them unsuitable for converting 100 TB of images to TFRecord format.

34
MCQmedium

A data science team needs to serve multiple versions of the same ML model on Vertex AI Endpoints for A/B testing. They want to gradually shift traffic from the current 'champion' model to a new 'challenger' model. Which feature should they use?

A.Deploy the challenger to a separate endpoint and use a proxy to split traffic.
B.Use Cloud Load Balancing with weighted backend services.
C.Deploy both models to the same endpoint and use traffic splitting.
D.Use Vertex AI Experiments to manage model versions.
AnswerC

Vertex AI Endpoints support deploying multiple model versions to one endpoint with traffic splitting, letting you route a percentage to the challenger and gradually shift it. This satisfies the A/B testing and champion-to-challenger migration requirement.

Why this answer

Vertex AI Endpoints natively support traffic splitting, allowing you to deploy multiple model versions (e.g., champion and challenger) to the same endpoint and assign a percentage of traffic to each. This enables gradual A/B testing without additional infrastructure, as the endpoint automatically routes requests based on the configured split. Option C is correct because it leverages this built-in feature, which is designed specifically for this use case.

Exam trap

The PMLE exam often tests the misconception that traffic splitting requires external load balancers or proxies, when in fact Vertex AI Endpoints provide this capability natively, and candidates may overlook the built-in feature in favor of more complex architectures.

How to eliminate wrong answers

Option A is wrong because deploying the challenger to a separate endpoint and using a proxy adds unnecessary complexity, latency, and management overhead; Vertex AI Endpoints already provide traffic splitting without external proxies. Option B is wrong because Cloud Load Balancing operates at the network layer (HTTP(S) or TCP/UDP) and is designed for distributing traffic across regional backends, not for splitting traffic between model versions on the same Vertex AI Endpoint; it would require separate endpoints and does not integrate with Vertex AI's model versioning. Option D is wrong because Vertex AI Experiments is a tool for tracking and comparing model training runs and hyperparameters, not for serving or routing live traffic; it has no mechanism to split traffic between deployed models.

35
Multi-Selecthard

A retail company deploys a new recommendation model alongside the current champion on Vertex AI Endpoints. They want to gradually shift traffic to the challenger while monitoring business metrics (conversion rate). Which two steps are required? (Choose 2)

Select 2 answers
A.Use Vertex AI Experiments to track the traffic split percentages.
B.Enable Cloud Memorystore to cache identical requests for both models.
C.Deploy the challenger model to the same endpoint as the champion with a separate deployed model.
D.Configure traffic split in the endpoint's traffic_split field (e.g., champion:90, challenger:10).
E.Use Cloud Monitoring to track custom metrics like conversion rate per model version.
AnswersC, D

Vertex AI Endpoints support multiple deployed models behind one endpoint, each with its own deployed model ID. Deploying the challenger alongside the champion on the same endpoint is the prerequisite that lets traffic be split between them.

Why this answer

Option C is correct because Vertex AI Endpoints support multiple deployed models behind a single endpoint, so the challenger must be deployed as a separate DeployedModel (with its own model ID and deployed_model_id) on the same endpoint as the champion to enable side-by-side serving. Option D is correct because gradual traffic shifting is implemented by setting the endpoint's traffic_split map, which assigns percentage weights keyed by deployed_model_id (e.g., champion deployed model: 90, challenger deployed model: 10), allowing controlled canary rollout. Option A is not required because Vertex AI Experiments tracks training/experiment runs and metrics, not live endpoint traffic splitting.

Option B is not required because Memorystore is a caching layer and does not perform traffic splitting or model comparison. Option E, while useful for observing conversion rate, is not a required step to shift traffic and is not marked correct; the question asks for the steps needed to perform the gradual shift itself.

Exam trap

This question tests the distinction between monitoring (which is optional after deployment) and the actual configuration steps required to shift traffic; candidates mistakenly select Cloud Monitoring (Option E) as a required step, but the question specifically asks for steps to 'gradually shift traffic,' which is accomplished by deploying to the same endpoint and setting the traffic split, not by monitoring after the fact.

36
MCQeasy

What is the primary purpose of Vertex AI Edge Manager?

A.To run batch predictions on edge devices
B.To deploy and manage ML models on edge devices at scale
C.To convert models to TensorFlow Lite automatically
D.To train models on edge devices using federated learning
AnswerB

Vertex AI Edge Manager deploys and manages ML models across fleets of edge devices at scale, handling distribution, versioning and monitoring. It targets edge hardware rather than cloud endpoints, which is precisely the capability the question asks about.

Why this answer

Vertex AI Edge Manager is specifically designed to deploy, monitor, and manage ML models on edge devices at scale. It handles model packaging, over-the-air updates, and health monitoring across fleets of edge devices, which is distinct from simply running batch predictions or converting model formats.

Exam trap

Google Cloud often tests the distinction between 'managing models at scale' (deployment, updates, monitoring) and 'running inference' or 'converting formats' — candidates confuse the operational management role with the execution or preprocessing steps.

How to eliminate wrong answers

Option A is wrong because batch predictions on edge devices are a use case, not the primary purpose; Vertex AI Edge Manager focuses on lifecycle management (deployment, updates, monitoring) rather than just executing predictions. Option C is wrong because model conversion to TensorFlow Lite is handled by tools like the TensorFlow Lite Converter or Vertex AI's model optimization services, not by Edge Manager itself. Option D is wrong because training on edge devices using federated learning is a separate paradigm (e.g., TensorFlow Federated) and is not a core function of Vertex AI Edge Manager, which manages already-trained models.

37
MCQeasy

Which Vertex AI feature allows you to reduce the size of a trained model to improve inference speed on edge devices without significant accuracy loss?

A.Vertex AI Model Optimization
B.Vertex AI Model Monitoring
C.Vertex AI Matching Engine
D.Vertex AI Continuous Training
AnswerA

Vertex AI Model Optimization applies quantisation and pruning to shrink trained models, reducing size and inference latency on constrained edge hardware. It meets the stem's requirement of faster edge inference without significant accuracy loss, unlike retraining or distillation approaches.

Why this answer

Vertex AI Model Optimization is the correct feature because it provides model quantization, pruning, and distillation techniques specifically designed to reduce model size and improve inference latency on edge devices. This service applies post-training quantization (e.g., FP32 to INT8) and structured weight pruning to shrink the model footprint while maintaining accuracy within acceptable thresholds, directly addressing the need for efficient deployment on resource-constrained hardware.

Exam trap

Google PMLE exams often test the distinction between 'optimization' (size/speed improvements) and 'monitoring' (observability), leading candidates to confuse Model Monitoring with performance tuning because both involve 'model performance' terminology.

How to eliminate wrong answers

Option B is wrong because Vertex AI Model Monitoring is used for detecting prediction drift, data skew, and feature attribution changes in deployed models, not for reducing model size or optimizing inference speed. Option C is wrong because Vertex AI Matching Engine is a vector similarity search service for large-scale embedding-based retrieval (e.g., recommendation systems), not a model compression or optimization tool. Option D is wrong because Vertex AI Continuous Training automates retraining pipelines based on new data or schedules, but it does not perform model size reduction or inference optimization.

38
Multi-Selecthard

An organization is deploying a mission-critical model on Vertex AI Endpoints. They need to ensure high availability and meet a strict SLO of 99.9% uptime. Which THREE steps should they take? (Choose 3)

Select 3 answers
A.Use Cloud CDN to cache responses.
B.Set minReplicas to at least 2 to ensure redundancy within a region.
C.Use a single large instance instead of multiple small ones.
D.Deploy the endpoint in multiple regions.
E.Configure health checks to detect and replace unhealthy instances.
AnswersB, D, E

Setting minReplicas to at least two keeps additional replicas running within the region, so a single node failure does not take the endpoint offline. This removes the single point of failure underpinning the 99.9% uptime SLO.

Why this answer

Option B is correct because setting minReplicas to at least 2 on a Vertex AI Endpoint ensures that multiple model server replicas run within the region, so a single replica failure or rolling update does not cause an outage and the 99.9% SLO can be maintained. Option D is correct because deploying the endpoint in multiple regions provides geographic redundancy, so a regional outage of Vertex AI or its underlying infrastructure does not take the mission-critical model offline. Option E is correct because configuring health checks (Vertex AI endpoint health checks/readiness probes) lets the service detect unhealthy replicas and route traffic away from or replace them, preserving availability.

Option A is not appropriate because Cloud CDN caches HTTP responses at the edge and does not address the availability of the Vertex AI prediction backend or its SLO. Option C is not appropriate because using a single large instance creates a single point of failure and provides no redundancy, which works against high availability.

Exam trap

PMLE often tests the misconception that a single large instance provides better availability than multiple small ones — candidates confuse vertical scaling (more power) with horizontal scaling (more redundancy), but only the latter provides fault tolerance.

39
MCQeasy

You have trained a scikit-learn model and saved it as a joblib file in Cloud Storage. You need to deploy this model to Vertex AI for online predictions with minimal effort. What should you do?

A.Upload the joblib file to Vertex AI Model Registry and deploy it directly without specifying a container.
B.Convert the scikit-learn model to TensorFlow SavedModel format and use the pre-built TensorFlow container.
C.Write a custom container that loads the joblib file and serves predictions using Flask, then deploy it to Vertex AI.
D.Use the pre-built scikit-learn container provided by Vertex AI and specify the model artifact path in Cloud Storage.
AnswerD

Vertex AI provides pre-built containers for popular frameworks like scikit-learn. You can deploy the model by specifying the pre-built container image and the path to the joblib file in Cloud Storage. This requires no custom code and is the fastest way to deploy a scikit-learn model for online predictions.

Why this answer

Vertex AI offers pre-built containers for scikit-learn that can load joblib files directly. By using the pre-built container and specifying the model artifact path, you can deploy the model with minimal effort. Writing a custom container or converting the model format adds unnecessary work, and Model Registry does not support containerless deployment.

Exam trap

The trap here is assuming that you need to write a custom container or convert the model format to deploy a scikit-learn model on Vertex AI.

40
MCQmedium

A company is deploying a new model version to an existing Vertex AI endpoint. They want to test the new version with 5% of traffic before fully rolling it out. What is the correct approach?

A.Create a new endpoint for the new version and update the client to call both endpoints.
B.Deploy the new version and set the minimum replicas to 0, then gradually increase.
C.Use Cloud Load Balancing to distribute traffic between two endpoints.
D.Deploy the new version as a separate model on the same endpoint and use the `traffic_split` parameter in the deployment request.
AnswerD

Deploying the new version as a separate model on the same Vertex AI endpoint and setting the `traffic_split` parameter to route 5% of requests to it directly satisfies the constraint of testing with a controlled fraction of live traffic before a full rollout. This mechanism uses the endpoint’s built-in traffic routing to allocate a precise percentage of inference requests to the new model version without requiring a separate endpoint or external load balancer.

Why this answer

Vertex AI endpoints support traffic splitting between multiple deployed models. By deploying the new model version to the same endpoint and setting `traffic_split` to 5% for the new version and 95% for the existing version, the endpoint automatically routes a corresponding proportion of inference requests to each model without any client-side changes.

Exam trap

The trap here is that candidates may confuse traffic splitting with scaling or load balancing, assuming that adjusting replicas or using an external load balancer is required, when Vertex AI's native `traffic_split` is the simplest and correct method for canary deployments.

How to eliminate wrong answers

Option A is wrong because creating a new endpoint and updating clients to call both endpoints introduces unnecessary complexity, latency, and risk of client misconfiguration; Vertex AI endpoints natively support traffic splitting, making this approach redundant. Option B is wrong because setting minimum replicas to 0 does not control traffic distribution; it only affects autoscaling behavior, and gradually increasing replicas does not route a specific percentage of traffic to the new version. Option C is wrong because Cloud Load Balancing operates at the network layer and cannot intelligently split traffic between two Vertex AI endpoints based on model version; it would require additional proxy logic and defeats the purpose of Vertex AI's built-in traffic management.

41
MCQeasy

You deploy a new version of a model to a Vertex AI endpoint and want to gradually shift traffic from the old version to the new version over 24 hours. The endpoint currently serves 100% traffic to the old version. What should you do?

A.Use Vertex AI Experiments to run an A/B test between the two versions.
B.Deploy the new version to a separate endpoint and update your client to use the new endpoint for a percentage of requests.
C.Update the endpoint to split traffic between the two model versions using the traffic split configuration.
D.Delete the old version and redeploy the new version with a different endpoint name, then update DNS.
AnswerC

Vertex AI endpoints support traffic split configuration across deployed model versions, letting you assign percentage weights to each. Adjusting these weights gradually shifts requests from old to new, achieving the controlled 24-hour rollout without redeployment.

Why this answer

Vertex AI endpoints support a built-in traffic split configuration that allows you to gradually shift traffic between model versions deployed to the same endpoint. By updating the endpoint's traffic split percentages (e.g., from 100% old / 0% new to 0% old / 100% new over 24 hours), you can achieve a smooth, controlled rollout without changing client code or managing multiple endpoints.

Exam trap

Google often tests the misconception that traffic splitting requires separate endpoints or client-side logic, when in fact Vertex AI provides a native traffic split configuration on a single endpoint.

How to eliminate wrong answers

Option A is wrong because Vertex AI Experiments is designed for tracking and comparing model training runs, not for managing production traffic splits or A/B testing at the serving layer. Option B is wrong because deploying to a separate endpoint and updating the client to split requests manually introduces unnecessary complexity, client-side changes, and potential inconsistency; Vertex AI's traffic split feature handles this natively at the server side. Option D is wrong because deleting the old version and redeploying with a different endpoint name, then updating DNS, would cause a complete traffic cutover (not gradual) and disrupt service during the DNS propagation period, which can take minutes to hours.

42
MCQmedium

A company uses Vertex AI Vector Search for similarity search. They have a dataset of 10 million 512-dimensional vectors. Which index type should they choose for lowest latency at high recall?

A.Brute-force (flat) index
B.Approximate nearest neighbor (ANN) index with Scann
C.Tree-based index
D.Hashing-based index
AnswerB

Scann's approximate nearest neighbour index trades exact search for graph-based traversal, delivering far lower query latency at high recall on large vector sets. For 10 million 512-dimensional vectors, this satisfies the lowest-latency-at-high-recall constraint better than brute-force exact search.

Why this answer

For a dataset of 10 million 512-dimensional vectors, a brute-force (flat) index would be far too slow for low-latency queries. Approximate Nearest Neighbor (ANN) with ScaNN (Scalable Nearest Neighbors) is specifically designed by Google for high-dimensional vector search, offering sub-linear query time while maintaining high recall through techniques like anisotropic quantization and tree-based partitioning. This makes it the optimal choice for balancing latency and recall at this scale.

Exam trap

The trap here is that candidates often assume brute-force is the only way to guarantee high recall, but the question explicitly asks for lowest latency at high recall, which is the exact trade-off that ANN indexes like ScaNN are designed to optimize.

How to eliminate wrong answers

Option A is wrong because a brute-force (flat) index computes exact distances to every vector, resulting in O(N) complexity per query, which is prohibitively slow for 10 million vectors and cannot achieve low latency. Option C is wrong because tree-based indexes (e.g., KD-trees, R-trees) suffer from the curse of dimensionality in high-dimensional spaces (512-D), where their performance degrades to near brute-force due to the sparsity of data. Option D is wrong because hashing-based indexes (e.g., LSH) typically require multiple hash tables to achieve high recall, leading to high memory usage and often lower recall compared to optimized ANN methods like ScaNN, especially for 512-dimensional vectors.

43
MCQmedium

An ML engineer needs to update a model deployed on a Vertex AI endpoint without downtime. They want to gradually shift traffic to the new version while monitoring for errors. What is the correct procedure?

A.Use a canary deployment by deploying to a separate endpoint and using a load balancer with weighted routing.
B.Deploy the new model to a new endpoint, then update DNS to point to the new endpoint.
C.Delete the old model and deploy the new one with the same endpoint.
D.Deploy the new model to the same endpoint with 0% traffic initially, then gradually increase traffic while monitoring.
AnswerD

Deploying the new model to the same endpoint at 0% traffic creates a second deployed model behind the shared endpoint, letting you shift traffic incrementally via the endpoint's traffic split. This satisfies the no-downtime constraint: the existing version keeps serving while you monitor error rates before promoting the new one.

Why this answer

Vertex AI supports deploying multiple models to the same endpoint and splitting traffic via a deployed model's traffic percentage. The correct canary procedure is to deploy the new model version to the existing endpoint with 0% traffic, then gradually increase its traffic share while monitoring error rates and latency. This avoids downtime and enables safe rollback by shifting traffic back to the old version.

Exam trap

The trap here is confusing Vertex AI's native traffic-splitting capability with generic load-balancer-based canary patterns — candidates who don't know Vertex AI endpoints support multiple models with weighted traffic pick the external load balancer answer.

How to eliminate wrong answers

Option A is wrong because Vertex AI endpoints natively support traffic splitting — introducing an external load balancer and separate endpoint adds unnecessary complexity and is not the standard procedure. Option B is wrong because DNS-based cutover is a blue/green pattern with no gradual traffic control and DNS propagation delays cause unpredictable behavior. Option C is wrong because deleting the old model before the new one is validated risks downtime and eliminates rollback capability.

44
MCQmedium

You need to serve a large embedding model for similarity search with low latency. The model was trained to generate 256-dimensional embeddings. You plan to use Vertex AI Vector Search. Which index type should you choose to balance accuracy and performance for a dataset with 10 million vectors?

A.Tree-based index
B.Approximate nearest neighbor (ANN) index using ScaNN
C.Brute-force index
D.Hash-based index
AnswerB

ScaNN builds an approximate nearest neighbour index, trading exact recall for substantially lower query latency and memory at 10 million vectors. This satisfies the stem's balance of accuracy and performance, where exact search would be too slow at that scale.

Why this answer

Vertex AI Vector Search uses ScaNN (Scalable Nearest Neighbors) as its underlying ANN algorithm, which is specifically designed for high-dimensional embeddings (like 256-d) and large-scale datasets (10M vectors). ScaNN balances accuracy and performance by employing anisotropic quantization and tree-based partitioning, making it the optimal choice for low-latency similarity search without requiring exhaustive comparison.

Exam trap

Candidates often mistakenly choose Tree-based index (Option A) because ScaNN uses tree-based partitioning internally, but a standalone tree index fails in high dimensions. Vertex AI Vector Search’s ScaNN combines tree partitioning with quantization to overcome the curse of dimensionality and balance accuracy and performance for 10 million 256-d vectors.

How to eliminate wrong answers

Option A is wrong because a pure tree-based index (e.g., KD-tree) suffers from the 'curse of dimensionality' at 256 dimensions, where performance degrades to near brute-force levels. Option C is wrong because a brute-force index computes exact distances for all 10M vectors, resulting in O(n) latency that is unacceptable for real-time serving. Option D is wrong because hash-based indexes (e.g., LSH) are typically used for approximate nearest neighbor search in lower dimensions or for specific distance metrics, but they are not natively supported as a primary index type in Vertex AI Vector Search, and they often require extensive tuning to match ScaNN's accuracy-latency trade-off.

45
MCQeasy

You have a Vertex AI endpoint serving a model that returns predictions in about 200 ms. During a marketing campaign, traffic increases tenfold for short bursts. You want the endpoint to handle the bursts without manual intervention and without over-provisioning for the entire day. What should you do?

A.Increase the machine type to a larger instance with more memory and vCPUs.
B.Configure autoscaling on the endpoint with a minimum replica count of one and a maximum replica count that covers peak load.
C.Deploy the model to a batch prediction job and schedule it to run every hour.
D.Enable request logging and set up an alert when latency exceeds a threshold.
AnswerB

Vertex AI endpoint autoscaling adjusts replicas based on utilization or request concurrency. Setting a low minimum keeps cost down during normal traffic, while a maximum that covers peak load allows the endpoint to scale out during bursts. This matches the requirement to handle tenfold spikes automatically without over-provisioning all day, because replicas are added only when needed and removed when traffic subsides.

Why this answer

Vertex AI endpoints support autoscaling based on metrics such as CPU utilization or request concurrency. By setting a low minimum replica count, you keep baseline cost low during normal periods. A maximum replica count sized for peak load allows the endpoint to scale out automatically when traffic increases tenfold.

This elastic behavior handles short bursts without manual intervention and avoids paying for peak capacity all day, which is exactly what the scenario requires.

Exam trap

The trap here is thinking that a larger machine type provides elasticity, when autoscaling adds replicas horizontally and is what actually handles burst traffic.

46
Multi-Selectmedium

You are deploying a model to a Vertex AI Endpoint and need to reduce inference latency for a real-time application. Which two actions should you take? (Choose two.)

Select 2 answers
A.Enable request-response logging to monitor latency metrics.
B.Optimize the model for inference using techniques such as quantization or pruning.
C.Choose a machine type with a GPU accelerator that matches the model's compute profile.
D.Enable autoscaling with a high maximum replica count to handle traffic spikes.
E.Set the endpoint's minReplicaCount to zero to save cost during idle periods.
AnswersB, C

Quantization reduces numerical precision, shrinking model size and memory bandwidth needs, which often lowers latency. Pruning removes redundant weights, reducing computation. Both techniques can speed up inference on suitable hardware with minimal accuracy loss, making them effective for latency-sensitive real-time serving when applied carefully.

Why this answer

Latency reduction comes from making each prediction faster, not from adding capacity. Using a GPU that fits the model's compute profile accelerates the math, and inference-specific optimizations like quantization or pruning reduce the work per prediction. Autoscaling, zero minimum replicas, and logging affect scalability, availability, or observability rather than per-request latency.

Exam trap

The trap here is equating scalability features like autoscaling with latency improvements, when they primarily affect throughput and availability.

47
MCQmedium

You are A/B testing a new model version (challenger) against the current version (champion) on Vertex AI. You want to gradually shift traffic from champion to challenger while measuring business metrics. Which approach should you use?

A.Deploy the challenger to a separate endpoint and use a load balancer to route a percentage of requests.
B.Use Cloud Armor to route traffic based on headers.
C.Deploy both models to the same endpoint and use the traffic split feature to allocate percentages.
D.Create a new endpoint for the challenger and gradually shift DNS records.
AnswerC

Deploying both models to one Vertex AI endpoint lets the traffic split feature assign a percentage to each deployed model ID, shifting requests gradually from champion to challenger. This directly satisfies the gradual traffic-shift requirement while both versions serve live predictions, so business metrics can be compared under real traffic.

Why this answer

Vertex AI endpoints natively support traffic splitting between multiple deployed models on the same endpoint, allowing you to assign a percentage of prediction traffic to the champion and challenger. This is the built-in, supported mechanism for gradual rollouts and A/B experiments, and it integrates with Vertex AI's monitoring so you can measure business metrics per model version.

Exam trap

The trap here is assuming you need external infrastructure (load balancers, DNS, Cloud Armor) to split traffic, when Vertex AI endpoints already provide native traffic splitting as a first-class feature.

How to eliminate wrong answers

Option A is wrong because deploying to separate endpoints and adding an external load balancer bypasses Vertex AI's native traffic split and requires custom routing logic that Vertex AI already provides. Option B is wrong because Cloud Armor is a WAF/DDoS protection service operating at the edge for HTTP(S) load balancing, not a model traffic-routing mechanism for Vertex AI endpoints. Option D is wrong because DNS-based shifting is coarse, slow to propagate (TTL-dependent), and cannot provide the fine-grained percentage control or per-model metrics that Vertex AI traffic split offers.

48
Multi-Selecthard

Your team has deployed a model on Vertex AI endpoints and you are planning an A/B test to compare a new challenger model (v2) against the current champion (v1). The test should measure business metrics such as click-through rate. Which THREE steps should you take to set up the A/B test correctly? (Choose 3 correct answers)

Select 3 answers
A.Deploy the challenger model (v2) to the same endpoint as the champion (v1).
B.Modify your application to log which model version served each prediction.
C.Create a new endpoint for v2 and gradually shift DNS traffic.
D.Use Vertex AI Experiments to compare model performance.
E.Set up a traffic split between v1 and v2, e.g., 90% v1 and 10% v2.
AnswersA, B, E

Deploying both models to the same endpoint is required for A/B testing, because Vertex AI traffic splitting operates across model versions co-deployed on one endpoint. This lets requests be routed proportionally to v1 and v2 for click-through-rate comparison.

Why this answer

Option A is correct because Vertex AI endpoints natively support deploying multiple models to the same endpoint, which is the prerequisite for serving both champion v1 and challenger v2 behind a single prediction URL. Option E is correct because once both models share an endpoint, you configure a traffic split (e.g., 90% to v1 and 10% to v2) via the deployed model's traffic percentage, which is exactly how Vertex AI routes online prediction requests between model versions for an A/B test. Option B is correct because to measure business metrics such as click-through rate per variant, the application must log which model version (or deployed model ID) served each prediction so that downstream clicks can be attributed to v1 or v2.

Option C is not appropriate because shifting DNS traffic is a coarse, infrastructure-level approach that does not use Vertex AI's built-in traffic splitting and cannot reliably target a percentage of prediction requests. Option D is not appropriate because Vertex AI Experiments is designed for tracking and comparing training/evaluation runs and metrics, not for routing live prediction traffic between deployed model versions.

Exam trap

The trap here is that candidates confuse Vertex AI Experiments (for training) with endpoint traffic splitting (for serving), and they incorrectly think creating separate endpoints with DNS shifting is a valid A/B testing method, when Vertex AI's native traffic splitting is the correct and simpler approach.

49
MCQmedium

A company needs to run batch predictions on 10 TB of data stored in Cloud Storage. The predictions should be written to BigQuery. Which approach should they use?

A.Export the model to Cloud Functions and trigger on file upload
B.Create a Vertex AI Batch Prediction job with GCS input and BigQuery output
C.Use Vertex AI Online Prediction with a batch job
D.Use Dataflow to read from GCS and write to BigQuery, calling the model for each record
AnswerB

Vertex AI Batch Prediction natively reads GCS input and writes results directly to BigQuery, satisfying both the 10 TB Cloud Storage source and the BigQuery destination without intermediate exports. This managed service handles large-scale batch inference, so no custom pipeline or data movement code is required.

Why this answer

Vertex AI Batch Prediction natively supports reading input from Cloud Storage and writing predictions directly to BigQuery, making it the most efficient and fully managed solution for large-scale batch inference on 10 TB of data. This approach avoids the complexity of custom infrastructure or per-record model calls, leveraging Vertex AI's optimized batch processing pipeline.

Exam trap

This question tests the distinction between batch and online prediction modes in Vertex AI. The trap is that candidates may confuse Vertex AI's batch prediction with using Dataflow or Cloud Functions, not realizing that Vertex AI natively supports BigQuery as a direct output destination for batch jobs.

How to eliminate wrong answers

Option A is wrong because Cloud Functions are designed for event-driven, lightweight processing and cannot handle 10 TB of data efficiently; exporting a model to Cloud Functions also lacks native batch prediction orchestration and BigQuery output support. Option C is wrong because Vertex AI Online Prediction is intended for real-time, low-latency inference on individual requests, not for batch jobs; there is no 'batch job' mode within online prediction. Option D is wrong because while Dataflow can read from GCS and write to BigQuery, calling the model for each record would require custom code and per-record inference, which is less efficient and more complex than using Vertex AI's built-in batch prediction with direct BigQuery output.

50
MCQmedium

You need to deploy a PyTorch model for online inference on Vertex AI but the model was trained using custom ops that are not natively supported. You want to use NVIDIA Triton Inference Server for optimisation. How should you proceed?

A.Convert the model to TFLite and deploy on an edge device.
B.Build a custom container with NVIDIA Triton Inference Server and deploy it to Vertex AI.
C.Export the model to ONNX and deploy using Vertex AI's built-in TensorFlow serving.
D.Use Vertex AI Model Optimisation to automatically quantise the model.
AnswerB

Building a custom container lets you bundle Triton with the required custom op libraries, then deploy it as a Vertex AI custom container model. This satisfies the stem's constraint: Triton is not a natively supported Vertex AI serving framework, and the custom ops demand dependencies the prebuilt PyTorch containers lack.

Why this answer

Vertex AI supports custom containers for prediction, so you can package NVIDIA Triton Inference Server with your PyTorch model and any custom ops/libraries it needs, then deploy that container as a Vertex AI Model. Triton natively supports PyTorch (via TorchScript), ONNX, TensorRT, and custom backends, making it the right choice when the model relies on non-standard operators that Vertex AI's pre-built PyTorch/TensorFlow containers cannot execute.

Exam trap

The trap here is assuming that exporting to ONNX or using Vertex AI Model Optimisation will magically make unsupported custom ops work; the exam tests whether you know that custom containers are the escape hatch for non-standard runtimes and operators.

How to eliminate wrong answers

Option A is wrong because TFLite targets mobile/edge deployment and cannot run arbitrary PyTorch custom ops, and it does not address online inference on Vertex AI. Option C is wrong because exporting to ONNX does not guarantee the custom ops are supported by TensorFlow Serving, and Vertex AI's built-in TF serving container cannot execute PyTorch custom operators. Option D is wrong because Vertex AI Model Optimisation performs quantization/pruning on supported model formats but does not add support for unsupported custom ops or change the serving runtime.

51
MCQmedium

An ML engineer needs to run batch predictions on 10 TB of data stored in BigQuery using a TensorFlow model. The predictions must be written to BigQuery. Which service should they use?

A.Create a Dataflow pipeline to read from BigQuery, run the model using Python, and write results to BigQuery.
B.Export BigQuery data to GCS, run batch prediction on GCS, then load results back to BigQuery.
C.Use Vertex AI online prediction with batch requests.
D.Use Vertex AI Batch Prediction with BigQuery source and sink.
AnswerD

Batch Prediction reads directly from BigQuery as the input source and writes predictions back to a BigQuery sink, avoiding 10 TB of manual export. This natively satisfies both the large-scale source and the required output destination.

Why this answer

Vertex AI Batch Prediction natively supports BigQuery as both input source and output sink, so the engineer can submit a batch job directly against the BigQuery table and have predictions written back to BigQuery without any data movement code. This is the most efficient and least error-prone option for 10 TB of data.

Exam trap

PMLE often tests whether candidates know Vertex AI Batch Prediction supports BigQuery natively — many pick the GCS export/import route or a custom Dataflow pipeline because they assume data must be staged in GCS first, which is not required.

How to eliminate wrong answers

Option A is wrong because building a custom Dataflow pipeline to load the TensorFlow model and run inference is unnecessary engineering overhead when Vertex AI Batch Prediction already handles BigQuery I/O natively, and Dataflow is not optimized for GPU/TPU model inference. Option B is wrong because exporting 10 TB to GCS, running batch prediction, then reloading to BigQuery adds significant latency, storage cost, and operational complexity compared to native BigQuery source/sink. Option C is wrong because Vertex AI online prediction is designed for low-latency, small-payload requests and has strict request size and QPS limits — it cannot handle 10 TB batch workloads efficiently.

52
MCQmedium

A company needs to serve a high-throughput prediction service with strict latency requirements. They want to minimize cold starts and ensure consistent performance. Which endpoint configuration is most appropriate?

A.Set min_replicas to an estimated baseline and max_replicas to a higher number
B.Set min_replicas and max_replicas equal to a fixed number
C.Set min_replicas to 0 and max_replicas to a high number
D.Do not set min_replicas; let Vertex AI automatically determine
AnswerA

Pinning min_replicas to an estimated baseline keeps warm capacity always available, eliminating cold starts, while max_replicas absorbs traffic spikes. This satisfies the stem's high-throughput, strict-latency and consistent-performance requirements, unlike autoscaling from zero, which reintroduces cold-start delays.

Why this answer

Setting min_replicas to an estimated baseline ensures that a minimum number of instances are always running, eliminating cold starts for baseline traffic. Setting max_replicas to a higher number allows the service to scale up to handle traffic spikes while maintaining consistent performance. This configuration balances cost and latency by avoiding the overhead of scaling from zero while still accommodating bursts.

Exam trap

Google Cloud often tests the misconception that setting min_replicas to 0 is cost-effective, but the trap here is that it ignores the strict latency requirement and the reality of cold start delays in model serving with Vertex AI.

How to eliminate wrong answers

Option B is wrong because setting min_replicas and max_replicas equal to a fixed number prevents any autoscaling, leading to either over-provisioning (waste) or under-provisioning (latency spikes) under variable load. Option C is wrong because setting min_replicas to 0 means the service can scale down to zero, causing cold starts on every request when traffic resumes, which violates the strict latency requirement. Option D is wrong because not setting min_replicas and relying on Vertex AI's automatic determination may result in the service scaling to zero or having unpredictable baseline capacity, introducing cold starts and inconsistent performance.

53
MCQmedium

An application serving predictions from a Vertex AI endpoint receives many identical requests within a short time window. The team notices redundant computation and wants to cache responses to reduce latency and cost. What is the recommended solution?

A.Deploy the model on a larger machine type to handle duplicate requests faster.
B.Enable Vertex AI endpoint caching by setting the `enable_cache` flag.
C.Implement a cache layer using Cloud Memorystore for Redis, hashing prediction requests.
D.Use Cloud CDN in front of the endpoint.
AnswerC

Cloud Memorystore for Redis supplies a low-latency shared cache; hashing each prediction request produces a deterministic key so identical requests return the stored response instead of recomputing. This directly removes the redundant computation and cost the stem describes.

Why this answer

Vertex AI does not provide built-in request caching; instead, the recommended pattern is to implement an external cache like Cloud Memorystore for Redis. By hashing the prediction request payload and using it as a cache key, identical requests within the short time window can be served from Redis, eliminating redundant model inference and reducing both latency and cost.

Exam trap

The trap here is that candidates assume Vertex AI has a native caching feature (like an `enable_cache` flag) because other Google Cloud services offer caching, but Vertex AI endpoints require an external cache layer like Cloud Memorystore for Redis.

How to eliminate wrong answers

Option A is wrong because deploying on a larger machine type increases throughput but does not eliminate redundant computation for identical requests; it still performs the same inference multiple times, wasting resources. Option B is wrong because Vertex AI endpoints do not support an `enable_cache` flag; this is a fictitious feature, and Vertex AI has no built-in request caching mechanism. Option D is wrong because Cloud CDN caches static content at the edge based on HTTP cache headers, but prediction requests are typically POST with dynamic payloads that are not cacheable by CDN, and CDN cannot inspect or hash request bodies for deduplication.

54
Multi-Selecthard

You are designing a batch prediction pipeline using Vertex AI. The input data is 50 TB in CSV format on GCS. The model requires feature engineering that involves complex transformations (e.g., datetime parsing, one-hot encoding). Which TWO services or steps should you include in your pipeline?

Select 2 answers
A.Use Cloud Functions to transform each file individually.
B.Use Cloud SQL to store intermediate results.
C.Run Vertex AI batch prediction job with GCS source pointing to the processed TFRecord files.
D.Use Dataflow to read CSV, perform feature engineering, and write to GCS in TFRecord format.
E.Use Dataflow to read CSV, perform feature engineering, and write to BigQuery.
AnswersC, D

Running a Vertex AI batch prediction job against GCS-hosted TFRecord files lets the pipeline consume the preprocessed, engineered features at scale, avoiding repeated transformation of 50 TB of raw CSV. TFRecord input suits large batch workloads efficiently.

Why this answer

Option D is correct because Dataflow is the managed, horizontally scalable Apache Beam service designed for large-scale ETL like 50 TB of CSV, and it can perform the required complex transformations (datetime parsing, one-hot encoding) before writing the engineered features to GCS in TFRecord format, which is an efficient binary format for ML training and prediction. Option C is correct because a Vertex AI batch prediction job can then consume those processed TFRecord files directly from GCS as its input source, letting the model run inference at scale without re-doing feature engineering. Option A is not appropriate because Cloud Functions is an event-driven, short-lived serverless compute service with limited memory and execution time, making it unsuitable for transforming 50 TB of data file-by-file.

Option B is not appropriate because Cloud SQL is a relational OLTP database, not a scalable store for massive intermediate ML feature data. Option E is not appropriate because writing engineered features to BigQuery does not produce the TFRecord input that Vertex AI batch prediction expects, and BigQuery is not the target format for this pipeline's model input.

Exam trap

Candidates often assume that any scalable service can handle batch processing, but Cloud Functions and Cloud SQL are unsuitable for 50TB. The trap is to think that more than two services are needed, but a Dataflow pipeline and Vertex AI batch prediction are sufficient.

55
MCQmedium

A company uses Vertex AI Vector Search (Matching Engine) for a product recommendation system. The product embeddings are updated hourly. Which index update method should they use to ensure low latency for new items?

A.Batch rebuild the index every hour
B.Use streaming updates to add new embeddings incrementally
C.Create a new index each hour and swap endpoints
D.Use brute-force index to simplify updates
AnswerB

Streaming updates let the index ingest new product embeddings incrementally without a full rebuild, so hourly-refreshed items become searchable almost immediately. This satisfies the low-latency requirement for new items, whereas batch rebuilds would delay their availability.

Why this answer

Vertex AI Vector Search supports streaming updates, allowing new embeddings to be added incrementally without rebuilding the entire index. This ensures low latency for new items by making them searchable almost immediately after update, which is critical for hourly refresh cycles where batch rebuilds would introduce significant delay.

Exam trap

The trap here is that candidates often assume batch rebuilds are the only reliable method for consistency, overlooking that streaming updates in Vertex AI Vector Search are designed specifically for low-latency incremental ingestion without sacrificing search quality.

How to eliminate wrong answers

Option A is wrong because batch rebuilding the index every hour incurs high latency and computational cost, as the entire index must be reconstructed from scratch, delaying availability of new items. Option C is wrong because creating a new index each hour and swapping endpoints is inefficient and introduces downtime during the swap, plus it requires managing multiple index versions unnecessarily. Option D is wrong because brute-force indices do not simplify updates; they perform exhaustive linear scans, which are slow and unscalable for large embedding sets, and they lack the optimized approximate nearest neighbor (ANN) search that Vector Search provides.

56
MCQeasy

A machine learning engineer wants to deploy a trained model to Vertex AI for online predictions. Which Vertex AI resource is required to serve the model and provide an endpoint URL?

A.Vertex AI Pipeline
B.Vertex AI Model Registry
C.Vertex AI Feature Store
D.Vertex AI Endpoint
AnswerD

An Endpoint is the managed Vertex AI resource that hosts the deployed model and exposes a serving URL for online predictions. Deploying the model to an Endpoint allocates compute and returns the endpoint URL the engineer needs to send prediction requests.

Why this answer

Vertex AI Endpoint is the required resource to deploy a trained model for online predictions, as it provides a dedicated endpoint URL that accepts prediction requests and routes them to the model. Without an endpoint, the model cannot be accessed via HTTP/HTTPS for real-time inference, which is the core requirement for online serving.

Exam trap

The trap here is that candidates confuse the Model Registry (which stores and versions models) with the actual serving infrastructure, assuming that registering a model automatically creates an endpoint, when in fact a separate Endpoint resource must be created and the model must be deployed to it.

How to eliminate wrong answers

Option A is wrong because Vertex AI Pipeline is used for orchestrating and automating ML workflows (e.g., training, evaluation), not for serving models or providing an endpoint URL. Option B is wrong because Vertex AI Model Registry is a central repository for managing model versions and metadata, but it does not itself expose an endpoint for predictions; models must be deployed to an endpoint for serving. Option C is wrong because Vertex AI Feature Store is designed for storing, serving, and sharing feature data for training and prediction, not for hosting models or providing inference endpoints.

57
MCQeasy

A startup is deploying a scikit-learn model to Vertex AI for online predictions. They want to minimize the effort required to containerize the model and ensure it can handle HTTP requests. What should they do?

A.Use the pre-built scikit-learn container provided by Vertex AI and deploy the model by specifying the model artifact in Cloud Storage.
B.Convert the scikit-learn model to TensorFlow SavedModel format and use the TensorFlow pre-built container for deployment.
C.Deploy the model using a custom container that includes the scikit-learn library and a simple HTTP server implemented with Python's http.server module.
D.Write a custom Flask application to load the model and expose a /predict endpoint, then build a Docker image and push it to Artifact Registry.
AnswerA

Vertex AI provides pre-built containers for popular frameworks like scikit-learn. These containers handle HTTP serving, request parsing, and model loading automatically. By simply pointing to the model artifact, the startup avoids containerization effort and ensures compatibility with Vertex AI's prediction API.

Why this answer

The pre-built scikit-learn container on Vertex AI eliminates the need for custom containerization and HTTP server code. It automatically handles model loading and prediction requests, making it the lowest-effort solution for deploying scikit-learn models.

Exam trap

The trap here is assuming that a custom container is always needed, overlooking the convenience and compatibility of Vertex AI's pre-built containers for common frameworks.

58
MCQeasy

Your Vertex AI endpoint receives many identical prediction requests (same input features). You want to cache responses to reduce latency and cost. Which Google Cloud service should you use?

A.Cloud Memorystore for Redis
B.Cloud CDN
C.Bigtable
D.Cloud Storage with object versioning
AnswerA

Memorystore for Redis provides a low-latency in-memory store that sits in front of the endpoint, letting identical feature vectors return cached predictions instead of re-invoking the model. This directly cuts both response latency and per-prediction cost, matching the stem's caching goal.

Why this answer

Cloud Memorystore for Redis is an in-memory data store that provides sub-millisecond latency, making it ideal for caching prediction responses. By caching identical prediction requests, you can reduce the number of calls to the Vertex AI endpoint, lowering latency and cost. Redis supports key-value storage with TTL, perfect for caching.

Exam trap

The trap is selecting Cloud CDN because it is a caching service, but CDN caches HTTP responses at the edge and is not suitable for caching dynamic API prediction results that require custom key logic.

How to eliminate wrong answers

Option B is wrong because Cloud CDN caches HTTP content at edge locations, but it is designed for static or cacheable web content, not for caching API prediction responses that may require authentication and dynamic inputs. Option C is wrong because Bigtable is a NoSQL database for large-scale analytical workloads, not a low-latency cache. Option D is wrong because Cloud Storage with object versioning is for storing objects with version history, not for caching API responses.

59
MCQeasy

Your company runs a high-traffic web application that serves the same machine learning model prediction for many identical requests (e.g., product recommendations for the same user profile). You want to reduce latency and load on the prediction endpoint by caching responses. Which Google Cloud service should you use?

A.Cloud CDN
B.Cloud Memorystore
C.Cloud Spanner
D.BigQuery
AnswerB

Cloud Memorystore provides a managed Redis or Memcached layer that stores prediction results keyed by request parameters, so identical requests return cached responses without invoking the endpoint. This cuts both latency and endpoint load, matching the repeated-prediction pattern described.

Why this answer

Cloud Memorystore (B) is correct because it provides a managed in-memory cache (Redis or Memcached) that can store the results of identical prediction requests, reducing latency and load on the prediction endpoint. By caching responses keyed on the user profile or request parameters, subsequent identical requests can be served directly from Memorystore in microseconds, avoiding redundant model inference.

Exam trap

The trap here is that candidates confuse caching at the edge (CDN) with caching at the application layer (Memorystore), assuming any cache service works for dynamic API responses, but Cloud CDN cannot cache POST requests or application-specific payloads without significant configuration and still lacks the fine-grained key-value semantics needed for identical prediction requests.

How to eliminate wrong answers

Option A (Cloud CDN) is wrong because it caches static content (e.g., images, CSS) at edge locations, not dynamic API responses for identical requests; it cannot cache POST request payloads or application-level prediction results without complex workarounds. Option C (Cloud Spanner) is wrong because it is a globally distributed relational database designed for transactional consistency and high availability, not for low-latency caching of ephemeral prediction responses. Option D (BigQuery) is wrong because it is a serverless data warehouse for analytical queries on large datasets, not a caching layer for real-time inference results.

60
MCQhard

An ML platform team deploys the same custom container to two Vertex AI endpoints: one for interactive scoring and one for nightly bulk scoring. The interactive endpoint must return predictions in under 200 ms and receives small single-record requests. The bulk endpoint sends large batched requests and tolerates seconds of latency. Both endpoints currently use the same machine type and the same container image, and the interactive endpoint frequently misses its latency target. Which change best resolves the interactive latency problem?

A.Increase the request timeout on the interactive endpoint and add more replicas.
B.Enable autoscaling on the bulk endpoint so it absorbs the interactive traffic as well.
C.Move both endpoints to a machine type with more vCPUs and more memory.
D.Deploy a container variant tuned for the interactive endpoint: load the model once at startup, keep it resident, and avoid per-request batch assembly, then serve it on a latency-optimized machine type.
AnswerD

Latency on small requests is dominated by fixed per-request overhead, so removing repeated model loading and batch assembly from the request path gives the largest gain. A container that loads the model once and keeps it warm in memory, paired with a machine type chosen for single-request speed rather than bulk throughput, directly targets the 200 ms SLO while leaving the bulk endpoint's throughput-optimized configuration untouched.

Why this answer

Interactive single-record scoring and bulk batched scoring have opposite performance profiles, so they should not share one serving configuration. The interactive path is dominated by fixed per-request overhead such as model loading and batch assembly, so a container that loads the model once, keeps it resident, and skips per-request batching removes that cost. Pairing it with a machine type selected for low-latency single requests, while leaving the bulk endpoint tuned for throughput, resolves the SLO breach without over-sizing the bulk tier.

Exam trap

The trap here is treating a latency SLO breach as a pure capacity problem and adding replicas or vCPUs, when the fixed per-request overhead of batch-oriented model loading is the actual cause.

61
MCQeasy

You are deploying a model to a Vertex AI endpoint and need to minimize latency for online predictions. Which machine type should you choose?

A.n1-standard-2 with NVIDIA Tesla T4
B.e2-standard-2
C.n1-standard-2
D.n1-highmem-2
AnswerA

NVIDIA Tesla T4 GPUs deliver low-latency inference for online predictions, and the n1-standard-2 shape supplies adequate vCPU and memory for a single model replica. This satisfies the stem's latency-minimisation constraint, since T4s accelerate matrix operations that CPU-only machine types would process far more slowly.

Why this answer

For online predictions with minimal latency, a machine type with a GPU is essential because it accelerates model inference. The n1-standard-2 with NVIDIA Tesla T4 provides GPU acceleration, which significantly reduces latency compared to CPU-only instances. The other options lack GPUs and are not optimized for low-latency inference.

Exam trap

The trap is overlooking the need for a GPU for low-latency inference and selecting a CPU-only machine type based on cost or memory, when the question explicitly asks to minimize latency.

How to eliminate wrong answers

Option B is wrong because e2-standard-2 is a general-purpose CPU-only machine type, which would result in higher latency for model inference. Option C is wrong because n1-standard-2 is CPU-only and lacks the parallel processing power of a GPU for inference. Option D is wrong because n1-highmem-2 is optimized for memory-intensive workloads but still CPU-only, so it does not minimize latency for predictions.

62
MCQmedium

A media company uses a Vertex AI endpoint to serve a video recommendation model. The model is updated weekly with new embeddings. They want to minimize downtime during model updates and ensure that the new model performs well before fully rolling it out. They also need to be able to revert quickly if issues arise. What should they do?

A.Deploy the new model to a new endpoint and use a load balancer to distribute traffic between the old and new endpoints.
B.Create a new endpoint for the new model and update the application to call the new endpoint after testing.
C.Use a batch prediction job to test the new model on a large dataset, then replace the existing model on the endpoint.
D.Deploy the new model as a new version on the same endpoint, set it to receive 0% of traffic initially, then gradually increase traffic while monitoring performance.
AnswerD

Deploying a new model version on the same endpoint allows for seamless traffic splitting without downtime. Starting with 0% traffic lets you validate the new model with a small percentage of live traffic or via direct prediction requests. Gradually increasing traffic enables A/B testing and performance monitoring. If issues arise, you can quickly shift traffic back to the previous version.

Why this answer

The optimal approach is to deploy the new model as a version on the existing endpoint and use traffic splitting to gradually shift traffic. This provides zero-downtime updates, allows performance validation with real traffic, and enables quick rollback by adjusting traffic percentages. Other methods either cause downtime, lack gradual rollout, or add unnecessary complexity.

Exam trap

The trap here is thinking that a new endpoint is required for testing, when Vertex AI endpoints support multiple model versions and traffic splitting natively.

63
MCQeasy

You need to run a batch prediction job on Vertex AI for a large dataset stored in BigQuery. The model expects CSV input. Which input format should you specify for the batch prediction job?

A.JSON Lines files on Cloud Storage
B.CSV files on Cloud Storage
C.TFRecord files on Cloud Storage
D.BigQuery table
AnswerB

For batch prediction, Vertex AI supports CSV files on Cloud Storage as an input format. Since the model expects CSV, providing CSV files on Cloud Storage directly matches the model's input requirements and avoids any format conversion or export step.

Why this answer

Vertex AI batch prediction supports multiple input formats, but the format must match what the model expects. Since the model expects CSV, specifying CSV files on Cloud Storage is the correct choice. This avoids any need for data conversion and ensures the batch prediction job can parse the input correctly.

Exam trap

The trap here is assuming that any supported input format works regardless of the model's expected input, when the format must match the model's signature.

64
MCQmedium

You have a champion model serving 100% traffic on a Vertex AI endpoint. You want to deploy a challenger model and gradually shift 10% of traffic to it for A/B testing. What is the correct approach?

A.Use Cloud Run to deploy both models and use Cloud Endpoints for traffic splitting.
B.Deploy the challenger on the same endpoint and use the traffic split parameter to allocate 10% traffic to it.
C.Deploy the challenger on a separate endpoint and use Cloud Armor to split traffic.
D.Create a new endpoint for the challenger and route 10% of requests via a load balancer.
AnswerB

Deploying the challenger to the same endpoint lets Vertex AI distribute prediction requests across both models using its built-in traffic split, which accepts percentage allocations per deployed model ID. Setting the challenger's split to 10 satisfies the gradual A/B requirement without a second endpoint, DNS changes, or client-side routing logic.

Why this answer

Vertex AI endpoints support traffic splitting by deploying multiple model versions and assigning traffic percentages. You deploy the challenger as a new deployed model on the same endpoint and set traffic split: champion 90%, challenger 10%.

65
MCQmedium

A team is deploying a large PyTorch model for online inference. They want to use NVIDIA Triton Inference Server to optimize serving performance. How can they integrate Triton with Vertex AI?

A.Package the model with Triton in a custom container and deploy it to Vertex AI
B.Vertex AI automatically uses Triton for all PyTorch models
C.Deploy the model to GKE and use Vertex AI as a frontend
D.Use a prebuilt Vertex AI PyTorch container that includes Triton
AnswerA

Vertex AI lets you supply a custom prediction container, so bundling the PyTorch model with Triton and its model repository into that image gives Triton full control of inference, enabling its batching and optimisation features on Vertex AI's managed endpoint infrastructure.

Why this answer

Vertex AI supports custom containers for prediction, so the standard integration pattern is to build a container that bundles the model with NVIDIA Triton Inference Server and deploy it as a Vertex AI Model with a custom prediction container. This lets Triton handle dynamic batching, model ensembles, and multi-framework serving while Vertex AI manages endpoints, autoscaling, and monitoring.

Exam trap

PMLE often tests the misconception that managed services like Vertex AI automatically apply third-party optimizers (Triton, TensorRT) to any framework — in reality, integration requires the candidate to package the optimizer inside a custom container.

How to eliminate wrong answers

Option B is wrong because Vertex AI does not automatically wrap PyTorch models with Triton — the default prebuilt PyTorch containers use TorchServe or a simple Flask-based handler, not Triton. Option C is wrong because deploying to GKE and using Vertex AI as a 'frontend' is not a supported integration pattern; Vertex AI endpoints run on managed infrastructure, not on customer GKE clusters as a frontend proxy. Option D is wrong because Google's prebuilt Vertex AI PyTorch containers do not include Triton Inference Server — Triton must be added by the customer in a custom container.

66
MCQmedium

You are deploying a custom PyTorch model to a Vertex AI Endpoint for real-time inference. The model artifact is stored in a Cloud Storage bucket. Your security team requires that the model be served from a container that runs as a non-root user and has no network access except to the Vertex AI prediction service. Which deployment configuration should you use?

A.Deploy the model using a custom container that runs as root but configure the endpoint to use a private VPC with no external IP.
B.Use the pre-built PyTorch container provided by Vertex AI and set the endpoint to use a private VPC with no external IP.
C.Deploy the model using a custom container that sets the USER instruction to a non-root user in the Dockerfile, and configure the endpoint to use a private VPC with no external IP.
D.Deploy the model using a custom container that sets the USER instruction to a non-root user, but allow the endpoint to have an external IP for easier debugging.
AnswerC

A custom container allows you to control the user and network settings. Setting USER to a non-root user in the Dockerfile satisfies the non-root requirement, and deploying the endpoint in a private VPC with no external IP restricts network access to only the Vertex AI prediction service. This meets both security constraints.

Why this answer

A custom container is necessary to control the user context, and deploying it in a private VPC with no external IP ensures network isolation. The pre-built container runs as root, and allowing an external IP breaks the network restriction. Thus, the configuration that combines a non-root custom container with a private VPC is the only one that satisfies both security requirements.

Exam trap

The trap here is assuming that Vertex AI's pre-built containers automatically run as non-root or that network isolation alone is sufficient without controlling the container user.

67
MCQmedium

A retail team must run nightly batch predictions over 20 TB of Parquet data stored in Cloud Storage using a custom PyTorch model registered in Vertex AI Model Registry. They want the job to finish within a fixed maintenance window and prefer not to manage the underlying compute. Which configuration should they use?

A.Schedule the model on a Vertex AI pipeline with a Dataflow step that calls the endpoint for each record, using streaming inserts to write results back to Cloud Storage.
B.Deploy the model to a Vertex AI endpoint with autoscaling and write a client script that submits all rows as individual online prediction requests.
C.Create a Vertex AI CustomJob that runs a PyTorch training script with the input path as a parameter, since CustomJob automatically performs batch inference when given a Parquet input.
D.Create a Vertex AI BatchPredictionJob with the model from Model Registry, set the input to the Cloud Storage Parquet path, and specify a machine type plus a starting and maximum replica count for the distributed workers.
AnswerD

Batch prediction reads directly from Cloud Storage, shards the input across workers, and scales horizontally according to the replica count you configure, all on managed infrastructure. Setting a maximum replica count lets the job finish within the maintenance window without the team provisioning or patching any compute themselves.

Why this answer

Batch prediction is purpose-built for large offline scoring: it reads files directly from Cloud Storage, distributes work across managed replicas, and writes sharded output without any endpoint. Specifying machine type and replica bounds gives the team control over throughput so the job completes inside the maintenance window while Vertex AI handles provisioning.

Exam trap

The trap here is assuming that a CustomJob or an online endpoint is needed for large-scale inference when the managed batch prediction service already covers it.

68
MCQmedium

A team has deployed a model on Vertex AI and wants to cache frequent identical prediction requests to improve latency and reduce cost. Which Google Cloud service should they use?

A.Cloud Bigtable
B.Cloud CDN
C.Cloud Memorystore
D.Cloud SQL
AnswerC

Cloud Memorystore offers a managed in-memory cache that the serving path can query before invoking the model, returning stored responses for repeated identical requests. This reduces endpoint invocations, lowering both latency and cost as the stem requires.

Why this answer

Cloud Memorystore provides a managed Redis or Memcached in-memory cache that can store frequent identical prediction requests and their responses, dramatically reducing latency and backend load. It is the standard Google Cloud service for application-level caching in front of Vertex AI endpoints.

Exam trap

PMLE often tests the distinction between caching layers — candidates may pick Cloud CDN for 'caching' without realizing CDN only caches HTTP responses at the edge and cannot key on arbitrary prediction payloads.

How to eliminate wrong answers

Option A is wrong because Cloud Bigtable is a wide-column NoSQL database optimized for high-throughput analytics and time-series data, not for low-latency key-value caching of prediction responses. Option B is wrong because Cloud CDN caches HTTP content at edge locations for static assets, not dynamic prediction API responses keyed by request payload. Option D is wrong because Cloud SQL is a relational database with millisecond-to-tens-of-milliseconds latency, far slower than in-memory caching and not designed for high-QPS cache workloads.

69
MCQmedium

A company wants to run batch predictions on millions of records stored in BigQuery. They need to preprocess the data (e.g., feature engineering) before feeding it to the model. Which approach is most scalable and cost-effective?

A.Use a large DataProc cluster to preprocess and run batch predictions.
B.Preprocess inline in the batch prediction job using a custom container.
C.Use a custom Python script on a Compute Engine instance.
D.Preprocess with Cloud Dataflow, output to Cloud Storage, then submit a Vertex AI batch prediction job.
AnswerD

Cloud Dataflow performs distributed preprocessing at scale, writing engineered features to Cloud Storage, which Vertex AI batch prediction reads directly. This decouples heavy transformation from the prediction job, satisfying the millions-of-records scalability and cost-efficiency constraint better than in-job preprocessing.

Why this answer

The most scalable and cost-effective because Cloud Dataflow (Apache Beam) provides serverless, auto-scaling preprocessing that handles large volumes of data efficiently, and Vertex AI batch predictions natively read from Cloud Storage, avoiding the need to manage infrastructure. This decouples preprocessing from prediction, allowing each to scale independently and minimizing costs by using ephemeral, pay-per-use resources.

Exam trap

A common mistake is assuming a single large cluster (Dataproc) or a single VM is sufficient for batch processing, when in fact serverless, auto-scaling services like Dataflow are more appropriate for large-scale, ephemeral preprocessing tasks.

How to eliminate wrong answers

Option A is wrong because Dataproc clusters require manual sizing, incur idle costs, and add operational overhead for a simple preprocessing task; it is overkill and less cost-effective than serverless options. Option B is wrong because preprocessing inline in a custom container for batch prediction tightly couples preprocessing with prediction, preventing independent scaling and making it harder to handle large-scale data transformations efficiently. Option C is wrong because a single Compute Engine instance cannot scale horizontally to process millions of records in a reasonable time, and managing failover, retries, and parallelization would require custom code, making it neither scalable nor cost-effective.

70
MCQmedium

A machine learning team deploys a PyTorch model for online prediction on Vertex AI using a custom container. They notice that the first few requests after scaling up experience high latency. What is the most likely cause and how should they mitigate it?

A.The endpoint is not configured for autoscaling; enable min_replica=0 to allow scale-to-zero.
B.The model file is corrupted; re-upload to Vertex AI Model Registry.
C.The container has a slow initialization; set initialDelaySeconds in the health probe to give more time before considering the pod ready.
D.Use a smaller machine type (n1-standard-2) to reduce startup overhead.
AnswerC

Slow container initialisation delays readiness, so the pod receives traffic before PyTorch and model weights finish loading, producing cold-start latency after scaling. Raising initialDelaySeconds postpones the readiness probe's first check, preventing premature traffic routing until initialisation completes. This directly addresses the scaling-induced latency constraint.

Why this answer

The high latency on the first few requests after scaling up is a classic symptom of a slow container initialization. By setting `initialDelaySeconds` in the health probe, you allow the container more time to start up and become ready before it receives traffic, preventing premature routing that causes timeouts or retries. This is a common tuning parameter for custom containers on Vertex AI, where model loading or dependency initialization can take several seconds.

Exam trap

The trap here is that candidates confuse slow initialization with autoscaling misconfiguration, assuming that scale-to-zero or smaller machines would fix the latency, when in fact the root cause is the readiness probe timing.

How to eliminate wrong answers

Option A is wrong because the problem occurs after scaling up, not from a cold start with zero replicas; setting min_replica=0 would actually worsen latency by requiring full cold starts. Option B is wrong because a corrupted model file would cause persistent prediction failures or errors, not just high latency on the first few requests after scaling. Option D is wrong because using a smaller machine type (n1-standard-2) would increase startup overhead and latency, not reduce it, as it provides fewer CPU and memory resources for initialization.

71
Multi-Selecthard

A team is building a batch prediction pipeline that processes raw data from Cloud Storage, performs complex preprocessing, and then runs predictions using a large model. The preprocessing step is compute-intensive and the prediction step is I/O-bound. Which TWO Google Cloud services should they combine to optimize cost and performance? (Choose 2)

Select 2 answers
A.Dataflow for preprocessing and writing results to Cloud Storage
B.Cloud Functions to preprocess data row by row
C.Cloud Run to serve the preprocessed data as an API
D.Vertex AI Batch Prediction with Cloud Storage source
E.Vertex AI Batch Prediction with BigQuery source
AnswersA, D

Dataflow provides autoscaling, distributed workers suited to the compute-intensive preprocessing stage, reading raw Cloud Storage data and writing transformed output back to Cloud Storage. This satisfies the stem's constraint by matching elastic compute to preprocessing while keeping the I/O-bound prediction stage separate.

Why this answer

Option A (Dataflow for preprocessing and writing results to Cloud Storage) is correct because Dataflow is a fully managed, autoscaling service built on Apache Beam that handles compute-intensive, parallel batch preprocessing over data in Cloud Storage, and it can write the transformed output back to Cloud Storage for the next stage. Option D (Vertex AI Batch Prediction with Cloud Storage source) is correct because it is the managed, I/O-bound batch inference service designed to read preprocessed files directly from Cloud Storage and run predictions on a large model without provisioning persistent serving infrastructure, which optimizes cost for batch workloads. Option B is wrong because Cloud Functions is event-driven and limited in execution time and memory, making it unsuitable for compute-intensive row-by-row preprocessing of large batch datasets.

Option C is wrong because Cloud Run serving preprocessed data as an API adds unnecessary request/response overhead and is designed for online serving, not batch pipeline handoff. Option E is wrong because the pipeline explicitly reads raw data from Cloud Storage and writes preprocessed results there, so a BigQuery source for batch prediction does not match the described data flow.

Exam trap

Google often tests the distinction between batch and online serving patterns, and the trap here is that candidates may choose Cloud Functions or Cloud Run for preprocessing because they are familiar serverless options, without realizing that Dataflow is purpose-built for large-scale, compute-intensive batch processing and that Vertex AI Batch Prediction is the correct service for offline inference at scale.

72
Multi-Selecteasy

Which TWO of the following can be used as input sources for Vertex AI batch prediction jobs? (Choose 2)

Select 2 answers
A.Cloud Firestore
B.Cloud SQL
C.BigQuery
D.Cloud Spanner
E.Cloud Storage
AnswersC, E

BigQuery is a supported input source for Vertex AI batch prediction jobs, letting you point the job directly at a table or query. This avoids exporting data to Cloud Storage first, satisfying the question's requirement for valid batch prediction input sources.

Why this answer

Option C (BigQuery) is correct because Vertex AI batch prediction jobs natively accept a BigQuery table as the input source, specified via the bigQuerySource field in the BatchPredictionJob, allowing you to run predictions directly against tabular data stored in BigQuery. Option E (Cloud Storage) is correct because Vertex AI batch prediction jobs also accept input files (such as JSONL, CSV, or TFRecord) stored in a Cloud Storage bucket, specified via the gcsSource field with a URI like gs://bucket/file.jsonl. Options A (Cloud Firestore), B (Cloud SQL), and D (Cloud Spanner) are not valid input sources for Vertex AI batch prediction; these are operational databases that are not directly supported as batch prediction inputs, so you would need to export their data to BigQuery or Cloud Storage first.

Exam trap

The trap here is that candidates often assume any Google Cloud database (like Firestore, Cloud SQL, or Spanner) can serve as a direct input source for batch predictions, but Vertex AI batch prediction only supports BigQuery and Cloud Storage as input sources, requiring data to be exported or staged in those services first.

73
MCQmedium

You have deployed a model to a Vertex AI Endpoint and need to perform a canary release of a new model version to 10% of traffic. You want to monitor the new version's performance before gradually increasing its traffic share. What should you do?

A.Deploy the new model as a separate endpoint and use a load balancer to split traffic.
B.Deploy the new model to the same endpoint with 100% traffic and monitor closely.
C.Create a new endpoint for the new model and use Vertex AI's traffic director to shift traffic.
D.Upload the new model as a new version and deploy it to the same endpoint with 10% traffic split.
AnswerD

Vertex AI Endpoints support multiple deployed models with configurable traffic splits. By deploying the new version to the same endpoint and assigning 10% of traffic, you can canary test it. You can then monitor metrics and gradually increase the split. This is the native, straightforward approach that integrates with Vertex AI monitoring and rollback.

Why this answer

Vertex AI Endpoints natively support deploying multiple model versions and splitting traffic by percentage. Deploying the new version to the same endpoint with a 10% split enables canary testing, and the split can be adjusted as confidence grows. Separate endpoints with external load balancing or a full cutover do not meet the controlled, gradual rollout requirement.

Exam trap

The trap here is thinking you need a separate endpoint or an external load balancer, when Vertex AI Endpoints already provide built-in traffic splitting.

74
MCQmedium

You are deploying a scikit-learn model to a Vertex AI endpoint for real-time inference. Prediction requests arrive as JSON payloads containing a single instance per request, and the model's predict method expects a pandas DataFrame with named columns. You want to avoid writing a custom container. Which approach should you take?

A.Upload the model artifact with a pre-built Vertex AI XGBoost container, because that container internally wraps every scikit-learn estimator and feeds it a DataFrame.
B.Upload the model artifact with a pre-built Vertex AI scikit-learn container and supply a custom prediction routine (predictor.py) that converts the decoded JSON dict into a pandas DataFrame before calling predict.
C.Upload the model artifact with a pre-built TensorFlow container and wrap the scikit-learn estimator in a tf.function so the container's serving signature handles DataFrame creation.
D.Upload the model artifact with a pre-built Vertex AI scikit-learn container and rely on its default Predictor, which automatically converts JSON objects into a pandas DataFrame with matching column names.
AnswerB

A custom prediction routine packaged as a Python source distribution lets you subclass the pre-built scikit-learn Predictor and override preprocess to build a DataFrame from the decoded instance. This keeps the managed runtime image while satisfying the model's DataFrame input contract, so no custom container or Dockerfile is needed.

Why this answer

A custom prediction routine gives you a managed pre-built runtime plus a Python package that overrides preprocessing, which is the supported way to adapt request payloads to a model's expected input format. Because the estimator requires named DataFrame columns, the override must build that DataFrame explicitly; the default Predictor performs no such conversion and other pre-built containers cannot load the estimator at all.

Exam trap

The trap here is assuming the pre-built scikit-learn container automatically converts JSON payloads into a named-column DataFrame for every estimator.

75
Multi-Selectmedium

You are optimizing a model for deployment on Vertex AI using NVIDIA Triton Inference Server. Which TWO actions can you take to improve inference performance?

Select 2 answers
A.Increase the number of model replicas to the maximum.
B.Use TensorRT to quantize the model to FP16 or INT8.
C.Disable model caching to reduce memory usage.
D.Enable dynamic batching in Triton to aggregate requests.
E.Use a larger machine type with more vCPUs.
AnswersB, D

TensorRT applies FP16 or INT8 quantisation, reducing weight precision and memory bandwidth while exploiting GPU tensor cores for faster matrix maths. This directly satisfies the stem's inference-performance constraint on Triton, where lower-precision kernels cut latency and raise throughput versus FP32 execution.

Why this answer

Option B is correct because TensorRT can quantize a model to FP16 or INT8, reducing precision and memory bandwidth requirements while leveraging NVIDIA GPU tensor cores, which directly lowers inference latency and increases throughput on Triton. Option D is correct because Triton's dynamic batching aggregates multiple incoming inference requests into a single batch at runtime, improving GPU utilization and throughput without requiring client-side batching. Option A is not appropriate because simply maxing out replicas can waste resources and does not guarantee better performance if the model is not compute-bound or if the machine lacks capacity.

Option C is wrong because disabling model caching forces reloads and increases latency rather than improving inference performance. Option E is not ideal because adding vCPUs does not help GPU-bound inference and may not improve Triton throughput.

Exam trap

Google often tests the misconception that simply adding more replicas or CPU resources will linearly improve inference performance, ignoring the GPU-bound nature of model serving and the importance of batching and precision optimization.

Page 1 of 2 · 102 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Serving and Scaling Models questions.