Courseiva

CCNA Serving and Scaling Models Questions

27 of 102 questions · Page 2/2 · Serving and Scaling Models · Answers revealed

76
MCQhard

Your team serves a model on a Vertex AI endpoint with autoscaling. During a flash sale, traffic jumps from 50 to 900 requests per second within one minute, and many requests time out with 429 responses before new replicas become ready. You want to absorb the burst with the least user-visible impact. What should you do?

A.Set a higher minReplicaCount so the endpoint always keeps enough warm capacity for peak traffic, and combine it with a lower autoscaling metric target to trigger scaling earlier.
B.Increase the endpoint's maxReplicaCount and rely on the autoscaler's default metrics to add capacity as quickly as possible.
C.Enable request logging on the endpoint so Cloud Logging captures the 429 responses and the autoscaler can use log volume as an additional scaling signal.
D.Deploy the model to a second endpoint and split traffic between the two endpoints so each handles roughly half of the incoming requests.
AnswerA

Keeping warm replicas sized for the burst removes the cold-start gap entirely, and lowering the autoscaling metric target makes the autoscaler add replicas before saturation rather than after. Together they absorb the sudden spike immediately while still allowing scale-down when traffic subsides, which is exactly what is needed here.

Why this answer

Bursts that outpace autoscaling reaction time must be met with pre-provisioned warm capacity. Setting minReplicaCount near peak demand guarantees replicas are already serving when the spike lands, and a lower metric target causes additional replicas to be requested earlier in the ramp. Merely changing the maximum, logging, or duplicating endpoints does not remove the delay between demand and available capacity.

Exam trap

The trap here is believing that raising maxReplicaCount makes autoscaling react fast enough to absorb a sudden burst.

77
MCQmedium

You have a TensorFlow model that you want to deploy on edge devices for real-time inference. The model was trained in Vertex AI. You need to convert it to a format suitable for on-device inference. Which approach should you use?

A.Export the model as a serialized TFX pipeline.
B.Export the model to a SavedModel and deploy it using Vertex AI Edge Manager.
C.Convert the model to TensorFlow Lite using the TensorFlow Lite converter.
D.Use Vertex AI Model Optimization to compile the model for edge devices.
AnswerC

TensorFlow Lite is Google's runtime for on-device inference, and the TensorFlow Lite converter transforms a trained TensorFlow SavedModel into the compact .tflite format with optimisations such as quantisation. This suits edge devices with limited compute and memory.

Why this answer

TensorFlow Lite is specifically designed for on-device inference on edge devices, offering optimized performance and reduced model size. The TensorFlow Lite converter transforms a TensorFlow model (e.g., from a SavedModel) into the FlatBuffer format (.tflite), which is lightweight and compatible with mobile and embedded platforms. This directly addresses the requirement for real-time inference on edge devices.

Exam trap

A common misconception is that Vertex AI services like Edge Manager or Model Optimization directly produce a deployable edge format, when in fact TensorFlow Lite conversion is the required final step for on-device inference.

How to eliminate wrong answers

Option A is wrong because a TFX pipeline is a production ML workflow framework for orchestrating training, validation, and deployment, not a model format for on-device inference; serializing it does not produce a deployable edge model. Option B is wrong because Vertex AI Edge Manager is a service for managing and deploying models to edge devices, but it expects models in a compatible format like TensorFlow Lite; exporting to SavedModel alone is insufficient without conversion, and Edge Manager itself does not perform the conversion. Option D is wrong because Vertex AI Model Optimization focuses on techniques like pruning and quantization to improve model efficiency, but it does not compile the model into a format suitable for on-device inference; the output still requires conversion to TensorFlow Lite for edge deployment.

78
MCQeasy

You need to serve multiple models on a single Vertex AI endpoint to reduce costs. How can you achieve this?

A.Use Cloud Run to serve each model separately.
B.Use Vertex AI Prediction with multi-model serving by deploying multiple models to one endpoint with traffic splits.
C.Package all models into a single container and deploy that container.
D.Deploy each model to its own endpoint and use a load balancer.
AnswerB

Multiple models can be deployed to a single endpoint, each receiving a portion of the traffic.

Why this answer

Vertex AI Prediction supports multi-model serving, allowing you to deploy multiple models to a single endpoint and use traffic splits to route a percentage of requests to each model. This reduces costs by sharing underlying infrastructure (e.g., compute resources) across models, rather than provisioning separate endpoints or containers for each model.

Exam trap

The trap here is that candidates often confuse multi-model serving with containerization, assuming that bundling models into a single container (Option C) is equivalent to Vertex AI's native multi-model support, but this ignores the need for traffic splitting and independent model lifecycle management.

How to eliminate wrong answers

Option A is wrong because Cloud Run serves each model as a separate service, which does not consolidate models onto a single endpoint and incurs additional costs for individual scaling and networking. Option C is wrong because packaging all models into a single container violates the principle of model isolation, complicates updates, and does not leverage Vertex AI's native traffic-splitting mechanism for granular control. Option D is wrong because deploying each model to its own endpoint and using a load balancer increases operational overhead and cost, as each endpoint requires separate compute resources, defeating the purpose of cost reduction.

79
MCQeasy

A data scientist wants to deploy a trained TensorFlow model to Vertex AI for online predictions. They need to serve predictions with low latency and want to leverage GPU acceleration. Which machine type should they select when creating the Vertex AI endpoint?

A.n1-standard-4 with 1 NVIDIA Tesla T4
B.n1-standard-4
C.e2-standard-4
D.n1-highmem-8
AnswerA

NVIDIA Tesla T4 GPUs attached to n1-standard-4 instances provide the GPU acceleration the stem demands, while n1-standard-4 supplies sufficient vCPU and memory for low-latency online inference of a TensorFlow model. Vertex AI supports this accelerator pairing directly, satisfying both the GPU and latency constraints when deploying the endpoint.

Why this answer

The n1-standard-4 machine type supports attaching GPUs such as the NVIDIA Tesla T4, which provides GPU acceleration for low-latency online predictions. Vertex AI endpoints require a machine type that allows GPU attachment, and the n1-series is one of the few families that supports GPUs, while the T4 offers a good balance of cost and performance for inference workloads.

Exam trap

The trap here is that candidates may assume any machine type can be paired with a GPU, but only specific series (like n1, n2, g2) support GPU attachment, and the e2 series explicitly does not, leading to a wrong selection if the GPU requirement is overlooked.

How to eliminate wrong answers

Option B is wrong because n1-standard-4 without a GPU does not provide GPU acceleration, so it cannot meet the requirement for low-latency predictions with GPU. Option C is wrong because e2-standard-4 does not support attaching GPUs at all; the e2 series is designed for cost-optimized CPU-only workloads. Option D is wrong because n1-highmem-8, while it can support GPUs, is over-provisioned in memory for typical inference tasks and does not include a GPU by default, so it would not satisfy the explicit need for GPU acceleration unless a GPU is attached, but the option as stated lacks the GPU specification.

80
MCQmedium

Your team has deployed a model to a Vertex AI endpoint and wants to route a small percentage of live traffic to a new model version for evaluation. You need to split traffic at the endpoint level without changing the client application. What should you do?

A.Create two separate endpoints and use a load balancer to distribute traffic between them.
B.Deploy both model versions to the same endpoint and set a traffic split percentage.
C.Deploy the new model version as a separate endpoint and update the client to call both endpoints.
D.Use a Vertex AI batch prediction job to send a percentage of live traffic to the new model.
AnswerB

Vertex AI Endpoints support deploying multiple models to the same endpoint and assigning a traffic split percentage to each deployed model. This allows a gradual rollout to a new version while keeping the client application pointed at a single endpoint URL, which matches the requirement exactly.

Why this answer

Deploying multiple models to a single endpoint and configuring a traffic split is the native Vertex AI mechanism for canary or A/B testing. The endpoint continues to expose one URL, so clients remain unchanged, and the split percentage controls how much live traffic reaches each model version. This enables safe evaluation of the new version with minimal risk.

Exam trap

The trap here is believing that a separate endpoint plus an external load balancer is equivalent to an endpoint traffic split, when the native feature avoids client changes and additional infrastructure.

81
MCQhard

You are deploying a PyTorch model on Vertex AI and want to use NVIDIA Triton Inference Server for optimal performance. You have built a custom container with Triton. Which serving configuration should you use?

A.Deploy the model on GKE with Triton and expose via Istio.
B.Use the prebuilt Vertex AI PyTorch prediction container and set environment variables to enable Triton.
C.Use Vertex AI Model Optimization to automatically convert the model to TensorRT and deploy with built-in server.
D.Upload your Triton container to Container Registry and specify it as the prediction container in Vertex AI Model.
AnswerD

Vertex AI accepts a custom container as the prediction container, so pushing the Triton image to Container Registry and referencing it lets Vertex AI route prediction traffic through Triton's dynamic batching and concurrent model execution, satisfying the requirement to serve PyTorch via Triton.

Why this answer

To deploy a custom Triton Inference Server container on Vertex AI, you upload your container to Artifact Registry (or Container Registry) and specify it as the prediction container when creating the Vertex AI Model resource. Vertex AI supports custom containers for prediction, allowing you to run Triton with your model artifacts. This is the standard approach for custom serving frameworks.

Exam trap

PMLE often tests the misconception that prebuilt Vertex AI containers can be toggled to use Triton via environment variables, when in fact a custom container is required.

How to eliminate wrong answers

Option A is wrong because deploying on GKE with Istio is a manual, self-managed approach that bypasses Vertex AI's managed prediction service; the question asks for the Vertex AI serving configuration. Option B is wrong because the prebuilt Vertex AI PyTorch prediction container does not include Triton, and environment variables cannot enable a server that is not installed. Option C is wrong because Vertex AI Model Optimization converts models to TensorRT and deploys with the built-in server, but it does not use Triton and may not support all PyTorch models.

82
MCQeasy

You deployed a model to a Vertex AI endpoint with minReplicas=0 and maxReplicas=5. After sending prediction requests, you notice the endpoint takes about 30 seconds to respond initially, but subsequent requests are fast. What is the most likely cause?

A.The model is too large for the machine type.
B.Cold start occurs because the endpoint scaled down to zero.
C.The VPC Service Controls are blocking the initial request.
D.The endpoint's autoscaling is misconfigured.
AnswerB

With minReplicas=0, Vertex AI scales the endpoint to zero replicas when idle, so the first request must provision a container and load the model — roughly 30 seconds. Once a replica is warm, subsequent requests hit it directly, which is why later latency drops sharply.

Why this answer

Vertex AI endpoints with minReplicas=0 scale down to zero when idle. The first request after a period of inactivity triggers a cold start, where the endpoint must provision a new VM instance and load the model, causing a ~30-second delay. Subsequent requests are fast because the instance remains warm and handles them without provisioning overhead.

Exam trap

Google often tests the distinction between cold start latency and persistent performance issues, so candidates may mistakenly attribute the initial delay to model size or network misconfiguration instead of recognizing the intentional scaling-to-zero behavior.

How to eliminate wrong answers

Option A is wrong because a model too large for the machine type would cause persistent latency or errors on every request, not just the first one after idle time. Option C is wrong because VPC Service Controls enforce network boundaries and would block all requests consistently, not just the initial one with a 30-second delay. Option D is wrong because the autoscaling configuration (minReplicas=0, maxReplicas=5) is correct for scaling to zero; the observed behavior is the expected cold start, not a misconfiguration.

83
MCQmedium

You are deploying a model to a Vertex AI endpoint that will serve predictions for a mobile application. The application sends a single request per user action and expects a response within 100 ms. The model is small and CPU-bound. You want to minimize cost while meeting the latency requirement. Which endpoint configuration should you choose?

A.Use a CPU-only machine type with minReplicaCount=1 and maxReplicaCount=5, and enable autoscaling based on CPU utilization.
B.Use a CPU-only machine type with minReplicaCount=5 and maxReplicaCount=5 to ensure high availability.
C.Use a CPU-only machine type with minReplicaCount=1 and maxReplicaCount=1.
D.Use a GPU-enabled machine type with minReplicaCount=1 and maxReplicaCount=1.
AnswerA

This configuration uses cost-effective CPU instances and allows the endpoint to scale out when CPU utilization rises, ensuring latency remains low under load. Setting minReplicaCount=1 keeps idle cost low, while maxReplicaCount=5 provides headroom. Autoscaling based on CPU is appropriate for a CPU-bound model. This balances cost and performance effectively.

Why this answer

For a small CPU-bound model with variable traffic, the optimal cost-latency trade-off is a CPU-only machine type with a low minimum replica count and autoscaling enabled. This keeps baseline cost low while allowing the endpoint to add replicas when CPU utilization increases, preserving the 100 ms latency target. GPU instances and fixed high replica counts unnecessarily increase cost without improving latency.

Exam trap

The trap here is assuming that GPUs or a fixed number of replicas are needed for low latency, when a small CPU model can meet the SLO with autoscaling.

84
MCQmedium

You are deploying a large language model on a Vertex AI endpoint. The model is loaded from a Cloud Storage bucket at container startup, which adds 3 minutes to each cold start. You want to reduce cold-start time and ensure predictable latency during scale-out. Which approach should you take?

A.Store the model artifacts in a custom container image and push it to Artifact Registry.
B.Use a larger machine type with more vCPUs and memory for each replica.
C.Set the endpoint's minReplicaCount to a high value so that replicas are always warm.
D.Enable request-response logging on the endpoint to monitor startup latency.
AnswerA

Baking the model into a custom container image pulls the artifacts when the container image is downloaded, which happens as part of standard node provisioning. This reduces the additional model-download step at startup, cutting cold-start time and making scale-out more predictable.

Why this answer

Embedding the model artifacts in a custom container image eliminates the separate download from Cloud Storage during container startup. The image layers are pulled by the node as part of standard container initialization, which is typically faster and more predictable than fetching a large model from a GCS bucket after the container starts.

Exam trap

The trap here is assuming that raising the minimum replica count eliminates cold starts entirely, when it only masks them for steady-state traffic and does nothing for scale-out events.

85
MCQhard

A company uses Vertex AI Matching Engine for a product recommendation system. They need to update the index with new product embeddings every hour, but the index is used for online queries with low latency. Which index update strategy should they use?

A.Use streaming updates to insert new embeddings incrementally
B.Use a hybrid approach with batch for daily full rebuild and streaming for hourly
C.Use batch updates to replace the index every hour
D.Recreate the index from scratch each hour
AnswerA

Streaming updates let Matching Engine insert or delete datapoints incrementally while the index stays queryable, so hourly embedding refreshes avoid the rebuild-and-redeploy cycle that batch updates require. This preserves the low-latency online serving the stem demands, since queries continue against the live index throughout.

Why this answer

Streaming updates in Vertex AI Matching Engine allow incremental insertion of new embeddings into an existing index without rebuilding it. This satisfies the requirement for hourly updates while maintaining low-latency online queries, as the index remains available and consistent during the update process.

Exam trap

Google often tests the misconception that batch updates are required for consistency or that streaming updates cannot handle frequent changes, leading candidates to choose hybrid or batch approaches when incremental streaming is both sufficient and optimal for low-latency online serving.

How to eliminate wrong answers

Option B is wrong because a hybrid approach with batch for daily full rebuild and streaming for hourly adds unnecessary complexity and cost; streaming updates alone suffice for hourly increments without needing a daily rebuild. Option C is wrong because batch updates replace the entire index, causing downtime or increased latency during the rebuild, which violates the low-latency online query requirement. Option D is wrong because recreating the index from scratch each hour is inefficient, time-consuming, and disrupts query availability, making it unsuitable for real-time serving.

86
Multi-Selecthard

A fintech company needs to deploy a TensorFlow model for real-time fraud detection with strict latency SLO (p99 < 100ms). They expect variable traffic with spikes. They also want to minimize cold-start latency. Which two configurations should they use? (Choose 2)

Select 2 answers
A.Set min_replicas = 0 to allow scale-to-zero and save costs.
B.Use a GPU-enabled machine type (e.g., N1 with T4) to accelerate inference.
C.Set min_replicas = 3 to keep a baseline of warm instances.
D.Enable Vertex AI Model Optimization for automatic quantization.
E.Use batch prediction instead of online prediction.
AnswersB, C

GPU acceleration (N1 with T4) cuts inference compute time, directly addressing the p99 < 100ms SLO that CPU inference on a TensorFlow fraud model would likely breach. It does not solve cold starts, so it pairs with warm replicas.

Why this answer

Option B is correct because a GPU-enabled machine type such as an N1 instance with an NVIDIA T4 accelerator provides the parallel compute throughput needed to keep TensorFlow inference within a p99 latency SLO under 100ms, which CPU-only serving often cannot guarantee for larger models. Option C is correct because setting min_replicas = 3 keeps a baseline of warm, already-loaded model instances, eliminating cold-start latency for the initial requests and giving the autoscaler headroom to absorb traffic spikes without waiting for new replicas to initialize. Option A is not appropriate because min_replicas = 0 enables scale-to-zero, which directly reintroduces cold-start latency and violates the strict p99 < 100ms SLO.

Option D is not selected because Vertex AI Model Optimization (quantization) is a model-compression technique that may reduce latency but is not a required configuration for meeting the SLO and can degrade accuracy. Option E is not selected because batch prediction processes data offline in bulk and cannot serve real-time fraud detection requests with sub-100ms latency.

Exam trap

A common misconception is that scale-to-zero (min_replicas = 0) is always cost-effective, but in latency-sensitive real-time inference, it introduces unacceptable cold-start delays, making baseline warm instances (min_replicas > 0) essential.

87
MCQmedium

Your team has deployed a model on Vertex AI endpoints. You need to monitor the prediction latency to ensure it meets a 99th percentile SLO of 500ms. You want to set up an alert if the latency exceeds this threshold. Which metric should you use?

A.The 99th percentile of the `prediction/online/response_latencies` metric.
B.The number of prediction requests that timeout.
C.Average prediction latency from the endpoint's logs.
D.The maximum prediction latency from the endpoint's monitoring dashboard.
AnswerA

The prediction/online/response_latencies metric records server-side latency for online predictions, and its 99th percentile aligns exactly with the stem's 500ms SLO threshold. Alerting on that percentile detects tail latency affecting the slowest 1% of requests, which averages or medians would obscure.

Why this answer

The `prediction/online/response_latencies` metric in Vertex AI provides a distribution of latency values, allowing you to query the 99th percentile directly. This aligns with the SLO requirement to monitor the tail latency, not the average or maximum, ensuring that the worst-case performance for 1% of requests stays under 500ms.

Exam trap

Google Cloud often tests the distinction between tail latency (percentiles) and central tendency (average) or extreme values (maximum), trapping candidates who confuse SLO monitoring with simple failure counts or averages.

How to eliminate wrong answers

Option B is wrong because the number of prediction requests that timeout is a count of failures, not a latency measurement; it does not capture the 99th percentile latency and would miss requests that complete but exceed 500ms. Option C is wrong because average prediction latency can mask high tail latencies; a low average could hide a significant number of requests exceeding 500ms, violating the SLO. Option D is wrong because the maximum prediction latency is a single extreme value, often an outlier due to cold starts or transient spikes, and does not represent the 99th percentile behavior required for the SLO.

88
MCQhard

You are serving a model on a Vertex AI endpoint that requires a GPU. The model is used for interactive predictions with a strict latency SLO. You notice that during peak hours, some requests time out because the endpoint's autoscaler is slow to add GPU replicas. Which action should you take to meet the SLO?

A.Enable request batching on the endpoint to process multiple requests per replica.
B.Set a higher minimum replica count on the endpoint to keep more GPU replicas warm during peak hours.
C.Reduce the model's input size by truncating features to lower per-request compute.
D.Switch to a CPU-based machine type to avoid GPU provisioning delays.
AnswerB

Increasing the minimum replica count ensures that more GPU replicas are already running and ready to serve traffic before the peak arrives. Since GPU replicas take longer to start, pre-warming them avoids the delay caused by slow autoscaling. This directly addresses the timeout issue by providing sufficient capacity during peak hours without waiting for the autoscaler to react, helping to meet the strict latency SLO.

Why this answer

GPU replicas take longer to provision and initialize than CPU replicas, so autoscaling may not react quickly enough during sudden peaks. Setting a higher minimum replica count keeps additional GPU replicas warm and ready before demand spikes. This pre-provisioned capacity absorbs peak traffic immediately, preventing request queuing and timeouts.

It addresses the slow autoscaler response directly, ensuring the endpoint can meet its strict latency SLO during peak hours.

Exam trap

The trap here is assuming that batching or reducing input size solves timeout issues, when the real cause is insufficient warm GPU capacity during peak demand.

89
Multi-Selectmedium

A media company is serving a video recommendation model on a Vertex AI Endpoint. The model receives a mix of requests: some require only a few features, while others require many features from a feature store. The team wants to reduce average latency and cost without retraining the model. Which TWO strategies should they use? (Choose two.)

Select 2 answers
A.Switch from online predictions to batch predictions for all requests.
B.Enable request batching on the endpoint to group multiple requests into a single prediction call.
C.Use a smaller machine type for the endpoint to reduce cost per replica.
D.Implement feature caching in the model container to avoid repeated feature store lookups for the same entities.
E.Increase the endpoint's maxReplicaCount to handle peak traffic.
AnswersB, D

Batching multiple requests into one prediction call can improve throughput and reduce per-request overhead, especially when the model can process a batch efficiently. This reduces average latency under load and can lower cost by using resources more effectively. It does not require retraining and is a serving-side optimization.

Why this answer

Request batching and feature caching are serving-side optimizations that reduce per-request overhead and repeated feature lookups. They lower average latency and can reduce cost by improving resource utilization, all without retraining the model. Other options either do not address the root causes or are unsuitable for real-time serving.

Exam trap

The trap here is assuming that scaling out with more replicas reduces average latency, when it primarily helps with peak load and may increase cost.

90
MCQhard

A fraud-detection model is deployed on a Vertex AI endpoint and must respond within 30 ms for 95% of requests. During testing, the team sees that p95 latency is dominated by feature retrieval from an external online store, not by model inference. They want to reduce latency without retraining the model. What should they do first?

A.Enable request-response logging on the endpoint to identify which feature lookups are slow, then retrain the model with fewer features.
B.Switch the endpoint to batch prediction mode so requests are grouped and feature lookups are amortized.
C.Increase the endpoint's maxReplicaCount so more replicas can handle feature lookups in parallel.
D.Co-locate the feature values with the model by using a low-latency in-memory feature store or embedding the features in the request path.
AnswerD

When feature retrieval dominates p95 latency, the most direct fix is to remove or shorten that network hop. Serving features from an in-memory store or including them in the request payload eliminates the external round trip and keeps the model's inference time as the only major component. This aligns with the no-retraining constraint.

Why this answer

The p95 latency is dominated by an external online feature store, so the serving architecture, not the model or replica count, is the limiting factor. Bringing features into the model's process via an in-memory store or passing them in the request removes the network round trip and directly attacks the dominant latency component.

Exam trap

The trap here is defaulting to horizontal scaling whenever latency is high, when the real issue is per-request external dependency latency that replication cannot fix.

91
MCQmedium

You are using Vertex AI Matching Engine for similarity search. Your index has 10 million embeddings of 512 dimensions. The query latency requirement is under 10ms for 99th percentile. Which index type should you choose?

A.Brute-force index with cosine distance.
B.Approximate Nearest Neighbor (ANN) index using the ScaNN algorithm.
C.A custom distance-based index using Cloud SQL.
D.A tree-based index from scikit-learn deployed as a custom container.
AnswerB

ScaNN's approximate nearest neighbour index partitions and quantises vectors, avoiding exhaustive distance computation across all 10 million embeddings. This keeps 99th-percentile query latency under 10ms, which exact brute-force search could not achieve at this scale.

Why this answer

The ScaNN (Scalable Nearest Neighbors) algorithm is specifically designed for high-dimensional, large-scale similarity search with strict latency requirements. With 10 million 512-dimensional embeddings, an ANN index like ScaNN can achieve sub-10ms query latency at the 99th percentile by trading a small amount of recall for dramatic speed improvements, which is exactly what Vertex AI Matching Engine optimizes for.

Exam trap

The trap here is that candidates assume brute-force is the only 'accurate' option and underestimate how severely the curse of dimensionality degrades tree-based and exact methods at 512 dimensions, leading them to pick A or D despite the explicit latency constraint.

How to eliminate wrong answers

Option A is wrong because a brute-force index computes exact distances against all 10 million embeddings, which for 512-dimensional vectors would require O(10M * 512) operations per query, far exceeding the 10ms latency target even with optimized hardware. Option C is wrong because Cloud SQL is a relational database not designed for vector similarity search; it lacks native support for high-dimensional distance computations and would require full table scans, making sub-10ms latency impossible at this scale. Option D is wrong because scikit-learn's tree-based indices (e.g., KD-Tree, Ball Tree) degrade to near-linear search in high dimensions (curse of dimensionality), performing no better than brute force for 512 dimensions, and deploying as a custom container adds unnecessary overhead without addressing the fundamental algorithmic limitation.

92
MCQmedium

You are using Vertex AI Vector Search with an approximate nearest neighbor index. You need to update the index with new data every hour. The updates must be available for queries immediately. Which update method should you use?

A.Recreate the index every hour using a scheduled job.
B.Batch update by creating a new index and deploying it.
C.Streaming updates using the streaming API.
D.Use a brute-force index that supports real-time updates.
AnswerC

Streaming updates via the streaming API mutate the deployed index in place, so newly added or removed datapoints become queryable within seconds. This satisfies the hourly refresh with immediate query visibility, unlike batch rebuilds that require reindexing and redeployment.

Why this answer

Vertex AI Vector Search supports streaming updates via its streaming API, which allows you to insert, update, or delete vectors in real time. This ensures that new data is immediately available for approximate nearest neighbor (ANN) queries without requiring index recreation or redeployment, meeting the requirement for hourly updates with instant query availability.

Exam trap

A common misconception is that updating an approximate nearest neighbor (ANN) index requires recreating and redeploying it. However, Vertex AI Vector Search provides a streaming API that enables real-time insert, update, and delete operations, making new data immediately available for queries without the overhead of full index rebuilds.

How to eliminate wrong answers

Option A is wrong because recreating the entire index every hour is inefficient and introduces downtime during the rebuild process, failing the requirement for immediate query availability. Option B is wrong because batch updating by creating a new index and deploying it involves a delay for building and deploying the index, so updates are not available immediately for queries. Option D is wrong because Vertex AI Vector Search does not offer a brute-force index that supports real-time updates; brute-force indices are typically used for exact nearest neighbor search and are not designed for real-time streaming updates in this service.

93
MCQhard

You are using Vertex AI Vector Search for a product recommendation system. Your index is updated with new embeddings every hour. To minimize query latency while keeping the index fresh, what should you do?

A.Use streaming updates to insert new embeddings into the deployed index.
B.Rebuild the entire index hourly as a batch job and redeploy it.
C.Create a new index each hour and use traffic splitting to gradually shift traffic.
D.Use a brute-force index instead of ANN to ensure accuracy after updates.
AnswerA

Streaming updates let Vector Search mutate the deployed index in place as new embeddings arrive, so hourly refreshes avoid the rebuild-and-redeploy cycle that batch updates require. This keeps the index fresh without taking it offline, directly satisfying the low-latency, hourly-freshness constraint in the stem.

Why this answer

Vertex AI Vector Search supports streaming updates that let you insert, update, or delete datapoints in a deployed index without rebuilding it, so queries stay low-latency while the index remains fresh. This is the intended mechanism for near-real-time freshness with minimal disruption. Batch rebuilds and redeployments cause downtime and latency spikes, which the question explicitly asks to avoid.

Exam trap

PMLE often tests the trade-off between index freshness and query latency, and candidates wrongly assume a full rebuild is required for updates, missing the streaming update capability.

How to eliminate wrong answers

Option B is wrong because rebuilding the entire index hourly and redeploying is expensive, slow, and causes query unavailability or latency spikes during the swap, defeating the goal of minimizing latency. Option C is wrong because creating a new index hourly with traffic splitting adds operational overhead and still requires full index builds, and traffic splitting is for model/index rollout, not hourly freshness. Option D is wrong because brute-force (exact) search scales poorly and dramatically increases query latency, the opposite of the requirement.

94
MCQmedium

You need to run batch predictions on a large dataset stored in BigQuery using a Vertex AI model. The dataset contains 10 million rows, and each prediction takes about 100ms. You want to minimize cost and execution time. What should you do?

A.Export the BigQuery data to CSV in GCS, then run a custom Dataflow pipeline to make predictions.
B.Use Vertex AI batch prediction with BigQuery as the source and sink.
C.Use Vertex AI online prediction and send all rows as separate requests.
D.Use a custom container running on Google Kubernetes Engine to perform inference.
AnswerB

Vertex AI batch prediction reads directly from BigQuery and writes results back there, avoiding data export and manual sharding. This satisfies the stem's cost and execution-time constraints by parallelising the 10 million rows across managed workers rather than issuing row-by-row online requests.

Why this answer

Vertex AI batch prediction natively supports BigQuery as both input and output, eliminating the need for data export or custom pipelines. For 10 million rows at 100ms each, batch prediction processes them in parallel across multiple machines, minimizing execution time while avoiding the per-node costs of online prediction or the overhead of managing Dataflow or GKE clusters.

Exam trap

Google Cloud exams often test the distinction between batch and online prediction, trapping candidates who overlook that batch prediction is purpose-built for large-scale, offline inference with native BigQuery integration, while online prediction is for real-time, low-latency use cases.

How to eliminate wrong answers

Option A is wrong because exporting to CSV and using Dataflow adds unnecessary complexity and cost; Vertex AI batch prediction can read directly from BigQuery, avoiding data movement and extra processing steps. Option C is wrong because online prediction is designed for low-latency, real-time requests on small payloads, and sending 10 million separate requests would be prohibitively expensive and slow due to per-request pricing and network overhead. Option D is wrong because running a custom container on GKE requires you to manage infrastructure, scaling, and fault tolerance, which is more costly and complex than using Vertex AI's managed batch prediction service.

95
MCQhard

A financial services firm serves a fraud-detection model on a Vertex AI endpoint that consumes features from a Vertex AI Feature Store online store. During a load test, prediction latency is acceptable, but the firm discovers that the model's feature values in production drift from the values used at training time because the training pipeline read from a BigQuery table with different transformation logic. The team wants the serving path to use the same feature definitions as training so online and offline values match. Which approach should they take?

A.Log all online feature values and retrain the model weekly on the logged production data.
B.Increase the online store's node count and shorten the feature value TTL so fresher values are served.
C.Route serving requests through the offline store and cache results in Memorystore for low latency.
D.Define the features once as Feature Store feature views with a shared transformation and serve the model from the online store while the training pipeline reads the same definitions from the offline store.
AnswerD

Feature Store provides a single definition for each feature, materialized to the online store for low-latency serving and to the offline store for training. When both paths derive from the same feature view and transformation, online and offline values are computed identically, which is the definition of training-serving skew prevention. This directly removes the divergent BigQuery logic that caused the drift.

Why this answer

Training-serving skew from divergent transformation logic is solved by unifying feature definitions, not by tuning storage or adding caching. Vertex AI Feature Store lets a feature view define a transformation once, then materialize the result to the online store for low-latency serving and to the offline store for training. Because both paths derive from the same definition, the values the model sees at inference match the values it was trained on, eliminating the drift at its source.

Exam trap

The trap here is treating a semantic mismatch between two transformation implementations as a latency or freshness problem, and reaching for storage scaling or caching instead of unifying the feature definition.

96
MCQeasy

Which API is recommended for high-throughput, low-latency online prediction requests to Vertex AI endpoints?

A.Cloud Functions
B.REST API
C.Cloud Pub/Sub
D.gRPC API
AnswerD

gRPC API uses HTTP/2 multiplexing and binary Protocol Buffers, cutting serialisation overhead and connection latency compared with REST's JSON over HTTP/1.1. This directly satisfies the stem's high-throughput, low-latency constraint for online prediction requests to Vertex AI endpoints.

Why this answer

gRPC API is recommended for high-throughput, low-latency online prediction requests to Vertex AI endpoints because it uses HTTP/2 for multiplexed streaming, binary serialization (Protocol Buffers), and supports bidirectional streaming, which reduces latency and improves throughput compared to REST. Vertex AI's prediction service natively supports gRPC for real-time inference, making it the optimal choice for latency-sensitive applications.

Exam trap

Google often tests the misconception that REST API is the default or only way to interact with cloud services, but the trap here is that for high-throughput, low-latency online predictions, gRPC is explicitly recommended over REST due to its performance advantages with Protocol Buffers and HTTP/2.

How to eliminate wrong answers

Option A is wrong because Cloud Functions is a serverless compute service for event-driven code, not an API for making prediction requests; it can invoke Vertex AI endpoints via REST or gRPC but is not itself an API protocol. Option B is wrong because REST API uses HTTP/1.1 with JSON serialization, which introduces higher latency and larger payload sizes compared to gRPC's binary Protocol Buffers, making it suboptimal for high-throughput, low-latency scenarios. Option C is wrong because Cloud Pub/Sub is a message queue for asynchronous, decoupled messaging, not designed for synchronous, low-latency online predictions; it adds queuing delay and is intended for batch or event-driven workflows.

97
MCQmedium

You are deploying a model to a Vertex AI Endpoint that will serve predictions to a global user base. You want to minimize latency for users in different regions while ensuring high availability. What should you do?

A.Deploy the model to multiple endpoints in the same region and use a global load balancer.
B.Deploy the model to a single endpoint in one region and use a global load balancer to route traffic.
C.Deploy the model to a single endpoint with a larger machine type to handle global traffic.
D.Deploy the model to multiple endpoints in different regions and use a global load balancer with latency-based routing.
AnswerD

Deploying the model to multiple regional endpoints brings the model closer to users, reducing network latency. A global load balancer with latency-based routing directs each user to the closest healthy endpoint, minimizing latency and providing high availability by failing over if a region becomes unavailable.

Why this answer

To minimize latency for a global user base, the model should be deployed in multiple regions so that users can be served from a nearby location. A global load balancer with latency-based routing ensures that each user is directed to the closest healthy endpoint, reducing latency and providing failover for high availability. Single-region deployments, even with larger machines or multiple endpoints in one region, cannot overcome geographic latency.

Exam trap

The trap here is focusing on capacity or redundancy within a single region instead of geographic distribution to reduce latency for a global user base.

98
MCQmedium

An organization wants to deploy a TensorFlow model on edge devices such as smartphones and IoT devices for offline inference. Which format should they export the model to?

A.ONNX format
B.TensorFlow Lite (TFLite)
C.SavedModel format
D.HDF5 format
AnswerB

TensorFlow Lite converts models to a compact FlatBuffer format with quantisation, cutting size and latency for constrained hardware. It provides the lightweight runtime that executes inference entirely on-device, satisfying the offline requirement for smartphones and IoT devices without server connectivity.

Why this answer

TensorFlow Lite (TFLite) is Google's lightweight runtime and model format specifically designed for on-device inference on mobile and embedded/IoT hardware. It produces a compact .tflite flatbuffer optimized for low latency, small binary size, and minimal memory, and supports hardware acceleration via delegates (GPU, NNAPI, Edge TPU). Exporting to TFLite is the standard path for offline inference on smartphones and IoT devices.

Exam trap

PMLE often tests the confusion between SavedModel (cloud/TF Serving) and TFLite (edge/mobile), so candidates who equate 'TensorFlow model export' with SavedModel pick the wrong format for constrained devices.

How to eliminate wrong answers

Option A is wrong because ONNX is an open interchange format primarily used for cross-framework portability (e.g., PyTorch to TensorRT), not the native optimized format for TensorFlow edge deployment. Option C is wrong because SavedModel is TensorFlow's full-fidelity serving format for TF Serving and cloud endpoints — it is too heavy for constrained edge devices. Option D is wrong because HDF5 (.h5) is the legacy Keras weight/architecture serialization format used for training checkpoints, not for optimized edge inference.

99
MCQmedium

You need to query a Vertex AI Vector Search index for nearest neighbours. The index is deployed on an endpoint. Which API method should you use to perform the query?

A.projects.locations.indexEndpoints.findNeighbors
B.projects.locations.indexes.match
C.projects.locations.indexes.query
D.projects.locations.endpoints.predict
AnswerA

findNeighbors is the IndexEndpoint method that queries a deployed index for nearest neighbours, accepting the query vector and returning neighbour IDs with distances. It is the correct API surface for querying an index already deployed to an endpoint.

Why this answer

The correct API method to query a deployed Vertex AI Vector Search index for nearest neighbors is `projects.locations.indexEndpoints.findNeighbors`. This method is specifically designed for vector similarity search against an index endpoint, returning the nearest neighbors for a given query vector. The other options either target the wrong resource (indexes instead of indexEndpoints) or use methods intended for different purposes like model prediction.

Exam trap

The exam often tests the distinction between model prediction endpoints and vector search endpoints, so the trap here is confusing the `predict` method (for model inference) with the `findNeighbors` method (for vector similarity search), leading candidates to incorrectly select option D.

How to eliminate wrong answers

Option B is wrong because `projects.locations.indexes.match` is not a valid API method; the correct method for matching against an index is `findNeighbors` on the index endpoint. Option C is wrong because `projects.locations.indexes.query` does not exist; the query operation for vector search is performed via the index endpoint, not directly on the index resource. Option D is wrong because `projects.locations.endpoints.predict` is used for online prediction from a deployed model, not for querying a vector search index.

100
MCQhard

A Vertex AI Endpoint hosts a model that must serve predictions with a strict 99th percentile latency under 100 ms. The model is a large TensorFlow model that processes images. During load testing, you observe that p99 latency spikes to 300 ms when batch size exceeds 1. You need to meet the latency SLO while maintaining reasonable throughput. What should you do?

A.Enable dynamic batching with a small maximum batch size and a short timeout.
B.Increase the batch size further to amortize overhead and improve throughput.
C.Increase the number of replicas and keep batch size at 1.
D.Switch to a CPU-only machine type with more cores to parallelize inference.
AnswerA

Dynamic batching groups multiple requests into a single inference call, improving throughput, while a small maximum batch size and short timeout bound the added latency. This keeps p99 under the SLO by preventing large batches that caused the 300 ms spike, and it balances throughput and latency better than fixed batch size of 1. It directly addresses the observed trade-off.

Why this answer

The latency spike is caused by large batch sizes. Dynamic batching with a small maximum and short timeout limits how long requests wait and how many are combined, keeping p99 under 100 ms while still batching enough to improve throughput. Simply adding replicas or changing hardware does not resolve the batch-size-induced tail latency.

Exam trap

The trap here is assuming that more replicas or a bigger batch size will fix latency, when the spike is specifically caused by large batches.

101
MCQmedium

You have a custom model deployed on a Vertex AI endpoint that receives online prediction requests. The model expects input features in a specific order, but clients sometimes send features in a different order. You want to ensure that the endpoint consistently receives correctly ordered features without modifying every client. What should you do?

A.Create a new endpoint for each client that expects the client's specific feature order.
B.Implement a preprocessing function in the model's serving signature that reorders the input features before prediction.
C.Configure the endpoint to reject requests with incorrectly ordered features and return an error message to clients.
D.Use a Vertex AI Feature Store to serve features to the model, ensuring that features are always in the correct order.
AnswerB

Adding a preprocessing step in the serving signature allows the endpoint to accept raw input and reorder features internally. This centralizes the logic, so clients don't need to change. Vertex AI custom containers can include a preprocessing function as part of the model's serving logic, ensuring consistent feature order regardless of client input. This is a robust solution that scales with the endpoint.

Why this answer

The most effective solution is to incorporate feature reordering into the model's serving logic. By adding a preprocessing function in the serving signature, the endpoint can accept features in any order and rearrange them correctly before inference. This centralizes the logic, avoids client-side changes, and ensures consistent predictions.

It is a standard practice in Vertex AI custom containers to include such preprocessing.

Exam trap

The trap here is thinking that Feature Store or multiple endpoints will solve the ordering problem, when the simplest fix is to handle reordering within the model's serving code.

102
MCQeasy

A healthcare analytics team needs to serve a model on Vertex AI to internal applications, but compliance requires that no prediction request or response payload ever be written to logs. They still want basic operational metrics such as request count and latency. What should they configure on the endpoint?

A.Delete the endpoint's service account permissions on Cloud Logging so the platform cannot write any log entries, including metrics.
B.Disable request-response logging on the deployed model while leaving endpoint monitoring and Cloud Monitoring metrics enabled.
C.Set the endpoint's traffic split to route all requests to a canary deployment that has logging disabled, while the primary deployment keeps logging enabled.
D.Enable request-response logging but set the sampling rate to a very small value so almost no payloads are captured.
AnswerB

Request-response logging is a separate, opt-in setting on the deployed model. Leaving it disabled means payloads are never captured, while the endpoint still emits standard metrics like request count, error rate, and latency percentiles to Cloud Monitoring, satisfying both the compliance constraint and the observability requirement.

Why this answer

Payload logging in Vertex AI is controlled by an explicit setting on each deployed model and is off unless enabled. Disabling it prevents request and response content from reaching Cloud Logging, while the endpoint continues to publish aggregate metrics such as request counts, error rates, and latency to Cloud Monitoring, which is exactly the separation this compliance scenario requires.

Exam trap

The trap here is conflating operational metrics with payload logging and assuming both must be turned off together to protect sensitive data.

← PreviousPage 2 of 2 · 102 questions total

Ready to test yourself?

Try a timed practice session using only Serving and Scaling Models questions.