20+ practice questions focused on Serving and Scaling Models — one of the most tested topics on the Google Professional Machine Learning Engineer exam. Each question includes a detailed explanation so you learn why the right answer is correct.
Start Serving and Scaling Models PracticeYou are deploying a new version of a model to a Vertex AI endpoint that already has a champion model serving 100% of traffic. You want to gradually shift traffic to the new version while monitoring for errors. Which approach should you use?
Explanation: Vertex AI endpoints support traffic splitting between model versions deployed to the same endpoint. By deploying the challenger to the same endpoint and setting an initial split (e.g., champion 90%, challenger 10%), you can gradually shift traffic while monitoring for errors. This approach uses the endpoint's built-in traffic management, avoiding the complexity and latency of external load balancers.
A company is using Vertex AI Prediction with a custom container that performs preprocessing before inference. The preprocessing step is CPU-intensive and the inference step uses a GPU. They want to minimize prediction latency while optimizing cost. Which architecture should they use?
Explanation: Running preprocessing and inference on the same GPU-backed instance (e.g., n1-standard-4 with T4) eliminates the network hop, serialization, and queueing overhead that would otherwise be introduced by splitting the pipeline across services. Because the preprocessing is CPU-intensive and the inference is GPU-bound, co-locating them lets the CPU work overlap with GPU execution on the same VM, minimizing end-to-end latency. It also avoids paying for two separate always-on resources, which is the cost-optimal choice for a tightly coupled preprocess-then-infer pipeline.
You have a Vertex AI endpoint with min_replica_count=2 and max_replica_count=10. You notice that during a traffic spike, the endpoint does not scale up quickly enough, causing increased latency. What should you do to improve autoscaling responsiveness?
Explanation: Reducing the target CPU utilization percentage (e.g., from the default 60% to a lower value like 40%) causes the autoscaler to trigger scale-up actions sooner, as the threshold for adding replicas is reached at a lower CPU load. This improves responsiveness during traffic spikes by initiating scaling earlier, reducing latency. The endpoint's min_replica_count=2 and max_replica_count=10 remain unchanged, so the scaling range is preserved.
An ML team wants to deploy multiple models (e.g., a recommender and a classifier) behind a single Vertex AI endpoint. The models have different resource requirements: the recommender needs GPU, the classifier needs high memory. How should they configure the endpoint?
Explanation: Vertex AI endpoints support deploying multiple models to the same endpoint, and each deployed model can have its own dedicated resources (machine type, accelerators, min/max replicas). This allows the recommender to use a GPU-backed machine while the classifier uses a high-memory CPU machine, all behind a single endpoint URL. Traffic is split or routed to the appropriate model based on the deployed model ID or traffic split configuration. This is the intended design for heterogeneous model serving.
You need to run a batch prediction job on Vertex AI using a model that requires custom preprocessing using a Python script. The preprocessing must be applied before inference. Which approach should you use?
Explanation: Dataflow (Apache Beam) is the recommended serverless service for distributed, scalable preprocessing of large datasets on Google Cloud. It can read raw data from GCS, apply custom Python preprocessing logic, and write the preprocessed results to GCS or BigQuery. The batch prediction job then reads the preprocessed data directly, avoiding the need to embed preprocessing in the prediction container or handle data on a single VM.
+15 more Serving and Scaling Models questions available
Practice all Serving and Scaling Models questions1. Baseline your knowledge
Start with 10 questions to gauge your current understanding of Serving and Scaling Models. This tells you whether you need a concept refresher or just practice.
2. Review every explanation
For each question — right or wrong — read the full explanation. Understanding why an answer is correct is more valuable than knowing the answer itself.
3. Focus on exam traps
Serving and Scaling Models questions on the PMLE frequently use trap wording. Look for subtle differences in answers that test your precision, not just general knowledge.
4. Reach 80% consistently
Do repeated sessions until you score 80%+ three times in a row. Then move to mixed-mode practice to test cross-topic recall under realistic conditions.
The exact number varies per candidate. Serving and Scaling Models is tested as part of the Google Professional Machine Learning Engineer blueprint. Practicing with targeted Serving and Scaling Models questions ensures you can handle any format or difficulty that appears.
Yes. Courseiva provides free PMLE practice questions across all exam topics and domains. The platform includes topic-based practice, mock exams, missed-question review, bookmarked questions, and readiness tracking — no account required.
Difficulty is subjective, but Serving and Scaling Models is a high-priority exam concept tested in multiple ways — direct recall, scenario analysis, and command-output interpretation. Consistent practice is the best way to build confidence.
Launch a full Serving and Scaling Models practice session with instant scoring and detailed explanations.
Start Serving and Scaling Models Practice →