Courseiva

PMLE · topic practice

Serving and Scaling Models practice questions

This domain covers deploying, optimizing, and operating models on Vertex AI: endpoints, custom containers, batch vs online prediction, autoscaling, GPUs/TPUs, and Vector Search indexes. Questions are scenario-based, asking you to choose the configuration, index type, or rollout strategy that meets latency, throughput, cost, and freshness constraints.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Serving and Scaling Models

What the exam tests

What to know about Serving and Scaling Models

Be able to deploy a model to a Vertex AI Endpoint, split traffic for canary rollout, pick the right Vector Search index for latency versus recall, and cut cold starts. The single most important thing: match the serving configuration to the stated latency, throughput, and freshness requirement.

Configuring Vertex AI Endpoints with traffic splits, autoscaling, and deployed model settings

Choosing Vector Search index types (tree-AH, brute force) and StreamUpdate for freshness

Reducing cold start with smaller images, model warming, and min replica counts

Serving custom containers via Artifact Registry and Vertex AI Prediction with GPUs

Watch out for

Common Serving and Scaling Models exam traps

  • ▸Assuming streaming index updates are free; they trade latency and cost for freshness versus batch rebuilds.
  • ▸Confusing traffic split (percentage routing) with model version deployment; both are needed for canary rollout.
  • ▸Ignoring min replica count and image size when diagnosing cold start latency on custom containers.

Practice set

Serving and Scaling Models questions

20 questions · select your answer, then reveal the explanation

You are deploying a new version of a model to a Vertex AI endpoint that already has a champion model serving 100% of traffic. You want to gradually shift traffic to the new version while monitoring for errors. Which approach should you use?

A company is using Vertex AI Prediction with a custom container that performs preprocessing before inference. The preprocessing step is CPU-intensive and the inference step uses a GPU. They want to minimize prediction latency while optimizing cost. Which architecture should they use?

You have a Vertex AI endpoint with min_replica_count=2 and max_replica_count=10. You notice that during a traffic spike, the endpoint does not scale up quickly enough, causing increased latency. What should you do to improve autoscaling responsiveness?

An ML team wants to deploy multiple models (e.g., a recommender and a classifier) behind a single Vertex AI endpoint. The models have different resource requirements: the recommender needs GPU, the classifier needs high memory. How should they configure the endpoint?

Question 5mediummultiple choice
Study the full Python automation breakdown →

You need to run a batch prediction job on Vertex AI using a model that requires custom preprocessing using a Python script. The preprocessing must be applied before inference. Which approach should you use?

An organization wants to deploy a model on edge devices (e.g., Android phones) for offline inference. They trained a model using TensorFlow. Which THREE steps should they take to prepare and deploy the model?

You want to deploy a TensorFlow model to a Vertex AI endpoint and enable online predictions. The model requires GPU for inference. Which machine type should you select when deploying the model?

Your Vertex AI endpoint is experiencing high latency during traffic spikes. You have set maxReplicas=10 and minReplicas=2. The CPU utilisation target is 60%. During spikes, the endpoint never scales beyond 4 replicas. What is the most likely reason?

Your team has built a low-latency similarity search service using Vertex AI Matching Engine (Vector Search). The index is updated daily with new embeddings. You need to serve the latest index without downtime. What is the correct deployment strategy?

You need to serve a model on an edge device with low latency and offline capability. Which approach should you use?

You have a Vertex AI endpoint with two deployed models: model A (champion) and model B (challenger). Traffic split is 90:10. You want to gradually increase model B's traffic to 50% over a week. What is the best way to update the traffic split?

You are deploying a model for real-time inference with strict latency requirements (<100ms P99). You want to autoscale based on custom metrics. Which TWO actions should you take? (Choose 2)

Your team is using Vertex AI Prediction for a large-scale NLP model (PyTorch, custom ops). The model currently runs on CPU but you want to optimise inference cost and performance. Which THREE approaches should you consider? (Choose 3)

You need to deploy a model for online predictions with low latency. You want to ensure that the endpoint can handle traffic bursts without cold start. Which TWO configurations should you set? (Choose 2)

A company deploys a model on Vertex AI Endpoints for real-time inference. They need to minimize latency for prediction requests that are identical to previous requests. Which approach should they use?

An ML engineer is optimizing a large model for deployment on Vertex AI with GPU acceleration. They want to reduce model size and improve inference latency without significant accuracy loss. Which tool should they use?

A company is deploying multiple models on a single Vertex AI endpoint to reduce costs. Each model has different traffic patterns. Which configuration should they use?

A company wants to use Vertex AI Vector Search for real-time product recommendations based on user embeddings. They need to update the index frequently with new product embeddings without significant downtime. Which TWO options should they consider? (Choose 2)

You have a Vertex AI endpoint serving a model with min replicas=2 and max replicas=10. You notice that during low traffic hours, the endpoint still runs 2 replicas, incurring costs. You want to reduce costs to zero when there is no traffic. What should you do?

You have a Vertex AI endpoint with autoscaling enabled. You notice that during traffic spikes, the endpoint takes a long time to scale up, causing prediction errors. What is the most effective solution?

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Serving and Scaling Models sessions

Start a Serving and Scaling Models only practice session

Every question in these sessions is drawn from the Serving and Scaling Models domain — nothing else.

Related practice questions

Related PMLE topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the PMLE exam test about Serving and Scaling Models?
Be able to deploy a model to a Vertex AI Endpoint, split traffic for canary rollout, pick the right Vector Search index for latency versus recall, and cut cold starts. The single most important thing: match the serving configuration to the stated latency, throughput, and freshness requirement.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Serving and Scaling Models questions in a focused session?
Yes — the session launcher on this page draws every question from the Serving and Scaling Models domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other PMLE topics?
Use the topic links above to move to related areas, or go back to the PMLE question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the PMLE exam covers. They are not copied from any real exam or dump site.