Courseiva

Generative AI Leader Fundamentals of Generative AI Practice Question

You are an ML engineer at a retail company. You have deployed a generative AI model on Vertex AI to generate product descriptions. The model uses a custom container and is deployed to a single endpoint. Recently, you noticed that inference latency has increased significantly during peak hours, causing timeouts. You have checked the logs and found that the CPU utilization on the deployed instances is consistently above 90% during peak hours. The model is currently deployed with a single machine type (n1-standard-4) and no scaling. You need to reduce latency without incurring excessive cost. What should you do?

⚠ Common exam trap

Candidates often assume adding a GPU (Option D) is always the best way to reduce inference latency, but for CPU-bound models with high utilization, scaling out with more replicas and a larger CPU machine is more cost-effective and directly addresses the bottleneck.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Change the machine type to n1-standard-8 and enable autoscaling with min replicas=1, max replicas=5

Upgrading to a larger machine (n1-standard-8) provides more CPU cores to handle the increased inference workload, while enabling autoscaling (min=1, max=5) allows the deployment to dynamically add replicas during peak hours to distribute the load and reduce latency. This combination addresses the high CPU utilization without over-provisioning during off-peak times, thus controlling cost.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Optimize the model using quantization and reduce the number of replicas

    Why it's wrong here

    Quantization shrinks model size and memory footprint but does not add compute capacity, and reducing replicas lowers throughput, worsening CPU saturation and timeouts. It is tempting as a cost-saving measure, and quantization would be right when memory or model size, not CPU load, is the bottleneck.

  • ✗

    Switch to batch prediction instead of online prediction

    Why it's wrong here

    Batch prediction processes queued jobs against stored data, returning results asynchronously; it cannot serve synchronous per-request product descriptions, so peak-hour timeouts persist. It is tempting because batch prediction cuts cost for large offline scoring runs, and would be right if latency-tolerant bulk generation were acceptable.

  • ✓

    Change the machine type to n1-standard-8 and enable autoscaling with min replicas=1, max replicas=5

    Why this is correct

    CPU is saturated above 90%, so the bottleneck is compute per replica. Doubling to n1-standard-8 gives each replica more vCPU headroom, and autoscaling to five replicas absorbs peak-hour concurrency while min replicas=1 keeps idle cost low.

  • ✗

    Add a GPU accelerator to the existing machine

    Why it's wrong here

    The bottleneck is CPU saturation from a custom container, so adding a GPU accelerator leaves the CPU-bound preprocessing path unchanged and raises cost. GPUs suit models whose matrix operations dominate and whose framework is CUDA-enabled; here autoscaling the n1-standard-4 replicas addresses the actual constraint.

About these practice questions

Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.