Courseiva

Generative AI Leader Practice Question: Business Strategies for Generative AI Solutions

A media company is using a generative AI model to create video captions. The model is deployed on Vertex AI with autoscaling. During peak hours, they observe high latency and request timeouts. Which action would most effectively address this issue?

⚠ Common exam trap

Many exam-takers confuse performance optimization (faster inference per request) with capacity planning (ensuring enough concurrent replicas), leading them to choose GPU upgrades or prompt tweaks instead of addressing the autoscaling configuration.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Increase the minimum number of replicas in the autoscaling configuration

Increasing the minimum number of replicas ensures that during peak hours, the model already has a baseline of warm instances ready to handle requests, reducing cold-start latency and preventing timeouts. Autoscaling can take time to spin up new replicas, so a higher minimum replica count directly mitigates the latency spike by pre-provisioning capacity.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Optimize the prompt to reduce output length

    Why it's wrong here

    Shortening prompts reduces per-request token generation, but the timeouts stem from insufficient serving capacity under peak concurrency, not from output length. It is tempting because prompt optimisation genuinely lowers latency and cost per call, making it the right choice when individual responses are slow rather than when the endpoint is saturated.

  • ✗

    Reduce the maximum number of replicas to limit resource usage

    Why it's wrong here

    Capping replicas removes the autoscaler's ability to add serving capacity, worsening queueing and timeouts during peaks. It is tempting as a cost-control measure to prevent runaway scaling, which is the correct action when budget overruns matter, but here the symptom is overload, so reducing replicas directly aggravates it.

  • ✗

    Switch to a GPU-based machine type for faster inference

    Why it's wrong here

    Vertex AI autoscaling already provisions replicas; switching to a GPU machine type changes per-request inference speed but does not add capacity, so timeouts from queueing persist. It is tempting when a single model's inference is genuinely compute-bound, which is the correct fix for slow individual predictions rather than overload.

  • ✓

    Increase the minimum number of replicas in the autoscaling configuration

    Why this is correct

    Autoscaling adds replicas only after load rises, so cold-start latency and timeouts persist during sudden peaks. Raising the minimum replica count keeps capacity warm, directly satisfying the peak-hour latency and timeout constraint by ensuring baseline throughput is always available.

About these practice questions

This Generative AI Leader question is part of Courseiva's 1,008-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.