Courseiva
hardMultiple Select

Generative AI Leader Practice Question: Deploying a generative AI application using…

A company is deploying a generative AI application using Vertex AI. They need to minimize latency for real‑time inference while maintaining high quality. Which TWO actions are most effective?

⚠ Common exam trap

The trap is assuming that larger models always provide better quality and that batching reduces latency; in reality, smaller models and shorter outputs reduce latency, while batching increases it.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a smaller model like Gemini Flash

Option B is correct because Gemini Flash is a lower-latency, lighter-weight model in the Gemini family, purpose-built for high-throughput, real-time tasks, so it directly reduces inference latency while still delivering strong quality for most generative AI use cases. Option D is correct because output tokens are generated sequentially (autoregressive decoding), so capping max output tokens to the shortest acceptable length proportionally cuts generation time and thus end-to-end latency. Option A is not appropriate because batching requests improves throughput and cost efficiency, not per-request real-time latency, and can actually add queuing delay. Option C is wrong because Gemini Pro is larger and slower than Flash, trading latency for quality rather than minimizing latency. Option E is wrong because raising temperature to 1.0 only affects randomness/creativity of sampling and does not reduce latency.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Batch multiple inference requests together

    Why it's wrong here

    Batching deliberately queues requests until a batch fills, which adds waiting time and raises per-request latency, directly contradicting the real-time requirement. It is tempting because batching maximises throughput and reduces cost per inference for offline or bulk scoring workloads where latency is irrelevant.

  • ✓

    Use a smaller model like Gemini Flash

    Why this is correct

    Gemini Flash trades parameter count for throughput, so each decoding step costs fewer FLOPs and returns tokens faster than Pro-class models. That directly cuts per-token inference latency, meeting the real-time constraint while quality remains adequate for many production tasks.

  • ✗

    Use Gemini Pro instead of Gemini Flash to ensure quality

    Why it's wrong here

    Gemini Pro carries higher inference latency than Flash, contradicting the real-time latency requirement. It is tempting because Pro offers greater capability, and would be correct where output quality outweighs response time, such as complex reasoning or batch generation rather than latency-sensitive serving.

  • ✓

    Reduce the max output tokens to the minimum acceptable length

    Why this is correct

    Capping max output tokens directly shortens the autoregressive decoding loop, since each generated token adds sequential latency. This satisfies the real-time inference constraint by bounding generation time, while the "minimum acceptable length" qualifier preserves output quality. Combined with streaming, it is the most direct latency lever available in Vertex AI.

  • ✗

    Increase the temperature to 1.0 for more creative outputs

    Why it's wrong here

    Temperature 1.0 increases output randomness, which does not reduce inference latency and can lower quality through less deterministic responses. It is tempting because raising temperature is the standard technique for creative or diverse generation tasks, such as brainstorming or marketing copy, where variety matters more than speed.

About these practice questions

One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.