Courseiva

Generative AI Leader Google Cloud's Generative AI Offerings Practice Question

A company is deploying a large language model on Vertex AI for real-time inference. They observe high latency and want to optimize. They have already enabled model caching. What next step should they take to reduce latency?

⚠ Common exam trap

Google often tests the misconception that adding more GPUs or increasing batch size reduces latency for real-time inference, when in fact these optimizations primarily improve throughput and can increase per-request latency.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Apply model quantization to reduce precision

Model quantization reduces the precision of the model's weights (e.g., from FP32 to INT8), which decreases memory footprint and accelerates computation on the hardware, directly lowering inference latency. Since Vertex AI already has model caching enabled, quantization is the next logical optimization step to reduce latency without requiring additional infrastructure changes.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Add more GPUs to the prediction endpoint

    Why it's wrong here

    Adding GPUs raises parallel throughput, not per-request latency; the model still executes the same sequential token generation. It is tempting because horizontal scaling suits batch or high-volume serving, but real-time latency needs a smaller distilled model, lower precision, or speculative decoding instead.

  • ✗

    Use a larger, more accurate model variant

    Why it's wrong here

    A larger model variant adds parameters and computation per token, increasing inference latency; accuracy gains do not offset the slower response. It is tempting because accuracy is often prioritised, and it would be correct where output quality outweighs speed, such as asynchronous batch processing rather than real-time serving.

  • ✗

    Increase the batch size for inference requests

    Why it's wrong here

    Increasing batch size groups more requests per inference pass, which raises throughput but adds queueing delay before each response, worsening latency for real-time serving. It is tempting because batching is a standard optimisation, and it would be correct for offline or bulk scoring where throughput matters more than per-request response time.

  • ✓

    Apply model quantization to reduce precision

    Why this is correct

    Quantization stores weights in lower precision, such as INT8, shrinking memory footprint and enabling faster matrix operations, which directly cuts inference latency. Caching is already applied, so reducing per-token compute is the remaining lever for real-time serving on Vertex AI.

About these practice questions

One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.