Courseiva

Generative AI Leader Fundamentals of Generative AI Practice Question

A machine learning engineer is building a text-to-image model using Vertex AI. They want to reduce inference latency. Which strategy is most effective?

⚠ Common exam trap

Many exam-takers confuse throughput optimization (batch processing) with latency reduction, or assume that more steps or higher resolution improve quality without considering the latency trade-off.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a smaller model variant

Using a smaller model variant directly reduces the number of parameters and computational operations required per inference pass, which lowers latency. In text-to-image models like Imagen or Stable Diffusion, the model size is the primary driver of forward-pass time, so a smaller variant (e.g., fewer layers or reduced latent dimensions) yields faster generation.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use a larger image resolution

    Why it's wrong here

    Higher resolution increases the number of pixels the diffusion process must denoise, directly raising inference latency rather than reducing it. It is tempting because resolution is a legitimate quality lever when output fidelity matters more than speed, but the stem explicitly targets latency, so this option moves the metric the wrong way.

  • ✓

    Use a smaller model variant

    Why this is correct

    A smaller model variant reduces the number of parameters and floating-point operations per forward pass, directly cutting compute per generated image. This satisfies the latency constraint because inference time scales with model size, unlike batching or caching, which improve throughput or repeat requests rather than single-pass generation speed.

  • ✗

    Enable batch processing

    Why it's wrong here

    Batch processing defers requests and returns results asynchronously, so it cannot lower per-request inference latency; it actually adds queueing delay. It is designed for high-volume, latency-tolerant workloads such as overnight scoring of stored data, where throughput and cost matter rather than immediate responses.

  • ✗

    Increase the number of inference steps

    Why it's wrong here

    Increasing inference steps adds denoising iterations, directly raising compute time per generated image, so latency worsens rather than improves. It is tempting because more steps genuinely raise image fidelity and prompt adherence in diffusion models — the right lever when output quality, not speed, is the constraint.

About these practice questions

One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.