Courseiva
mediumMultiple Select

Generative AI Leader Practice Question: A machine learning engineer wants to reduce the…

A machine learning engineer wants to reduce the latency of a Gemini-based chatbot running in production. Which TWO strategies would be MOST effective?

⚠ Common exam trap

A common misconception in Google exams is that streaming reduces total latency, but it only improves perceived latency (time-to-first-token) while total processing time remains unchanged.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Switch from Gemini Pro to Gemini Flash

Option A is correct because Gemini Flash is a lighter, lower-latency model variant than Gemini Pro, so switching to it directly reduces the time the model takes to generate each response. Option B is correct because max_output_tokens caps the number of tokens the model generates, and since generation time scales roughly with output length, lowering this value shortens the response and thus the overall latency. Option C is not correct because streaming improves perceived latency by delivering tokens incrementally, but it does not reduce the actual total generation time. Option D is not correct because fine-tuning adapts the model to a task and does not inherently lower inference latency. Option E is not correct because raising temperature to 1.0 affects randomness and creativity, not speed, and can even make outputs longer or less predictable.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Switch from Gemini Pro to Gemini Flash

    Why this is correct

    Gemini Flash is a smaller, optimised model variant with substantially lower inference latency than Gemini Pro, so swapping models reduces per-request response time. This directly addresses the production chatbot's latency constraint without changing the application logic.

  • ✓

    Reduce the max_output_tokens parameter

    Why this is correct

    Capping max_output_tokens directly shortens the decode phase, since each generated token requires a separate forward pass through the model. Fewer output tokens means less sequential autoregressive computation, which is the dominant latency contributor for long responses. This satisfies the stem's production latency constraint without altering model quality or requiring infrastructure changes.

  • ✗

    Enable streaming mode

    Why it's wrong here

    Streaming returns tokens incrementally, so perceived time-to-first-token drops, but total generation latency is unchanged. It is tempting because users see output sooner, yet the stem asks for reduced latency, which streaming masks rather than removes.

  • ✗

    Fine-tune the model on the specific task

    Why it's wrong here

    Fine-tuning alters model weights for task accuracy, not inference speed; a tuned model still runs the same forward pass. It is tempting because fine-tuning reduces prompt length, but that is an indirect prompt-engineering effect, not the mechanism this scenario needs.

  • ✗

    Increase the temperature to 1.0

    Why it's wrong here

    Temperature controls output randomness, not token generation speed, so raising it to 1.0 adds variability without cutting latency. It is tempting because sampling parameters feel like tuning knobs, but temperature is chosen for creativity, not throughput.

About these practice questions

Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.