Courseiva

Generative AI Leader Google Cloud's Generative AI Offerings Practice Question

A company is building a document summarization tool using Vertex AI Gemini API. They notice that the model sometimes returns incomplete summaries that miss key points. Which approach is most likely to improve summary quality without increasing token usage significantly?

⚠ Common exam trap

In Google Cloud exams, a common misconception is that increasing model size or token limits directly improves output quality, when in fact prompt engineering—specifically system instructions—is a more efficient and cost-effective lever for controlling model behavior.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Refine the system instruction to specify the desired summary format and key elements to include

Refining the system instruction directly addresses the root cause of incomplete summaries by providing explicit guidance on the desired output format and key elements to include. This approach improves the model's adherence to the task without increasing the number of input or output tokens, as it only modifies the instruction text, not the document length or generation limits.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Refine the system instruction to specify the desired summary format and key elements to include

    Why this is correct

    Refining the system instruction specifies the required summary format and key elements, steering the model to cover them without enlarging the input. This improves completeness at negligible token cost, unlike approaches that add lengthy examples or extra context.

  • ✗

    Increase the context window to include more of the document

    Why it's wrong here

    Enlarging the context window feeds more document text into every request, directly increasing input token usage — the opposite of the stated constraint. Context expansion suits cases where relevant content is genuinely excluded; here the document already fits, so summarisation instructions or grounding are needed.

  • ✗

    Switch to a larger Gemini model (e.g., from 1.0 Pro to 1.5 Pro)

    Why it's wrong here

    A larger Gemini model increases per-token cost and latency, contradicting the requirement to avoid significant token growth, and model size does not guarantee coverage of missed points. Larger models suit complex reasoning tasks; here prompt refinement or grounding addresses omission without changing token economics.

  • ✗

    Increase the max output token limit to allow longer summaries

    Why it's wrong here

    Raising max output tokens lets the model write longer, but incompleteness stems from prompt or context handling, not truncation, and longer outputs consume more tokens. This setting suits tasks where responses are genuinely cut off mid-sentence, not summaries that omit key points.

About these practice questions

Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.