Courseiva
hardMultiple Choice

Generative AI Leader Practice Question: A data scientist is using Vertex AI RAG Engine to…

A data scientist is using Vertex AI RAG Engine to build a question-answering system over a large corpus of technical manuals. Users report that answers are often verbose and include irrelevant details. Which configuration change is MOST likely to improve answer conciseness?

⚠ Common exam trap

The trap is confusing retrieval quality with generation length — candidates pick embedding model changes or chunk size, but the direct control on answer verbosity is the number of retrieved documents (k) fed into the prompt.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Decrease the number of retrieved documents (k) in the retrieval step

In a RAG pipeline, the number of retrieved documents (k) directly controls how much context is fed to the LLM. Reducing k limits the amount of potentially irrelevant material the model sees, which typically makes answers shorter and more focused. This is the most direct lever for conciseness.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the chunk size of documents in the index

    Why it's wrong here

    Larger chunks pack more surrounding text into each retrieved passage, giving the model extra material to echo as irrelevant detail. It is tempting because bigger chunks appear to supply richer context, but the fix is smaller, more precise chunks or reranking, not larger ones.

  • ✗

    Use a different embedding model for retrieval

    Why it's wrong here

    The embedding model governs retrieval relevance, not the generator's phrasing, so swapping it leaves verbose answers unchanged even if better chunks are fetched. It is tempting because irrelevant details suggest poor retrieval, yet conciseness is controlled at generation, via prompts or output limits.

  • ✓

    Decrease the number of retrieved documents (k) in the retrieval step

    Why this is correct

    Lowering k reduces the volume of retrieved context passed to the generative model, directly limiting the extraneous material it can draw into the response. Since verbosity stems from irrelevant retrieved chunks, fewer documents constrain the model to the most pertinent passages, improving conciseness without altering the underlying corpus or prompt.

  • ✗

    Increase the max output token limit

    Why it's wrong here

    Raising the max output token limit permits longer generations, directly enabling the verbosity reported. It is tempting because token limits are the obvious length control, but the correct lever is a lower limit or a prompt instructing brevity, not a higher ceiling.

About these practice questions

One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.