hardMultiple Choice
Generative AI Leader Practice Question: A data scientist is using Vertex AI RAG Engine to…
A data scientist is using Vertex AI RAG Engine to build a question-answering system over a large corpus of technical manuals. Users report that answers are often verbose and include irrelevant details. Which configuration change is MOST likely to improve answer conciseness?
⚠ Common exam trap
The trap is confusing retrieval quality with generation length — candidates pick embedding model changes or chunk size, but the direct control on answer verbosity is the number of retrieved documents (k) fed into the prompt.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Decrease the number of retrieved documents (k) in the retrieval step
In a RAG pipeline, the number of retrieved documents (k) directly controls how much context is fed to the LLM. Reducing k limits the amount of potentially irrelevant material the model sees, which typically makes answers shorter and more focused. This is the most direct lever for conciseness.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the chunk size of documents in the index
Why it's wrong here
Larger chunks pack more surrounding text into each retrieved passage, giving the model extra material to echo as irrelevant detail. It is tempting because bigger chunks appear to supply richer context, but the fix is smaller, more precise chunks or reranking, not larger ones.
- ✗
Use a different embedding model for retrieval
Why it's wrong here
The embedding model governs retrieval relevance, not the generator's phrasing, so swapping it leaves verbose answers unchanged even if better chunks are fetched. It is tempting because irrelevant details suggest poor retrieval, yet conciseness is controlled at generation, via prompts or output limits.
- ✓
Decrease the number of retrieved documents (k) in the retrieval step
Why this is correct
Lowering k reduces the volume of retrieved context passed to the generative model, directly limiting the extraneous material it can draw into the response. Since verbosity stems from irrelevant retrieved chunks, fewer documents constrain the model to the most pertinent passages, improving conciseness without altering the underlying corpus or prompt.
- ✗
Increase the max output token limit
Why it's wrong here
Raising the max output token limit permits longer generations, directly enabling the verbosity reported. It is tempting because token limits are the obvious length control, but the correct lever is a lower limit or a prompt instructing brevity, not a higher ceiling.
Go deeper
Related to this question
About these practice questions
One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.