Generative AI Leader Practice Question: Techniques to Improve Generative AI Model Output
A user provides a long document as context for a question-answering task, but the model outputs irrelevant answers. What is the most likely cause?
⚠ Common exam trap
Google often tests the misconception that safety filters or temperature settings are the primary cause of irrelevant outputs, when in fact the context window limit is the most direct and common technical constraint in long-document QA tasks.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The document exceeds the model's context window, truncating important details.
The most common cause of irrelevant answers when a long document is provided is that the document exceeds the model's fixed context window (e.g., 8K tokens for PaLM 2, 128K for Gemini 1.0 Pro, or up to 1M for Gemini 1.5 Pro). When the input is truncated, critical details needed for accurate retrieval and generation are lost, leading to off-target responses. This is a fundamental limitation of transformer architectures, which cannot attend to tokens beyond their maximum sequence length.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
The document exceeds the model's context window, truncating important details.
Why this is correct
Transformer models process a fixed token budget; text beyond the context window is silently truncated. With a long document, the crucial passages answering the question may fall outside that window, so the model responds from incomplete context, producing irrelevant output.
- ✗
Safety filters are blocking the relevant response.
Why it's wrong here
Safety filters suppress disallowed content, producing refusals or omissions, not coherent but off-topic answers; irrelevant output indicates the context exceeded the model's effective window or was poorly attended. It is tempting because filters do alter responses, and would be correct if the model returned a refusal or blocked message.
- ✗
The model's temperature is too low, making it deterministic.
Why it's wrong here
Low temperature reduces sampling randomness, making output focused and repeatable; it does not cause a model to ignore supplied context. It is tempting because temperature governs creativity, and would be correct if the model were drifting off-topic through excessive randomness at high settings.
- ✗
The model is not generating any tokens.
Why it's wrong here
Zero token generation yields an empty response, whereas the scenario describes answers being produced, merely irrelevant ones. It is tempting because truncation and quota limits can halt output, and would be correct if the model returned nothing at all rather than unrelated text.
Go deeper
Related to this question
About these practice questions
One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.