AI-102 Implement generative AI solutions Practice Question
You are configuring an Azure OpenAI deployment for a generative AI solution that summarizes long legal contracts. Users report that summaries sometimes omit clauses near the end of documents. The documents are up to 120 pages. You need to improve completeness without changing the model. What should you do?
⚠ Common exam trap
The trap here is assuming that a larger output token limit lets the model consider more of the input document.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Split each document into overlapping chunks, summarize each chunk, and then combine the summaries in a final aggregation step.
Long documents can exceed the model context window, causing later sections to be dropped. Chunking with overlap preserves boundary content, and summarizing chunks before aggregating covers the whole document. Output length, temperature, and streaming do not affect how much input text the model can attend to, so they cannot fix missing clauses.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the max_tokens parameter of the completion request to its maximum value.
Why it's wrong here
Max tokens controls the length of the generated response, not how much input the model can consider. Raising it may produce longer summaries but does not prevent the model from missing clauses that fall outside the input context window, so the completeness problem remains.
- ✗
Enable streaming responses so the client receives partial summaries as they are generated.
Why it's wrong here
Streaming changes delivery timing of the response, not the amount of source text the model processes. The model still receives only what fits in the context window, so clauses beyond that limit remain unseen. Streaming improves perceived latency but does not improve summary completeness.
- ✓
Split each document into overlapping chunks, summarize each chunk, and then combine the summaries in a final aggregation step.
Why this is correct
Chunking with overlap ensures content near boundaries is not lost, and a map-reduce style aggregation summarizes each part before combining them. This keeps each request within the context window while covering the entire document, which directly addresses omitted clauses near the end without changing the model.
- ✗
Raise the temperature so the model explores more of the document content.
Why it's wrong here
Temperature affects randomness in token selection, not input coverage. Higher temperature makes output less predictable and can introduce inaccuracies, but it does not give the model access to text beyond the context window. It cannot recover clauses that were truncated from the prompt.
Go deeper
Related to this question
About these practice questions
This AI-102 question is part of Courseiva's 761-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Microsoft exam blueprint
This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.