AI-102 Implement generative AI solutions Practice Question
You are building a generative AI feature in an application using the Azure OpenAI SDK. The feature must stream tokens to the user as they are generated and must also capture the full response for logging. Which implementation approach satisfies both requirements?
⚠ Common exam trap
The trap here is believing that a streaming response cannot also be captured in full, when accumulating the delta chunks reconstructs the complete message.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Set stream=true on the chat completions request, iterate over the streamed chunks to display deltas, and accumulate the delta content to form the complete response for logging.
Streaming chat completions deliver delta content across multiple chunk objects. Rendering those deltas satisfies the progressive display requirement, and concatenating them rebuilds the exact full message for logging. A single streaming request therefore meets both needs, whereas duplicate calls, partial logging, or multiple completions fail one requirement or the other.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Disable streaming and set n=2 so the model returns two completions, displaying one and logging the other.
Why it's wrong here
The n parameter requests multiple independent completions, which are separate generations, not a display copy and a log copy of the same output. Streaming remains disabled, so tokens are not delivered incrementally. This approach increases cost and still fails the real-time streaming requirement.
- ✗
Make two separate non-streaming requests: one to display the response and one to log it.
Why it's wrong here
Two non-streaming requests waste tokens and can produce differing outputs because generation is probabilistic. Neither request streams, so the user waits for the full response, failing the streaming requirement. Duplicating calls also doubles cost and latency without guaranteeing identical logged content.
- ✗
Set stream=true and log only the first chunk received from the stream.
Why it's wrong here
The first chunk typically contains the role and initial delta tokens, not the whole message. Logging only that chunk loses most of the generated content, so the logging requirement fails. Streaming does not change the fact that the complete response must be assembled from all chunks.
- ✓
Set stream=true on the chat completions request, iterate over the streamed chunks to display deltas, and accumulate the delta content to form the complete response for logging.
Why this is correct
Streaming returns incremental chat completion chunk objects containing delta content. Displaying each delta gives the user progressive output, while concatenating deltas reconstructs the full message for logging. This single request satisfies both real-time display and complete capture without a second call.
Go deeper
Related to this question
About these practice questions
One of 761 original AI-102 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Microsoft exam blueprint
This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.