AI-102 Implement generative AI solutions Practice Question
A developer is building a generative AI application with Azure OpenAI Service. The application must stream partial responses to users as tokens are generated, rather than waiting for the complete response. Which parameter should the developer set in the chat completion request?
⚠ Common exam trap
Candidates often confuse parameters that shape output content, such as temperature or max_tokens, with the parameter that changes how output is delivered over the network.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Set 'stream' to true.
Streaming is controlled by the stream parameter, which when true causes the service to emit incremental delta chunks over a persistent connection. This lets the application display tokens as they are generated, reducing perceived latency while keeping the same model and prompt configuration.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Set 'temperature' to 0.
Why it's wrong here
Temperature adjusts randomness in token selection, making output more deterministic when lowered. It has no effect on whether the response is delivered incrementally or as a single payload. Setting it to zero will not enable streaming and may reduce creativity, which is unrelated to the streaming requirement.
- ✓
Set 'stream' to true.
Why this is correct
Setting stream to true enables server-sent events so the service returns incremental chunks as tokens are produced. The client can render partial text immediately, which improves perceived latency for long generations. This is the documented parameter for streaming chat completions in Azure OpenAI Service.
- ✗
Set 'max_tokens' to a large value.
Why it's wrong here
Max_tokens caps the length of the generated response. Increasing it allows longer outputs but still returns the full completion in one response by default. It does not change the transport behavior, so users would still wait for the entire generation before seeing any text.
- ✗
Set 'n' to a value greater than 1.
Why it's wrong here
The n parameter controls how many alternative completions the model returns for a single prompt. It does not stream tokens incrementally; instead it produces multiple full responses. Using it increases token consumption and does not satisfy the requirement to deliver partial output as generation proceeds.
Go deeper
Related to this question
About these practice questions
This AI-102 question is part of Courseiva's 761-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Microsoft exam blueprint
This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.