AI-102 Implement generative AI solutions Practice Question
You are preparing an Azure OpenAI deployment for a production generative AI application that must stream responses to a web front end and must limit the cost of overly long conversations. You need to configure parameters that control response length and streaming behavior. Which two parameters should you set? (Choose two.)
⚠ Common exam trap
The trap here is treating sampling parameters like temperature or top_p as cost controls, when only the completion-length limit actually bounds how many tokens are billed.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
stream
Two distinct needs are stated: bounding the cost of long conversations and streaming responses to a web front end. The max_tokens parameter limits how many tokens each completion may contain, providing the cost ceiling. The stream parameter switches the API to incremental server-sent delivery, satisfying the front-end streaming requirement. Sampling parameters such as temperature, top_p, and frequency_penalty influence output content rather than length or delivery.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
frequency_penalty
Why it's wrong here
Frequency penalty reduces the likelihood of repeating tokens that have already appeared, which can make output more varied. It has no effect on how many tokens are generated overall or on whether the response is streamed. It is a style control and does not address the cost-limiting or streaming requirements in the scenario.
- ✗
top_p
Why it's wrong here
Top_p, or nucleus sampling, restricts sampling to the smallest set of tokens whose cumulative probability exceeds the threshold. It shapes which tokens are candidates and thus output diversity, but it neither caps completion length nor enables streaming. It is a sampling control, not a cost or delivery control.
- ✗
temperature
Why it's wrong here
Temperature controls the randomness of sampling and therefore the creativity or determinism of the output. While it influences response character, it does not bound response length or enable incremental delivery, so it does not satisfy either stated requirement. Adjusting it is a quality-tuning decision rather than a cost or streaming control.
- ✓
stream
Why this is correct
Setting stream to true causes the service to return the completion incrementally as server-sent events rather than as one blocking payload. This is exactly what a web front end needs to display tokens as they are produced, improving perceived latency. It does not change token accounting or cost, so it pairs naturally with a length limit.
- ✓
max_tokens
Why this is correct
max_tokens caps the number of tokens the model may generate in the completion, which directly bounds the cost of each response and prevents runaway output. For a production application that must limit the expense of long conversations, setting this parameter is the standard control. It does not affect the input prompt length, which is a separate consideration.
Go deeper
Related to this question
About these practice questions
This AI-102 question is part of Courseiva's 761-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Microsoft exam blueprint
This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.