Courseiva

AI-102 Implement generative AI solutions Practice Question

You are preparing an Azure OpenAI deployment for a production generative AI application that must stream responses to a web front end and must limit the cost of overly long conversations. You need to configure parameters that control response length and streaming behavior. Which two parameters should you set? (Choose two.)

⚠ Common exam trap

The trap here is treating sampling parameters like temperature or top_p as cost controls, when only the completion-length limit actually bounds how many tokens are billed.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

stream

Two distinct needs are stated: bounding the cost of long conversations and streaming responses to a web front end. The max_tokens parameter limits how many tokens each completion may contain, providing the cost ceiling. The stream parameter switches the API to incremental server-sent delivery, satisfying the front-end streaming requirement. Sampling parameters such as temperature, top_p, and frequency_penalty influence output content rather than length or delivery.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    frequency_penalty

    Why it's wrong here

    Frequency penalty reduces the likelihood of repeating tokens that have already appeared, which can make output more varied. It has no effect on how many tokens are generated overall or on whether the response is streamed. It is a style control and does not address the cost-limiting or streaming requirements in the scenario.

  • ✗

    top_p

    Why it's wrong here

    Top_p, or nucleus sampling, restricts sampling to the smallest set of tokens whose cumulative probability exceeds the threshold. It shapes which tokens are candidates and thus output diversity, but it neither caps completion length nor enables streaming. It is a sampling control, not a cost or delivery control.

  • ✗

    temperature

    Why it's wrong here

    Temperature controls the randomness of sampling and therefore the creativity or determinism of the output. While it influences response character, it does not bound response length or enable incremental delivery, so it does not satisfy either stated requirement. Adjusting it is a quality-tuning decision rather than a cost or streaming control.

  • ✓

    stream

    Why this is correct

    Setting stream to true causes the service to return the completion incrementally as server-sent events rather than as one blocking payload. This is exactly what a web front end needs to display tokens as they are produced, improving perceived latency. It does not change token accounting or cost, so it pairs naturally with a length limit.

  • ✓

    max_tokens

    Why this is correct

    max_tokens caps the number of tokens the model may generate in the completion, which directly bounds the cost of each response and prevents runaway output. For a production application that must limit the expense of long conversations, setting this parameter is the standard control. It does not affect the input prompt length, which is a separate consideration.

About these practice questions

This AI-102 question is part of Courseiva's 761-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Microsoft exam blueprint

This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.