You are preparing an Azure OpenAI deployment for a production generative AI application that must stream responses to a web front end and must limit the cost of overly long conversations. You need to configure parameters that control response length and streaming behavior. Which two parameters should you set? (Choose two.)
Setting stream to true causes the service to return the completion incrementally as server-sent events rather than as one blocking payload. This is exactly what a web front end needs to display tokens as they are produced, improving perceived latency. It does not change token accounting or cost, so it pairs naturally with a length limit.
Why this answer
Two distinct needs are stated: bounding the cost of long conversations and streaming responses to a web front end. The max_tokens parameter limits how many tokens each completion may contain, providing the cost ceiling. The stream parameter switches the API to incremental server-sent delivery, satisfying the front-end streaming requirement.
Sampling parameters such as temperature, top_p, and frequency_penalty influence output content rather than length or delivery.
Exam trap
The trap here is treating sampling parameters like temperature or top_p as cost controls, when only the completion-length limit actually bounds how many tokens are billed.