Courseiva

AI-102 Implement generative AI solutions Practice Question

You are building a generative AI feature in an application using the Azure OpenAI SDK. The feature must stream tokens to the user as they are generated and must also capture the full response for logging. Which implementation approach satisfies both requirements?

⚠ Common exam trap

The trap here is believing that a streaming response cannot also be captured in full, when accumulating the delta chunks reconstructs the complete message.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Set stream=true on the chat completions request, iterate over the streamed chunks to display deltas, and accumulate the delta content to form the complete response for logging.

Streaming chat completions deliver delta content across multiple chunk objects. Rendering those deltas satisfies the progressive display requirement, and concatenating them rebuilds the exact full message for logging. A single streaming request therefore meets both needs, whereas duplicate calls, partial logging, or multiple completions fail one requirement or the other.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Disable streaming and set n=2 so the model returns two completions, displaying one and logging the other.

    Why it's wrong here

    The n parameter requests multiple independent completions, which are separate generations, not a display copy and a log copy of the same output. Streaming remains disabled, so tokens are not delivered incrementally. This approach increases cost and still fails the real-time streaming requirement.

  • ✗

    Make two separate non-streaming requests: one to display the response and one to log it.

    Why it's wrong here

    Two non-streaming requests waste tokens and can produce differing outputs because generation is probabilistic. Neither request streams, so the user waits for the full response, failing the streaming requirement. Duplicating calls also doubles cost and latency without guaranteeing identical logged content.

  • ✗

    Set stream=true and log only the first chunk received from the stream.

    Why it's wrong here

    The first chunk typically contains the role and initial delta tokens, not the whole message. Logging only that chunk loses most of the generated content, so the logging requirement fails. Streaming does not change the fact that the complete response must be assembled from all chunks.

  • ✓

    Set stream=true on the chat completions request, iterate over the streamed chunks to display deltas, and accumulate the delta content to form the complete response for logging.

    Why this is correct

    Streaming returns incremental chat completion chunk objects containing delta content. Displaying each delta gives the user progressive output, while concatenating deltas reconstructs the full message for logging. This single request satisfies both real-time display and complete capture without a second call.

About these practice questions

One of 761 original AI-102 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Microsoft exam blueprint

This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.