Courseiva
Software Development →mediumMultiple Choice

NCA-GENL Software Development Practice Question

A developer is building a document-summarization service on NVIDIA NIM for LLMs and wants to stream partial tokens to the client while the model is still generating. The NIM endpoint exposes an OpenAI-compatible /chat/completions route. Which request parameter should the developer set to receive incremental token deltas rather than one complete response body?

⚠ Common exam trap

The trap here is assuming that an Accept header or a parameter like n or logprobs changes response timing, when only the stream field actually enables incremental token delivery.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Set "stream": true in the JSON request body and consume the server-sent event chunks.

Streaming on an OpenAI-compatible NIM chat completions endpoint is controlled by the stream boolean in the request body, which switches the transport to server-sent events carrying incremental delta objects. The other parameters alter candidate count, content negotiation, or scoring metadata, and none of them cause tokens to be emitted before generation completes.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Set "logprobs": true and reconstruct partial text from the token log probabilities.

    Why it's wrong here

    Requesting log probabilities returns scoring metadata alongside the completed text, not partial output. The response is still delivered as one payload after the model finishes, so the developer would have per-token confidence values but no way to render tokens to the user before the whole summarization job completes.

  • ✗

    Set the SSE header "Accept: text/event-stream" on the request without changing the body.

    Why it's wrong here

    Negotiating a content type does not by itself switch the endpoint into streaming mode. The OpenAI-compatible chat completions route decides its response framing from the stream field in the JSON body, so a header alone yields a normal single JSON response and the client still receives nothing until generation completes.

  • ✗

    Set "n": 2 so the endpoint returns two candidate completions the client can merge.

    Why it's wrong here

    The n parameter controls how many independent completion candidates are generated and billed, not how a single completion is delivered. Requesting multiple candidates still returns one complete JSON body after generation finishes, so the client would wait for the full summary and receive duplicate outputs it must then reconcile.

  • ✓

    Set "stream": true in the JSON request body and consume the server-sent event chunks.

    Why this is correct

    The OpenAI-compatible chat completions schema used by NIM for LLMs accepts a boolean stream field. When it is true, the server returns text/event-stream data chunks containing incremental choices delta objects terminated by a data: [DONE] sentinel, which is exactly the incremental token delivery the summarization UI needs.

About these practice questions

One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.