Courseiva
Software Development →mediumMultiple Choice

NCA-GENL Software Development Practice Question

An application requires streaming responses from a deployed LLM. Which communication protocol is most suitable for minimizing latency and ensuring efficient data delivery in a real-time generative AI application?

⚠ Common exam trap

Candidates often confuse gRPC with standard HTTP/1.1 REST APIs, forgetting that REST lacks native bidirectional streaming capabilities and incurs higher overhead for continuous token transmission in real-time generative applications.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

gRPC with server-side streaming.

Streaming generative responses requires a protocol that supports persistent connections and low-overhead message framing. gRPC with server-side streaming is the industry standard for high-performance AI inference, as it facilitates efficient binary serialization and maintains low latency across the network. Using this protocol is critical for user-facing applications where perceived latency is directly tied to the speed at which text tokens appear on the screen.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    HTTP/1.1 REST without streaming support.

    Why it's wrong here

    HTTP/1.1 REST typically waits for the full response to be generated before returning, which introduces significant latency for generative tasks. This batch-oriented behavior is unsuitable for applications needing real-time interactivity, as the user must wait for the entire completion before seeing any output token.

  • ✓

    gRPC with server-side streaming.

    Why this is correct

    gRPC uses HTTP/2 for transport, which supports server-side streaming. This allows tokens to be sent back as soon as they are generated, minimizing the time-to-first-token and creating a fluid, real-time experience for the end-user. It is the preferred method for high-performance communication in modern AI stacks.

  • ✗

    Standard SMTP email protocols.

    Why it's wrong here

    SMTP is designed for asynchronous email transmission and lacks the low-latency capabilities required for interactive AI services. Using this protocol for model inference would introduce extreme delays and is architecturally fundamentally mismatched with the needs of a generative AI interface that requires real-time token delivery.

  • ✗

    Synchronous SQL query polling.

    Why it's wrong here

    Polling a database for status updates is inefficient and adds unnecessary network round-trips. It consumes excessive resources on both the client and server sides, resulting in poor responsiveness that contradicts the requirement for high-performance, real-time streaming of generative AI model outputs to the client application.

About these practice questions

One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.