NCA-GENL Software Development Practice Question
An application requires streaming responses from a deployed LLM. Which communication protocol is most suitable for minimizing latency and ensuring efficient data delivery in a real-time generative AI application?
⚠ Common exam trap
Candidates often confuse gRPC with standard HTTP/1.1 REST APIs, forgetting that REST lacks native bidirectional streaming capabilities and incurs higher overhead for continuous token transmission in real-time generative applications.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
gRPC with server-side streaming.
Streaming generative responses requires a protocol that supports persistent connections and low-overhead message framing. gRPC with server-side streaming is the industry standard for high-performance AI inference, as it facilitates efficient binary serialization and maintains low latency across the network. Using this protocol is critical for user-facing applications where perceived latency is directly tied to the speed at which text tokens appear on the screen.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
HTTP/1.1 REST without streaming support.
Why it's wrong here
HTTP/1.1 REST typically waits for the full response to be generated before returning, which introduces significant latency for generative tasks. This batch-oriented behavior is unsuitable for applications needing real-time interactivity, as the user must wait for the entire completion before seeing any output token.
- ✓
gRPC with server-side streaming.
Why this is correct
gRPC uses HTTP/2 for transport, which supports server-side streaming. This allows tokens to be sent back as soon as they are generated, minimizing the time-to-first-token and creating a fluid, real-time experience for the end-user. It is the preferred method for high-performance communication in modern AI stacks.
- ✗
Standard SMTP email protocols.
Why it's wrong here
SMTP is designed for asynchronous email transmission and lacks the low-latency capabilities required for interactive AI services. Using this protocol for model inference would introduce extreme delays and is architecturally fundamentally mismatched with the needs of a generative AI interface that requires real-time token delivery.
- ✗
Synchronous SQL query polling.
Why it's wrong here
Polling a database for status updates is inefficient and adds unnecessary network round-trips. It consumes excessive resources on both the client and server sides, resulting in poor responsiveness that contradicts the requirement for high-performance, real-time streaming of generative AI model outputs to the client application.
About these practice questions
One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.