NCA-GENL Software Development Practice Question
A developer is building a customer-support assistant on NVIDIA NIM microservices. After a model update, responses that previously arrived in under 300 ms now take over two seconds, and the streaming client shows a long pause before the first token appears. GPU utilization is low and the prompt template was not changed. Which action should the developer take first to diagnose the regression?
⚠ Common exam trap
The trap here is assuming that low GPU utilization automatically means the fix is more GPUs or tensor parallelism, when the real question is which inference phase became slow.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Compare the time to first token (TTFT) and inter-token latency measured before and after the update to isolate whether the delay is in prefill or decode.
The reported symptom is a latency regression with low GPU utilization and unchanged prompts, so the developer must localize the delay before changing deployment configuration. Instrumenting time to first token and inter-token latency cleanly separates prompt prefill from token decode and reveals which phase regressed. Batching, parallelism, and disabling streaming either mask the signal or alter behavior, so measurement comes first.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Compare the time to first token (TTFT) and inter-token latency measured before and after the update to isolate whether the delay is in prefill or decode.
Why this is correct
Splitting latency into time to first token and inter-token latency immediately separates prefill (prompt processing) from decode (token generation). A long pause before the first token with fast subsequent tokens points at prefill, while slow steady output points at decode. This measurement requires no code change and directly targets the reported symptom, making it the correct first diagnostic step for the NIM deployment.
- ✗
Increase the tensor-parallel degree of the NIM container so the model weights are sharded across more GPUs.
Why it's wrong here
Tensor parallelism shards weights and attention heads across GPUs and can raise throughput, but applying it before measuring where the latency lives risks a costly redeploy that may not touch the bottleneck. Low GPU utilization does not by itself prove the model needs more shards; it can also indicate host-side or scheduling overhead. Diagnosis must precede this configuration change.
- ✗
Switch the client from streaming to a single non-streaming request so the full response is timed end to end.
Why it's wrong here
Removing streaming discards the very signal that distinguishes prefill from decode: the pause before the first token. A non-streaming wall-clock measurement collapses both phases into one number and hides where the two seconds are spent. It also changes the client behavior under test, so the comparison against the previous 300 ms baseline becomes invalid.
- ✗
Enable dynamic batching in the NIM microservice so concurrent requests are grouped into larger inference batches.
Why it's wrong here
Dynamic batching improves aggregate throughput under concurrency, but it can actually add latency to individual requests because a request waits for the batch window to close. The scenario describes a single slow response with low GPU utilization, which is not a throughput starvation symptom. Turning on batching first would confound the measurement rather than explain the regression.
About these practice questions
One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.