Courseiva
Software Development →hardMultiple Choice

NCA-GENL Software Development Practice Question

A developer is profiling a TensorRT-LLM serving deployment and notices that throughput collapses once concurrent requests exceed a small number of users, even though GPU compute utilization stays low. The model uses paged KV cache and continuous batching. Which factor most likely explains the bottleneck?

⚠ Common exam trap

The trap here is equating low GPU utilization with a compute problem, when low utilization plus a low concurrency ceiling usually points to KV cache memory capacity.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The KV cache memory pool is too small, so the scheduler cannot admit more concurrent sequences and requests queue while the GPU idles.

Paged KV cache allocates GPU memory in blocks per sequence, and the in-flight batch size is bounded by how many blocks the memory pool can hold. When that pool is too small, the scheduler queues new requests and the GPU sits underutilized, which matches the described symptom. Missing batching, thermal throttling, and GPU tokenization do not explain low compute utilization with a hard concurrency ceiling.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The tokenizer runs on the GPU and competes with the model for SM cycles, capping the batch size.

    Why it's wrong here

    Tokenization is a lightweight CPU-side preprocessing step and does not consume model SMs or set the batch size. Even if tokenization were slow, it would appear as host-side latency, not as a GPU concurrency ceiling tied to KV cache admission. This misattributes the bottleneck to a component that does not govern batching.

  • ✗

    TensorRT-LLM lacks continuous batching support and therefore serializes every request regardless of the KV cache size.

    Why it's wrong here

    TensorRT-LLM does implement continuous batching and paged KV cache; claiming otherwise misstates the product. Serialization would show high per-request latency but would not be caused by a missing feature. The observed low compute utilization points to a memory admission limit rather than absent batching, so this explanation does not fit the evidence.

  • ✗

    The GPU is thermally throttled, which reduces clock speed and therefore limits throughput at high concurrency.

    Why it's wrong here

    Thermal throttling raises compute utilization per unit of work and typically shows elevated temperatures and reduced clocks, but it does not selectively cap concurrency while leaving SMs idle. The scenario reports low compute utilization, which is inconsistent with throttling. This is a hardware hypothesis that does not match the software-level symptom.

  • ✓

    The KV cache memory pool is too small, so the scheduler cannot admit more concurrent sequences and requests queue while the GPU idles.

    Why this is correct

    With paged KV cache, each active sequence consumes blocks from a fixed GPU memory pool. If the pool is undersized, the scheduler must limit the number of in-flight sequences, so additional requests wait even though SM compute is underused. Low compute utilization combined with throughput saturation at low concurrency is the signature of KV cache capacity, not arithmetic throughput.

About these practice questions

One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.