Courseiva
Software Development →hardMultiple Select

NCA-GENL Software Development Practice Question

A developer is tuning a TensorRT-LLM deployment of a long-context chat model and observes that GPU memory is exhausted under concurrent requests, causing requests to be rejected. They want to reduce KV cache memory pressure without retraining the model. (Choose two.)

⚠ Common exam trap

The trap here is treating increased batching or disabled batching as memory optimizations, when both actually change concurrency in ways that do not reduce the KV cache footprint per sequence.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable paged KV cache so cache blocks are allocated dynamically instead of reserving a contiguous buffer per sequence.

KV cache memory dominates long-context serving, so the effective levers are how cache memory is allocated and how many bytes each cached token consumes. Paged KV cache removes fragmentation by allocating blocks on demand, and KV cache quantization reduces per-token storage precision. Both are runtime or build-time configuration changes that need no retraining, unlike architectural pruning, and both increase the number of concurrent sequences the GPU can hold.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Reduce the model's hidden dimension by pruning attention heads and rebuild the engine from the pruned checkpoint.

    Why it's wrong here

    Pruning attention heads changes the model architecture and requires modifying and retraining or at least recalibrating the checkpoint, which the scenario explicitly excludes. It also risks degrading quality in ways unrelated to the memory goal. The developer needs a runtime memory reduction, not an architectural change to the network, so this option violates the stated constraint.

  • ✓

    Enable paged KV cache so cache blocks are allocated dynamically instead of reserving a contiguous buffer per sequence.

    Why this is correct

    Paged KV cache breaks the cache into blocks allocated on demand and shared through a block manager, which greatly reduces internal and external fragmentation compared with reserving a maximum-length contiguous buffer for every sequence. Under concurrent long-context traffic this raises the number of sequences that fit in the same memory budget, directly relieving the exhaustion the developer observes without any retraining.

  • ✓

    Apply KV cache quantization so cached keys and values are stored in a lower-precision format such as INT8 or FP8.

    Why this is correct

    Quantizing the KV cache stores keys and values at reduced precision, cutting the bytes per cached token roughly in half or better depending on the format. Because the cache is the dominant memory consumer in long-context serving, this directly increases how many concurrent sequences fit in GPU memory. It is a runtime configuration change and requires no retraining of the underlying model weights.

  • ✗

    Disable continuous batching so each request is processed to completion before the next one begins.

    Why it's wrong here

    Disabling continuous batching serializes requests, so each sequence still needs its full cache allocation while others queue. It reduces concurrency rather than memory per sequence and hurts throughput substantially. The exhaustion stems from cache footprint under concurrency, and processing requests one at a time does not shrink that footprint, it merely hides the problem behind a longer queue.

  • ✗

    Increase the maximum number of batched tokens so more requests are processed simultaneously in a single forward pass.

    Why it's wrong here

    Raising the maximum batched tokens enlarges the working set of activations and cache that the engine must hold at once, which worsens memory pressure rather than relieving it. The developer's problem is insufficient cache capacity under concurrency, so expanding the simultaneous workload would push the GPU further past its limit and increase request rejections.

About these practice questions

One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.