Courseiva
Model Optimization →hardMultiple Choice

NCP-GENL Model Optimization Practice Question

Exhibit

Error: [TRT-LLM] KV Cache block allocation failed. Current allocation: 80% capacity. Request rejected to prevent OOM. Optimization status: KV_CACHE_ENABLED=True, PAGED_ATTENTION=False.

Refer to the exhibit. The deployment is facing memory allocation errors during peak load. Based on the error log, what is the most effective configuration change to resolve the issue while keeping the model architecture constant?

⚠ Common exam trap

Candidates often suggest increasing GPU VRAM or reducing batch sizes, which are reactive measures. They overlook PagedAttention, which is the specific architectural solution for KV cache fragmentation and memory OOM errors.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable PagedAttention in the TensorRT-LLM runtime.

The error log indicates that the system is running out of memory because the KV cache is allocated statically, which often leads to fragmentation or over-allocation. Enabling PagedAttention is the standard NVIDIA-recommended solution for this scenario. It allows the system to manage KV cache memory dynamically, reclaiming space from finished requests and efficiently allocating blocks for new tokens, thereby preventing the out-of-memory (OOM) errors caused by static allocation.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Switch to a smaller model size.

    Why it's wrong here

    Switching to a smaller model changes the architecture, which was ruled out by the requirement to keep the model constant. Furthermore, simply reducing the model size does not fix the fundamental issue of memory fragmentation in the KV cache, which is the root cause of the observed OOM errors.

  • ✗

    Increase the GPU clock frequency.

    Why it's wrong here

    Increasing the GPU clock frequency improves raw compute performance but does nothing to solve memory allocation failures. OOM errors are a result of memory capacity exhaustion or fragmentation, not computational speed. Adjusting clock speeds will likely increase power consumption and heat without addressing the memory management limitation shown.

  • ✓

    Enable PagedAttention in the TensorRT-LLM runtime.

    Why this is correct

    PagedAttention is designed to solve exactly this type of KV cache allocation failure. By moving from static allocation to paged block allocation, the system can utilize memory more efficiently and pack more requests into the same GPU memory footprint, effectively eliminating the OOM errors caused by inflexible memory management.

  • ✗

    Reduce the batch size to one.

    Why it's wrong here

    Reducing the batch size to one will lower memory usage, but it effectively cripples the throughput of the inference service. This is a workaround, not an optimization. Enabling PagedAttention is a superior technical solution because it allows for high throughput while maintaining stable memory management, rather than sacrificing performance.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

About these practice questions

This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.