NCP-GENL Model Optimization Practice Question
A team is deploying a Llama 2 13B model with NVIDIA TensorRT-LLM on a single A100 40GB GPU. They need to serve 32 concurrent requests with a maximum sequence length of 4096 tokens. They observe that the GPU runs out of memory during inference. Which configuration parameter should they adjust to control the maximum GPU memory allocated for the KV cache?
⚠ Common exam trap
The trap here is assuming that max_batch_size directly limits KV cache memory, when in fact the KV cache pool is separately governed by kv_cache_free_gpu_memory_fraction.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
kv_cache_free_gpu_memory_fraction
The KV cache in TensorRT-LLM is allocated from a memory pool whose size is controlled by kv_cache_free_gpu_memory_fraction. Lowering this fraction reduces the KV cache footprint, resolving out-of-memory errors when serving multiple long sequences. Other parameters like max_batch_size or max_num_tokens influence scheduling but do not directly bound the KV cache memory pool.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
max_batch_size
Why it's wrong here
max_batch_size determines the maximum number of requests processed in parallel, but it does not directly control the memory pool for the KV cache. Increasing it may worsen memory pressure, while decreasing it reduces throughput. The KV cache memory is separately managed by the TensorRT-LLM runtime and is not bounded by this parameter alone.
- ✗
max_num_tokens
Why it's wrong here
max_num_tokens sets the maximum number of tokens processed in a single iteration, influencing compute scheduling and temporary activation memory, not the persistent KV cache allocation. Adjusting it can affect latency and throughput but does not directly cap the KV cache memory, which is determined by sequence lengths and batch size.
- ✓
kv_cache_free_gpu_memory_fraction
Why this is correct
kv_cache_free_gpu_memory_fraction specifies the fraction of free GPU memory that TensorRT-LLM can use for the KV cache. Lowering this value reduces the KV cache size, preventing out-of-memory errors when serving long sequences with multiple concurrent requests. This parameter directly controls the memory pool dedicated to the KV cache, making it the correct adjustment.
- ✗
max_input_len
Why it's wrong here
max_input_len defines the maximum input sequence length the model can accept. While it impacts KV cache size indirectly by limiting sequence length, it does not control the GPU memory allocation for the KV cache. Reducing it would truncate inputs and degrade quality, whereas the memory issue is best solved by tuning the KV cache memory fraction.
Visual reference
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.