NCP-GENL Model Deployment Practice Question
Exhibit
config_policy: - name: "kv_cache_management" type: "dynamic" memory_limit: 0.8 eviction_policy: "lru"
Refer to the exhibit. What is the implication of setting the memory_limit to 0.8 in the context of an LLM inference service?
⚠ Common exam trap
Candidates often mistakenly believe the 0.8 memory limit triggers a system-wide shutdown or error, failing to recognize it as a threshold for the LRU eviction policy used in KV cache management.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
It will evict old KV cache entries when 80% limit is reached.
A memory limit of 0.8 indicates that the system will reserve up to 80% of the allocated memory for the KV cache. Once this limit is reached, the Least Recently Used (LRU) policy will begin evicting older sequences to make room for new ones. This helps prevent hard OOM crashes, but users of the evicted sequences will experience errors or forced re-computations when trying to continue their generation tasks.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
It guarantees that the system will never crash due to memory.
Why it's wrong here
While setting a memory limit helps prevent complete system OOM, it does not guarantee stability. If the system is under extreme load, it may still encounter issues. Furthermore, evicting active sequences causes logical errors for the end-user, which is a different kind of failure that can disrupt application reliability.
- ✓
It will evict old KV cache entries when 80% limit is reached.
Why this is correct
The 0.8 setting acts as a cap on the memory footprint of the KV cache. When usage reaches 80% of the assigned memory, the LRU policy triggers the eviction of the least recently used entries, allowing the system to continue operation without crashing, albeit at the cost of losing older sequence data.
- ✗
It forces the GPU to run at 80% of its clock speed.
Why it's wrong here
The memory limit parameter relates to memory allocation, not clock speed. Clock speed is governed by power and thermal management settings. Attempting to manage memory via clock speed would be ineffective and would only lead to performance degradation without solving the fundamental issue of VRAM capacity management for LLMs.
- ✗
It expands the memory capacity by 20% using swap space.
Why it's wrong here
The configuration does not interact with system swap space. Swap space is a host-level OS feature that is generally too slow for real-time GPU inference. The memory limit simply constrains the allocation policy within the VRAM already available to the application, preventing it from consuming all available physical device memory.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.