Courseiva
Model Deployment →hardMultiple Choice

NCP-GENL Model Deployment Practice Question

Exhibit

config_policy:
- name: "kv_cache_management"
  type: "dynamic"
  memory_limit: 0.8
  eviction_policy: "lru"

Refer to the exhibit. What is the implication of setting the memory_limit to 0.8 in the context of an LLM inference service?

⚠ Common exam trap

Candidates often mistakenly believe the 0.8 memory limit triggers a system-wide shutdown or error, failing to recognize it as a threshold for the LRU eviction policy used in KV cache management.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

It will evict old KV cache entries when 80% limit is reached.

A memory limit of 0.8 indicates that the system will reserve up to 80% of the allocated memory for the KV cache. Once this limit is reached, the Least Recently Used (LRU) policy will begin evicting older sequences to make room for new ones. This helps prevent hard OOM crashes, but users of the evicted sequences will experience errors or forced re-computations when trying to continue their generation tasks.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    It guarantees that the system will never crash due to memory.

    Why it's wrong here

    While setting a memory limit helps prevent complete system OOM, it does not guarantee stability. If the system is under extreme load, it may still encounter issues. Furthermore, evicting active sequences causes logical errors for the end-user, which is a different kind of failure that can disrupt application reliability.

  • ✓

    It will evict old KV cache entries when 80% limit is reached.

    Why this is correct

    The 0.8 setting acts as a cap on the memory footprint of the KV cache. When usage reaches 80% of the assigned memory, the LRU policy triggers the eviction of the least recently used entries, allowing the system to continue operation without crashing, albeit at the cost of losing older sequence data.

  • ✗

    It forces the GPU to run at 80% of its clock speed.

    Why it's wrong here

    The memory limit parameter relates to memory allocation, not clock speed. Clock speed is governed by power and thermal management settings. Attempting to manage memory via clock speed would be ineffective and would only lead to performance degradation without solving the fundamental issue of VRAM capacity management for LLMs.

  • ✗

    It expands the memory capacity by 20% using swap space.

    Why it's wrong here

    The configuration does not interact with system swap space. Swap space is a host-level OS feature that is generally too slow for real-time GPU inference. The memory limit simply constrains the allocation policy within the VRAM already available to the application, preventing it from consuming all available physical device memory.

About these practice questions

This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.