Courseiva
Workload Management →hardMultiple Choice

NCP-AIO Workload Management Practice Question

An administrator manages a shared NVIDIA cluster where several teams run inference services. One team's pods are being evicted repeatedly, and DCGM metrics show the node's GPUs are healthy but memory on the devices is nearly exhausted. The team insists their model fits. Which action should the administrator take FIRST to identify the cause?

⚠ Common exam trap

The trap here is treating GPU memory exhaustion as a hardware fault rather than as contention between co-located workloads.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Inspect the pod specs and running processes for GPU memory held by other containers on the same node.

Healthy GPUs with near-exhausted device memory in a shared cluster usually indicate that another container on the same node is holding memory on the same device. The fastest way to confirm this is to inspect pod specifications and running processes to see which workloads share the GPU. Replacing hardware, changing host memory limits, or tuning telemetry do not address the allocation conflict and delay the correct diagnosis.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the pod's memory limit in the deployment manifest and restart it.

    Why it's wrong here

    The eviction is driven by GPU device memory pressure, not by the container's host memory limit. Raising a CPU or host memory limit does not change how much framebuffer memory the process allocates on the device. The pod would still contend for the same physical GPU memory, so this change would not resolve the eviction and could mask the real co-tenancy problem.

  • ✗

    Replace the affected GPUs and re-run the job on fresh hardware.

    Why it's wrong here

    DCGM already reports the devices as healthy, so hardware replacement is unjustified and disruptive. Memory exhaustion on a healthy device points to allocation behavior, not failure. Swapping GPUs would consume maintenance windows, invalidate the current reproduction, and likely reproduce the same eviction pattern immediately, leaving the actual cause unidentified and the team frustrated.

  • ✗

    Lower the DCGM sampling interval so metrics capture the spike more precisely.

    Why it's wrong here

    Adjusting telemetry frequency changes how finely utilization is recorded but does not alter memory consumption or eviction behavior. Higher-resolution metrics might confirm the spike, yet the administrator already has enough evidence that device memory is exhausted. Time spent tuning sampling delays the diagnostic step of finding which process or neighbor container is holding the memory.

  • ✓

    Inspect the pod specs and running processes for GPU memory held by other containers on the same node.

    Why this is correct

    When device memory is nearly exhausted but hardware is healthy, the most likely cause is co-tenancy: other containers on the same node are holding GPU memory. Because the device plugin allocates whole GPUs by default, multiple pods can land on the same device only when sharing is explicitly enabled, so checking pod specs and active processes reveals whether a neighbor is consuming the memory the team expects to have.

About these practice questions

This NCP-AIO question is part of Courseiva's 309-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.