NCP-AIO Administration Practice Question
An administrator is managing an NVIDIA DGX SuperPOD used for large-scale AI training. The cluster uses a Slurm workload manager. The administrator needs to ensure that jobs are scheduled only on nodes with healthy GPUs and that failed GPUs are automatically drained from the pool. Which integration should be configured to achieve this?
⚠ Common exam trap
The trap here is assuming that Base Command Manager or the Container Toolkit alone can handle GPU health monitoring; they manage deployment and resource allocation, but DCGM is specifically designed for health checks and integration with schedulers like Slurm.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
NVIDIA Data Center GPU Manager (DCGM) with the Slurm health check plugin
DCGM with the Slurm health check plugin is the correct integration because it continuously monitors GPU health and can automatically drain nodes with failed GPUs from Slurm, ensuring jobs run only on healthy hardware. Other options lack the necessary health monitoring and automatic remediation capabilities.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
NVIDIA Base Command Manager with Slurm integration
Why it's wrong here
Base Command Manager is a cluster management tool that can deploy and manage Slurm, but it does not inherently provide GPU health monitoring or automatic node draining based on GPU failures. While it simplifies cluster provisioning, the specific health-check and drain behavior requires DCGM or similar monitoring integration, which is not the primary function of Base Command Manager alone.
- ✗
NVIDIA Fleet Command with Slurm edge scheduling
Why it's wrong here
Fleet Command is a cloud-based platform for managing edge AI deployments, not for on-premises Slurm clusters. It does not integrate with Slurm for GPU health monitoring or node draining. Using it in a DGX SuperPOD context would be inappropriate and would not fulfill the requirement of automatic draining based on GPU health.
- ✓
NVIDIA Data Center GPU Manager (DCGM) with the Slurm health check plugin
Why this is correct
DCGM provides comprehensive GPU health monitoring, and its integration with Slurm via the health check plugin allows automatic detection of unhealthy GPUs. When DCGM identifies a failed GPU, the plugin can drain the node from Slurm, preventing new jobs from being scheduled on it. This directly addresses the requirement to schedule only on healthy GPUs and automatically remove failed ones.
- ✗
NVIDIA Container Toolkit with Slurm's --gres flag
Why it's wrong here
The NVIDIA Container Toolkit enables containers to access GPUs, and Slurm's --gres flag requests GPU resources for jobs. However, this combination does not monitor GPU health or automatically drain nodes with failed GPUs. It only handles resource allocation; health monitoring and automatic remediation require additional components like DCGM.
About these practice questions
One of 309 original NCP-AIO practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.