Courseiva
Administration →mediumMultiple Choice

NCP-AIO Administration Practice Question

An administrator is managing an NVIDIA DGX SuperPOD used for large-scale AI training. The cluster uses a Slurm workload manager. The administrator needs to ensure that jobs are scheduled only on nodes with healthy GPUs and that failed GPUs are automatically drained from the pool. Which integration should be configured to achieve this?

⚠ Common exam trap

The trap here is assuming that Base Command Manager or the Container Toolkit alone can handle GPU health monitoring; they manage deployment and resource allocation, but DCGM is specifically designed for health checks and integration with schedulers like Slurm.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

NVIDIA Data Center GPU Manager (DCGM) with the Slurm health check plugin

DCGM with the Slurm health check plugin is the correct integration because it continuously monitors GPU health and can automatically drain nodes with failed GPUs from Slurm, ensuring jobs run only on healthy hardware. Other options lack the necessary health monitoring and automatic remediation capabilities.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    NVIDIA Base Command Manager with Slurm integration

    Why it's wrong here

    Base Command Manager is a cluster management tool that can deploy and manage Slurm, but it does not inherently provide GPU health monitoring or automatic node draining based on GPU failures. While it simplifies cluster provisioning, the specific health-check and drain behavior requires DCGM or similar monitoring integration, which is not the primary function of Base Command Manager alone.

  • ✗

    NVIDIA Fleet Command with Slurm edge scheduling

    Why it's wrong here

    Fleet Command is a cloud-based platform for managing edge AI deployments, not for on-premises Slurm clusters. It does not integrate with Slurm for GPU health monitoring or node draining. Using it in a DGX SuperPOD context would be inappropriate and would not fulfill the requirement of automatic draining based on GPU health.

  • ✓

    NVIDIA Data Center GPU Manager (DCGM) with the Slurm health check plugin

    Why this is correct

    DCGM provides comprehensive GPU health monitoring, and its integration with Slurm via the health check plugin allows automatic detection of unhealthy GPUs. When DCGM identifies a failed GPU, the plugin can drain the node from Slurm, preventing new jobs from being scheduled on it. This directly addresses the requirement to schedule only on healthy GPUs and automatically remove failed ones.

  • ✗

    NVIDIA Container Toolkit with Slurm's --gres flag

    Why it's wrong here

    The NVIDIA Container Toolkit enables containers to access GPUs, and Slurm's --gres flag requests GPU resources for jobs. However, this combination does not monitor GPU health or automatically drain nodes with failed GPUs. It only handles resource allocation; health monitoring and automatic remediation require additional components like DCGM.

About these practice questions

One of 309 original NCP-AIO practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.