Courseiva
Administration →hardMultiple Choice

NCP-AIO Administration Practice Question

An administrator is configuring NVIDIA Base Command Manager to manage a cluster of DGX nodes. They want to ensure that when a node's GPU temperature exceeds a defined threshold, the node is automatically drained and an alert is sent to the operations team. Which combination of Base Command Manager features should the administrator configure to achieve this?

⚠ Common exam trap

Watch out — candidates often confuse Base Command Manager's native health and alerting features with external monitoring tools like Prometheus or Kubernetes node problem detector.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Health checks with a custom script that triggers a node drain via the Base Command Manager API, and an alert rule that sends a notification.

Base Command Manager provides health checks that can execute custom scripts on nodes. A script can monitor GPU temperature and invoke the Base Command Manager API to drain the node, while an alert rule notifies the team. This leverages native features for automated remediation and alerting, unlike external monitoring stacks or Kubernetes-specific tools.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Health checks with a custom script that triggers a node drain via the Base Command Manager API, and an alert rule that sends a notification.

    Why this is correct

    Base Command Manager health checks can run custom scripts periodically. A script can query GPU temperature and, if over threshold, call the Base Command Manager API to drain the node. An alert rule then sends a notification. This provides automated remediation and alerting as required.

  • ✗

    NVIDIA GPU Operator with node problem detector, and a Kubernetes taint that evicts pods.

    Why it's wrong here

    The GPU Operator and node problem detector are Kubernetes-centric tools, not Base Command Manager features. Base Command Manager is a separate cluster management solution. Using Kubernetes taints would not integrate with Base Command Manager's native alerting and drain mechanisms.

  • ✗

    Prometheus Alertmanager with a webhook that drains the node, and Grafana dashboards for visualization.

    Why it's wrong here

    Prometheus and Grafana are monitoring tools, not native Base Command Manager features. While they can be integrated, the scenario asks for Base Command Manager features. The built-in health checks and alerting are more direct and do not require external components.

  • ✗

    Base Command Manager job scheduler with a preemption policy that kills jobs when temperature is high.

    Why it's wrong here

    While Base Command Manager includes a job scheduler, preemption policies are designed for resource contention, not hardware health. Killing jobs does not drain the node or send alerts. The health check and alerting features are the correct tools for temperature-based remediation.

About these practice questions

This NCP-AIO question is part of Courseiva's 309-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.