NCP-AIO Administration Practice Question
An administrator is configuring NVIDIA Base Command Manager to manage a cluster of DGX nodes. They want to ensure that when a node's GPU temperature exceeds a defined threshold, the node is automatically drained and an alert is sent to the operations team. Which combination of Base Command Manager features should the administrator configure to achieve this?
⚠ Common exam trap
Watch out — candidates often confuse Base Command Manager's native health and alerting features with external monitoring tools like Prometheus or Kubernetes node problem detector.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Health checks with a custom script that triggers a node drain via the Base Command Manager API, and an alert rule that sends a notification.
Base Command Manager provides health checks that can execute custom scripts on nodes. A script can monitor GPU temperature and invoke the Base Command Manager API to drain the node, while an alert rule notifies the team. This leverages native features for automated remediation and alerting, unlike external monitoring stacks or Kubernetes-specific tools.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Health checks with a custom script that triggers a node drain via the Base Command Manager API, and an alert rule that sends a notification.
Why this is correct
Base Command Manager health checks can run custom scripts periodically. A script can query GPU temperature and, if over threshold, call the Base Command Manager API to drain the node. An alert rule then sends a notification. This provides automated remediation and alerting as required.
- ✗
NVIDIA GPU Operator with node problem detector, and a Kubernetes taint that evicts pods.
Why it's wrong here
The GPU Operator and node problem detector are Kubernetes-centric tools, not Base Command Manager features. Base Command Manager is a separate cluster management solution. Using Kubernetes taints would not integrate with Base Command Manager's native alerting and drain mechanisms.
- ✗
Prometheus Alertmanager with a webhook that drains the node, and Grafana dashboards for visualization.
Why it's wrong here
Prometheus and Grafana are monitoring tools, not native Base Command Manager features. While they can be integrated, the scenario asks for Base Command Manager features. The built-in health checks and alerting are more direct and do not require external components.
- ✗
Base Command Manager job scheduler with a preemption policy that kills jobs when temperature is high.
Why it's wrong here
While Base Command Manager includes a job scheduler, preemption policies are designed for resource contention, not hardware health. Killing jobs does not drain the node or send alerts. The health check and alerting features are the correct tools for temperature-based remediation.
About these practice questions
This NCP-AIO question is part of Courseiva's 309-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.