NCP-AIO Administration Practice Question
An AI operations team is using NVIDIA DCGM (Data Center GPU Manager) to monitor a cluster of A100 GPUs. They want to set up proactive health checks to detect and mitigate GPU issues before they cause job failures. Which two DCGM features should they configure? (Choose two.)
⚠ Common exam trap
The trap here is selecting monitoring metrics like profiling as health checks, when they only provide performance data, not issue detection and remediation.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
DCGM policy management for automated remediation
DCGM health checks with periodic diagnostics and DCGM policy management for automated remediation are the two features that directly enable proactive health monitoring and mitigation. Health checks detect issues, and policies define actions to take when issues are found. Together, they form a proactive health management system. Other options are either monitoring metrics or management features that do not directly address proactive health.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
DCGM policy management for automated remediation
Why this is correct
DCGM policy management allows administrators to define policies that trigger actions when health violations occur, such as resetting a GPU or cordoning a node. This enables automated mitigation, reducing downtime. It works in conjunction with health checks. Configuring policies ensures that detected issues are handled promptly, which is essential for proactive health management in a large GPU cluster.
- ✗
DCGM profiling metrics for real-time utilization
Why it's wrong here
DCGM profiling metrics provide real-time data on GPU utilization, memory bandwidth, and other performance counters. While valuable for performance tuning, they do not directly detect or mitigate hardware issues. They are monitoring metrics, not health checks. The scenario asks for proactive health checks to detect and mitigate GPU issues, so profiling metrics alone are insufficient.
- ✗
DCGM group configuration for multi-node synchronization
Why it's wrong here
DCGM group configuration allows grouping GPUs across nodes for coordinated operations, such as running diagnostics on a group. However, it is a management feature, not a health check. While groups can be used to run health checks, the core features for proactive health are the checks themselves and policies. Group configuration alone does not detect or mitigate issues.
- ✓
DCGM health checks with periodic diagnostics
Why this is correct
DCGM health checks can run periodic diagnostics on GPUs to detect issues like ECC errors, thermal problems, and power anomalies. By configuring these checks, administrators can receive alerts and take corrective action before failures occur. This is a core feature for proactive health monitoring in AI operations, directly addressing the requirement to detect and mitigate GPU issues early.
- ✗
DCGM API integration with Prometheus for alerting
Why it's wrong here
Integrating DCGM with Prometheus allows exporting metrics and setting up alerts based on thresholds. This is useful for monitoring and alerting, but it is not a DCGM feature per se; it is an integration. The scenario asks for DCGM features to configure, and while alerting is part of proactive health, the core DCGM features are health checks and policies. Prometheus integration is an external tool, not a built-in DCGM feature for health checks and mitigation.
About these practice questions
One of 309 original NCP-AIO practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.