NCP-AIO Troubleshooting and Optimization Practice Question
An AI operations team is using NVIDIA DCGM to monitor a cluster of GPUs. They want to set up alerts for when GPUs are running at high temperatures for extended periods. Which DCGM feature should they use?
⚠ Common exam trap
Test-takers frequently confuse DCGM diagnostics (point-in-time testing) with continuous monitoring and alerting, which requires integration with a metrics platform.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
DCGM health checks with custom thresholds and alerting via DCGM exporter to Prometheus.
DCGM health checks allow defining thresholds for GPU metrics, including temperature. When combined with the DCGM exporter and Prometheus, teams can create alerting rules that trigger on sustained high temperatures. This is the recommended method for cluster-wide monitoring. Other options either misuse DCGM diagnostics, assume non-existent features, or use less scalable methods.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
DCGM's built-in email alerting system configured via dcgm.conf.
Why it's wrong here
DCGM does not have a built-in email alerting system. It provides metrics and health checks, but alerting is typically handled by external systems like Prometheus Alertmanager. This option describes a non-existent feature. Administrators must integrate DCGM with a monitoring stack for alerts.
- ✗
DCGM diagnostics with the '-r' flag to run a comprehensive test and report temperature violations.
Why it's wrong here
DCGM diagnostics are used for validating GPU health at a point in time, not for continuous monitoring. The '-r' flag runs a specific test, but it does not provide ongoing alerting. This option misuses the diagnostics tool for a monitoring task. It would not detect extended periods of high temperature in real-time.
- ✗
NVIDIA-smi's '--query-gpu=temperature.gpu' with a cron job to check and send alerts.
Why it's wrong here
While nvidia-smi can query temperature, it is not designed for scalable cluster monitoring. Using cron jobs is rudimentary and lacks the flexibility of DCGM's health checks and integration with Prometheus. This approach would be manual and less efficient. It does not leverage DCGM's capabilities for sustained condition alerting.
- ✓
DCGM health checks with custom thresholds and alerting via DCGM exporter to Prometheus.
Why this is correct
DCGM provides health checks that can be configured with thresholds for temperature and other metrics. The DCGM exporter can expose these metrics to Prometheus, which supports alerting rules based on sustained conditions. This combination allows for proactive monitoring and alerting on high temperatures over time. It is the standard approach for cluster-wide GPU monitoring.
About these practice questions
This NCP-AIO question is part of Courseiva's 309-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.