Courseiva

NCP-AIO Troubleshooting and Optimization Practice Question

An AI operations team is using NVIDIA DCGM to monitor a cluster of GPUs. They want to set up alerts for when GPUs are running at high temperatures for extended periods. Which DCGM feature should they use?

⚠ Common exam trap

Test-takers frequently confuse DCGM diagnostics (point-in-time testing) with continuous monitoring and alerting, which requires integration with a metrics platform.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

DCGM health checks with custom thresholds and alerting via DCGM exporter to Prometheus.

DCGM health checks allow defining thresholds for GPU metrics, including temperature. When combined with the DCGM exporter and Prometheus, teams can create alerting rules that trigger on sustained high temperatures. This is the recommended method for cluster-wide monitoring. Other options either misuse DCGM diagnostics, assume non-existent features, or use less scalable methods.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    DCGM's built-in email alerting system configured via dcgm.conf.

    Why it's wrong here

    DCGM does not have a built-in email alerting system. It provides metrics and health checks, but alerting is typically handled by external systems like Prometheus Alertmanager. This option describes a non-existent feature. Administrators must integrate DCGM with a monitoring stack for alerts.

  • ✗

    DCGM diagnostics with the '-r' flag to run a comprehensive test and report temperature violations.

    Why it's wrong here

    DCGM diagnostics are used for validating GPU health at a point in time, not for continuous monitoring. The '-r' flag runs a specific test, but it does not provide ongoing alerting. This option misuses the diagnostics tool for a monitoring task. It would not detect extended periods of high temperature in real-time.

  • ✗

    NVIDIA-smi's '--query-gpu=temperature.gpu' with a cron job to check and send alerts.

    Why it's wrong here

    While nvidia-smi can query temperature, it is not designed for scalable cluster monitoring. Using cron jobs is rudimentary and lacks the flexibility of DCGM's health checks and integration with Prometheus. This approach would be manual and less efficient. It does not leverage DCGM's capabilities for sustained condition alerting.

  • ✓

    DCGM health checks with custom thresholds and alerting via DCGM exporter to Prometheus.

    Why this is correct

    DCGM provides health checks that can be configured with thresholds for temperature and other metrics. The DCGM exporter can expose these metrics to Prometheus, which supports alerting rules based on sustained conditions. This combination allows for proactive monitoring and alerting on high temperatures over time. It is the standard approach for cluster-wide GPU monitoring.

About these practice questions

This NCP-AIO question is part of Courseiva's 309-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.