Courseiva
Administration →easyMultiple Choice

NCP-AIO Administration Practice Question

An AI operations team needs to monitor GPU health and utilization across a fleet of DGX nodes from a single dashboard. They want per-GPU metrics such as power, temperature, utilization, and ECC errors, and they want to retain historical data for capacity planning. Which NVIDIA tool is purpose-built to collect and expose these GPU telemetry metrics for centralized monitoring?

⚠ Common exam trap

Many exam-takers confuse a profiler or a manual command-line utility with a continuous telemetry pipeline, when only DCGM with its exporter provides centralized, historical GPU health metrics.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

NVIDIA DCGM (Data Center GPU Manager) with the DCGM exporter for Prometheus.

DCGM is NVIDIA's purpose-built data center GPU monitoring and management tool. It gathers health and performance metrics including power, temperature, utilization, and ECC errors, and the DCGM exporter makes them available to Prometheus for centralized dashboards and historical retention. This aligns exactly with the team's need for fleet-wide GPU telemetry and capacity planning data.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    NVIDIA DCGM (Data Center GPU Manager) with the DCGM exporter for Prometheus.

    Why this is correct

    DCGM is the NVIDIA tool designed for data center GPU monitoring and management. It collects health, utilization, power, temperature, and ECC metrics, and the DCGM exporter exposes them to Prometheus for centralized dashboards and long-term retention. This directly matches the requirement for fleet-wide GPU telemetry and historical capacity planning data.

  • ✗

    NVIDIA Base Command Manager's job scheduler logs.

    Why it's wrong here

    Base Command Manager logs relate to cluster provisioning, job scheduling, and workload management, not to per-GPU telemetry such as power, temperature, or ECC errors. They do not expose GPU health metrics in a form suitable for a centralized monitoring dashboard. Using scheduler logs for GPU telemetry would miss the detailed hardware metrics that DCGM is designed to collect and export.

  • ✗

    NVIDIA Nsight Systems for profiling GPU kernels and collecting timeline traces.

    Why it's wrong here

    Nsight Systems is a profiling tool for analyzing application performance and kernel timelines, not a fleet-wide telemetry and monitoring solution. It does not continuously collect power, temperature, or ECC metrics from many nodes, nor does it provide a centralized dashboard with historical retention. It is used for deep performance analysis of specific workloads, not for operational monitoring of a GPU fleet.

  • ✗

    NVIDIA CUDA Toolkit's `nvidia-smi` command run manually on each node.

    Why it's wrong here

    `nvidia-smi` provides a point-in-time view of GPU state on a single node, but it is not a centralized telemetry system. Running it manually does not aggregate metrics across a fleet, does not retain history, and does not feed a dashboard. While useful for quick checks and troubleshooting, it is not the purpose-built solution for continuous fleet-wide GPU monitoring and capacity planning.

About these practice questions

Courseiva writes every NCP-AIO question from scratch — 309 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.