NCP-AIO Administration Practice Question
An administrator is responsible for monitoring a large-scale AI cluster with hundreds of NVIDIA GPUs. They need to collect telemetry data such as GPU utilization, temperature, and power consumption from all nodes and store it centrally for analysis and alerting. Which NVIDIA tool should they use to collect and export GPU metrics to a monitoring system like Prometheus?
⚠ Common exam trap
It's easy for candidates to confuse nvidia-smi with a scalable monitoring solution, when it is only a per-node command-line tool without native Prometheus integration.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
NVIDIA DCGM Exporter
For centralized monitoring of GPU metrics in a large cluster, NVIDIA DCGM Exporter is the appropriate tool. It leverages NVIDIA Data Center GPU Manager (DCGM) to collect a comprehensive set of metrics from each GPU and exposes them via an HTTP endpoint that Prometheus can scrape. This allows administrators to aggregate metrics across all nodes, create dashboards in Grafana, and set up alerts based on thresholds.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
NVIDIA Nsight Systems
Why it's wrong here
NVIDIA Nsight Systems is a performance analysis tool for profiling GPU-accelerated applications, not for cluster-wide telemetry collection. It captures detailed traces of application execution on a single node, and while it can provide deep insights, it is not intended for continuous metric export to Prometheus or centralized monitoring across a large cluster.
- ✓
NVIDIA DCGM Exporter
Why this is correct
NVIDIA DCGM Exporter is a Prometheus exporter that collects GPU metrics using NVIDIA Data Center GPU Manager (DCGM) and exposes them in a format that Prometheus can scrape. It provides a wide range of metrics, including utilization, temperature, power, and memory usage, making it ideal for centralized monitoring of large GPU clusters.
- ✗
nvidia-smi
Why it's wrong here
nvidia-smi is a command-line utility for querying GPU status on a single node, but it does not provide a scalable way to export metrics to a centralized monitoring system. While it can output metrics in XML or CSV, it lacks the built-in integration with Prometheus and would require custom scripting to aggregate data across hundreds of nodes, making it unsuitable for this scenario.
- ✗
NVIDIA Base Command Manager
Why it's wrong here
NVIDIA Base Command Manager is a cluster management and monitoring tool that provides a web interface for viewing GPU metrics, but it is not designed to export metrics to external systems like Prometheus. It offers its own dashboard and alerting, but for integration with a centralized monitoring system, a dedicated exporter like DCGM Exporter is required.
About these practice questions
One of 309 original NCP-AIO practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.