NCP-AIO Troubleshooting and Optimization Practice Question
An AI engineer observes that a training job on an NVIDIA DGX H100 system is experiencing significant performance degradation. The GPU utilization is high, but the throughput remains low. Which tool should be used first to identify if the bottleneck is related to data loading or PCIe bandwidth saturation?
⚠ Common exam trap
Candidates often select NVIDIA-SMI or Nsight Compute, failing to realize that Nsight Systems is the correct tool for identifying system-wide bottlenecks between the CPU, data pipeline, and GPU.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
NVIDIA Nsight Systems
NVIDIA Nsight Systems is the primary tool for analyzing system-wide performance, including CPU-GPU interactions and data transfer bottlenecks. Identifying whether a bottleneck occurs in the data pipeline or within the hardware interconnects is critical for optimizing training speed. By visualizing timelines, engineers can pinpoint if the GPU is starving for data or if the PCIe bus is congested, enabling targeted remediation for large-scale distributed training clusters.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
NVIDIA-SMI
Why it's wrong here
NVIDIA-SMI provides real-time monitoring of power, temperature, and basic utilization metrics. While useful for high-level health checks, it lacks the deep trace capabilities required to analyze data loading bottlenecks or PCIe saturation patterns. It cannot visualize the kernel execution timeline or I/O wait times effectively.
- ✗
NVIDIA DCGM
Why it's wrong here
DCGM is designed for cluster-wide health monitoring, diagnostics, and telemetry collection at scale. While it can detect hardware errors or overheating, it is not optimized for tracing the granular execution path or data loading latency issues that typically cause performance degradation in specific training jobs.
- ✓
NVIDIA Nsight Systems
Why this is correct
Nsight Systems provides a system-wide view of CPU and GPU activities, allowing engineers to correlate data transfer operations with kernel execution. This visibility is essential for identifying stalls, data starvation, or PCIe bus congestion, which are the most common causes of low throughput despite high GPU utilization.
- ✗
NVIDIA Nsight Compute
Why it's wrong here
Nsight Compute is an interactive kernel profiler designed for deep analysis of individual GPU kernels. It focuses on instruction-level performance, warp occupancy, and memory access patterns within the GPU. It is not intended for high-level troubleshooting of data loading pipelines or system-wide PCIe traffic analysis.
About these practice questions
This NCP-AIO question is part of Courseiva's 309-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.