NCP-AIO Troubleshooting and Optimization Practice Question
Which of the following is the primary indicator of PCIe bus saturation when profiling a training job on an NVIDIA DGX system?
⚠ Common exam trap
Candidates often confuse PCIe saturation with general GPU memory bandwidth bottlenecks. They mistakenly assume high GPU utilization is required to identify a bus issue, ignoring that starvation results in low utilization.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Low GPU duty cycle and high PCIe throughput utilization.
PCIe bus saturation occurs when the GPU's requirement for data exceeds the transfer capacity of the bus. This leads to GPU starvation, where the compute cores idle while waiting for data. Recognizing the specific correlation between low GPU utilization and high PCIe throughput is essential for identifying I/O-bound bottlenecks and optimizing data pipelines for large-scale training jobs.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
High SM occupancy but low memory bandwidth usage.
Why it's wrong here
This pattern indicates a compute-bound workload, not a PCIe bottleneck. When the SM occupancy is high, the GPU is busy with arithmetic. If the memory bandwidth is low, it means the workload is computationally intensive relative to its data needs, which is the opposite of a PCIe bottleneck.
- ✓
Low GPU duty cycle and high PCIe throughput utilization.
Why this is correct
If the GPU duty cycle is low, it means the GPU is waiting for data. If the PCIe throughput utilization is simultaneously high, it indicates that the bus is fully saturated with data transfers, confirming that the GPU is stalled because the data cannot arrive fast enough.
- ✗
High GPU temperature and low clock speeds.
Why it's wrong here
These are symptoms of thermal throttling, not PCIe saturation. When a GPU overheats, the hardware firmware reduces the clock frequency to protect the components. This is a cooling-related hardware issue, not a data communication bottleneck across the PCIe bus architecture.
- ✗
Frequent ECC errors in the GPU memory logs.
Why it's wrong here
ECC errors indicate physical memory degradation or cosmic ray events in the VRAM. They are unrelated to the transfer speeds across the PCIe bus. While ECC errors can lead to performance hits or crashes, they are not a reliable indicator of bus saturation.
About these practice questions
This NCP-AIO question is part of Courseiva's 309-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.