An AI operations engineer is troubleshooting an inference service on an NVIDIA A100 GPU that shows intermittent stalls. The monitoring dashboard reports GPU utilization at 100%, but request throughput is far below the validated baseline. Running nvidia-smi dmon reveals the SM utilization is high while memory controller utilization is low. Which action should the engineer take first to identify the bottleneck?
Nsight Systems captures a timeline of CPU and GPU activity, revealing kernel serialization, gaps, and synchronization stalls that inflate utilization without producing throughput. Because memory controller utilization is low, the bottleneck is likely execution serialization or CPU-side latency, which Nsight Systems can pinpoint. This is the correct first step to diagnose the cause before making configuration changes.
Why this answer
The high SM utilization with low memory controller utilization suggests the GPU is busy but not doing useful work, often due to kernel serialization or CPU-GPU synchronization stalls. Profiling with Nsight Systems provides the timeline needed to see gaps and serialization. Only after identifying the specific stall should configuration changes be considered, making profiling the correct first step.
Exam trap
The trap here is assuming that 100% GPU utilization always means the GPU is efficiently processing work, when it can indicate stalls or serialization.