Courseiva

NCP-AIO Troubleshooting and Optimization Practice Question

An AI operations engineer is troubleshooting an inference service on an NVIDIA A100 GPU that shows intermittent stalls. The monitoring dashboard reports GPU utilization at 100%, but request throughput is far below the validated baseline. Running nvidia-smi dmon reveals the SM utilization is high while memory controller utilization is low. Which action should the engineer take first to identify the bottleneck?

⚠ Common exam trap

The trap here is assuming that 100% GPU utilization always means the GPU is efficiently processing work, when it can indicate stalls or serialization.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Profile the inference process with Nsight Systems to inspect kernel execution and identify serialization gaps.

The high SM utilization with low memory controller utilization suggests the GPU is busy but not doing useful work, often due to kernel serialization or CPU-GPU synchronization stalls. Profiling with Nsight Systems provides the timeline needed to see gaps and serialization. Only after identifying the specific stall should configuration changes be considered, making profiling the correct first step.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Switch the inference precision from FP16 to FP32 to improve numerical stability.

    Why it's wrong here

    Changing precision affects numerical accuracy and may alter performance, but it does not help identify why the GPU is stalled. FP32 typically reduces throughput and increases memory usage, which is counterproductive when the goal is to find the bottleneck. This action is unrelated to the observed serialization symptoms.

  • ✗

    Enable MPS (Multi-Process Service) to allow concurrent kernel execution from multiple processes.

    Why it's wrong here

    MPS is designed for concurrent execution across multiple processes, not for diagnosing a single inference service with serialized kernels. Enabling it here would add complexity without addressing the low memory controller utilization or the stalls. Profiling should come first to confirm whether concurrency is even the issue.

  • ✓

    Profile the inference process with Nsight Systems to inspect kernel execution and identify serialization gaps.

    Why this is correct

    Nsight Systems captures a timeline of CPU and GPU activity, revealing kernel serialization, gaps, and synchronization stalls that inflate utilization without producing throughput. Because memory controller utilization is low, the bottleneck is likely execution serialization or CPU-side latency, which Nsight Systems can pinpoint. This is the correct first step to diagnose the cause before making configuration changes.

  • ✗

    Increase the batch size in the inference server configuration to improve GPU occupancy.

    Why it's wrong here

    Increasing batch size may raise throughput temporarily but does not diagnose the underlying serialization or latency issue. If the GPU is already reporting 100% utilization but low memory activity, blindly increasing batch size can worsen latency without fixing the root cause. This action skips the necessary profiling step and may mask the real problem.

About these practice questions

One of 309 original NCP-AIO practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.