Courseiva

NCP-AIO Troubleshooting and Optimization Practice Question

A production inference service running on NVIDIA GPUs exhibits periodic latency spikes every few minutes, correlating with CPU-side stalls and low GPU utilization during those intervals. Profiling with Nsight Systems shows large gaps between kernel launches and frequent cudaMalloc/cudaFree calls. Which action best addresses the root cause?

⚠ Common exam trap

The trap here is assuming that reducing kernel execution time or adding CPU threads will fix latency spikes caused by launch overhead, rather than addressing the orchestration inefficiency itself.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable CUDA Graph capture for the inference sequence and reuse the captured graph for repeated executions.

The profile shows CPU-side stalls and frequent cudaMalloc/cudaFree, indicating launch and allocation overhead. CUDA Graphs capture the entire sequence and replay it with minimal CPU involvement, eliminating both the allocation churn and the launch gaps. This directly reduces latency spikes and improves GPU utilization, making it the most effective remedy for the described symptoms.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Switch the model to use TensorRT with FP16 precision to reduce kernel execution time.

    Why it's wrong here

    Reducing kernel execution time with FP16 may lower average latency but does not eliminate the periodic CPU stalls and allocation overhead. The gaps between kernels would still cause spikes. While TensorRT can help overall performance, it is not the targeted fix for the CPU-side orchestration issue shown in the profile.

  • ✓

    Enable CUDA Graph capture for the inference sequence and reuse the captured graph for repeated executions.

    Why this is correct

    CUDA Graphs capture a sequence of kernel launches and memory operations into a single graph, then replay it with minimal CPU overhead. This eliminates repeated cudaMalloc/cudaFree and launch gaps, smoothing latency spikes. It directly targets the CPU-side stalls and low GPU utilization observed, making it the appropriate optimization for this scenario.

  • ✗

    Increase the number of worker threads on the CPU to better overlap preprocessing with GPU execution.

    Why it's wrong here

    Adding CPU worker threads may improve preprocessing throughput but does not address the repeated cudaMalloc/cudaFree calls or kernel launch gaps. Those are GPU-side orchestration overheads. More threads could even increase contention and memory allocation churn, worsening the latency spikes. The root cause remains the launch and allocation overhead.

  • ✗

    Enable Multi-Process Service (MPS) to allow concurrent kernel execution from multiple processes.

    Why it's wrong here

    MPS is designed for concurrent execution across processes, not for reducing launch overhead within a single process. In this scenario, the service likely runs in one process, and MPS would not address the cudaMalloc/cudaFree churn or launch gaps. It could introduce additional complexity without solving the latency spikes.

About these practice questions

Courseiva writes every NCP-AIO question from scratch — 309 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.