NCP-AIO Troubleshooting and Optimization Practice Question
A machine learning engineer is optimizing a recommendation model for inference on an NVIDIA T4 GPU. The model uses dynamic input shapes, and profiling shows that kernel launch overhead is a significant contributor to latency. Which optimization technique should be applied to reduce this overhead?
⚠ Common exam trap
The trap here is assuming that mixed precision or larger batch sizes reduce launch overhead, but they target different bottlenecks and do not minimize the number of CPU-GPU interactions.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Convert the model to TensorRT with dynamic shapes and enable CUDA graphs during inference.
CUDA graphs reduce kernel launch overhead by capturing a sequence of operations and replaying them as a single graph. TensorRT with dynamic shapes can leverage CUDA graphs to maintain flexibility while minimizing overhead. This directly addresses the profiling finding, making it the most effective optimization for the described scenario.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Convert the model to TensorRT with dynamic shapes and enable CUDA graphs during inference.
Why this is correct
CUDA graphs capture a sequence of kernel launches and replay them with a single launch, drastically reducing launch overhead. TensorRT supports dynamic shapes and can build engines that use CUDA graphs. This is ideal for models with variable input shapes where launch overhead is high. The combination reduces CPU overhead and improves latency.
- ✗
Use mixed precision (FP16) to reduce the number of kernels executed.
Why it's wrong here
Mixed precision reduces memory bandwidth and can speed up computations, but it does not reduce the number of kernel launches. Kernel launch overhead is related to the CPU-GPU interaction, not the precision. While FP16 may improve performance, it does not directly target the overhead described.
- ✗
Enable NVIDIA MPS (Multi-Process Service) to allow concurrent kernel execution.
Why it's wrong here
MPS allows multiple processes to share the GPU and can improve utilization in multi-process scenarios, but it does not reduce launch overhead for a single process. The overhead is per-kernel launch from the CPU, and MPS does not change that. This option is irrelevant to the specific problem of launch overhead in a single model inference.
- ✗
Increase the batch size to amortize kernel launch overhead across more samples.
Why it's wrong here
Increasing batch size can improve throughput but does not reduce the per-kernel launch overhead; it merely spreads it over more samples. For latency-sensitive inference with dynamic shapes, this may increase latency. The issue is the overhead itself, not the number of samples, so this does not address the root cause.
About these practice questions
One of 309 original NCP-AIO practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.