NCP-GENL GPU Acceleration and Optimization Practice Question
An engineer is deploying a large language model for real-time inference on an NVIDIA A100 GPU. The model uses FP16 weights but the inference server must handle variable-length input sequences. Profiling shows that the GPU spends significant time on memory-bound operations and that kernel launches are frequent. Which optimization is most appropriate to reduce latency while maintaining accuracy?
⚠ Common exam trap
The trap here is focusing solely on precision reduction as the default latency fix, while overlooking launch overhead and batching strategies that can yield larger gains with no accuracy cost.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use NVIDIA TensorRT with dynamic batching and CUDA graphs to capture the inference step.
TensorRT with dynamic batching and CUDA graphs targets the two identified issues: variable-length sequences are handled by dynamic shapes and batching, while CUDA graphs eliminate launch overhead by capturing the entire inference step. This approach maintains FP16 accuracy and improves latency without the risks of quantization.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the number of CUDA streams to parallelize kernel execution.
Why it's wrong here
Multiple streams can overlap independent work, but the inference of a single model with variable-length inputs often has dependencies that limit parallelism. Frequent kernel launches would still incur overhead, and memory-bound operations would not necessarily speed up. Streams add complexity and may not reduce latency for this serialized workload.
- ✗
Enable NVIDIA GPUDirect Storage to accelerate data loading.
Why it's wrong here
GPUDirect Storage speeds up data transfers between storage and GPU memory, which is irrelevant for inference where inputs are typically already in memory or arrive via network. The bottleneck is within the inference computation itself, not data ingestion from storage. This would not address memory-bound kernels or launch overhead.
- ✓
Use NVIDIA TensorRT with dynamic batching and CUDA graphs to capture the inference step.
Why this is correct
TensorRT optimizes the graph and supports dynamic shapes for variable-length sequences. Dynamic batching groups requests to improve GPU utilization, while CUDA graphs capture the sequence of kernel launches into a single graph, drastically reducing launch overhead. This combination directly addresses both memory-bound inefficiencies and frequent launches without sacrificing accuracy.
- ✗
Convert the model to INT8 using TensorRT with calibration.
Why it's wrong here
INT8 quantization can reduce memory bandwidth and increase throughput, but it requires careful calibration to maintain accuracy. For a model already in FP16, moving to INT8 may introduce accuracy loss that is unacceptable for some real-time applications. Moreover, the scenario emphasizes memory-bound operations and frequent kernel launches; INT8 alone does not address launch overhead or variable sequence handling.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.