Courseiva

NCP-GENL GPU Acceleration and Optimization Practice Question

An engineer is deploying a large language model for real-time inference on an NVIDIA A100 GPU. The model uses FP16 weights but the inference server must handle variable-length input sequences. Profiling shows that the GPU spends significant time on memory-bound operations and that kernel launches are frequent. Which optimization is most appropriate to reduce latency while maintaining accuracy?

⚠ Common exam trap

The trap here is focusing solely on precision reduction as the default latency fix, while overlooking launch overhead and batching strategies that can yield larger gains with no accuracy cost.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use NVIDIA TensorRT with dynamic batching and CUDA graphs to capture the inference step.

TensorRT with dynamic batching and CUDA graphs targets the two identified issues: variable-length sequences are handled by dynamic shapes and batching, while CUDA graphs eliminate launch overhead by capturing the entire inference step. This approach maintains FP16 accuracy and improves latency without the risks of quantization.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the number of CUDA streams to parallelize kernel execution.

    Why it's wrong here

    Multiple streams can overlap independent work, but the inference of a single model with variable-length inputs often has dependencies that limit parallelism. Frequent kernel launches would still incur overhead, and memory-bound operations would not necessarily speed up. Streams add complexity and may not reduce latency for this serialized workload.

  • ✗

    Enable NVIDIA GPUDirect Storage to accelerate data loading.

    Why it's wrong here

    GPUDirect Storage speeds up data transfers between storage and GPU memory, which is irrelevant for inference where inputs are typically already in memory or arrive via network. The bottleneck is within the inference computation itself, not data ingestion from storage. This would not address memory-bound kernels or launch overhead.

  • ✓

    Use NVIDIA TensorRT with dynamic batching and CUDA graphs to capture the inference step.

    Why this is correct

    TensorRT optimizes the graph and supports dynamic shapes for variable-length sequences. Dynamic batching groups requests to improve GPU utilization, while CUDA graphs capture the sequence of kernel launches into a single graph, drastically reducing launch overhead. This combination directly addresses both memory-bound inefficiencies and frequent launches without sacrificing accuracy.

  • ✗

    Convert the model to INT8 using TensorRT with calibration.

    Why it's wrong here

    INT8 quantization can reduce memory bandwidth and increase throughput, but it requires careful calibration to maintain accuracy. For a model already in FP16, moving to INT8 may introduce accuracy loss that is unacceptable for some real-time applications. Moreover, the scenario emphasizes memory-bound operations and frequent kernel launches; INT8 alone does not address launch overhead or variable sequence handling.

About these practice questions

One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.