Courseiva

NCP-GENL GPU Acceleration and Optimization Practice Question

An engineer is optimizing a Transformer-based LLM for inference on an NVIDIA A100 GPU. The model uses FP16 precision, but during generation, the GPU's Tensor Cores are underutilized, and latency is higher than expected. Profiling reveals that many small matrix multiplications are executed sequentially. Which technique is most effective to improve Tensor Core utilization and reduce latency?

⚠ Common exam trap

The trap here is focusing on quantization or CUDA graphs, which address different bottlenecks, rather than recognizing that small sequential operations require fusion to improve Tensor Core efficiency.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Apply kernel fusion to combine multiple small operations into a single kernel.

The underutilization of Tensor Cores stems from many small, sequential matrix multiplications. Kernel fusion combines these into larger operations, allowing Tensor Cores to process more data per instruction and reducing kernel launch overhead. This directly improves utilization and lowers latency, making it the most effective technique among the options for this specific bottleneck.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Enable CUDA graphs to capture and replay the sequence of kernel launches.

    Why it's wrong here

    CUDA graphs reduce CPU launch overhead by capturing a sequence of kernels and replaying them with a single launch. This can help if the workload is launch-bound, but the profiling indicates small matrix multiplications are the issue. CUDA graphs do not fuse kernels or improve Tensor Core utilization for small ops; they only streamline launch. Therefore, it is less effective than fusion for this scenario.

  • ✗

    Use NVIDIA TensorRT to quantize the model to INT8.

    Why it's wrong here

    INT8 quantization can accelerate inference and reduce memory bandwidth, but it may degrade accuracy and requires calibration. More importantly, it does not directly address the underutilization caused by small sequential matrix multiplications. While TensorRT can apply optimizations like fusion, the primary benefit here is not quantization itself. Thus, it is not the most effective immediate technique for the described bottleneck.

  • ✗

    Increase the batch size to amortize kernel launch overhead.

    Why it's wrong here

    Increasing batch size can improve throughput by amortizing overhead, but in latency-sensitive scenarios like real-time generation, it increases latency per token. Moreover, the issue is small matrix multiplications that may not fully utilize Tensor Cores even with larger batches. The core problem is the sequential execution of small ops, so batching alone does not address the underutilization.

  • ✓

    Apply kernel fusion to combine multiple small operations into a single kernel.

    Why this is correct

    Kernel fusion merges multiple small matrix multiplications and element-wise operations into a single kernel, reducing launch overhead and enabling better Tensor Core utilization by processing larger, combined workloads. This is especially effective in Transformer inference where many small GEMMs occur. By fusing operations, the GPU can execute them more efficiently, lowering latency and increasing utilization of Tensor Cores.

About these practice questions

Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.