Courseiva

NCP-GENL GPU Acceleration and Optimization Practice Question

An engineer is deploying a large language model using NVIDIA TensorRT-LLM on an A100 GPU. The model uses multi-head attention with a sequence length of 4096. During inference, the GPU's Tensor Cores are underutilized, and the kernel launch overhead is high due to many small operations. Which optimization should be applied to improve Tensor Core utilization and reduce overhead?

⚠ Common exam trap

The trap here is focusing on batch size or precision when the core issue is kernel fragmentation and launch overhead, which require fusion.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable kernel fusion by using TensorRT-LLM's fused multi-head attention plugin.

The underutilization of Tensor Cores and high kernel launch overhead stem from many small operations in multi-head attention. TensorRT-LLM's fused multi-head attention plugin combines these operations into a single optimized kernel, reducing overhead and improving Tensor Core utilization. This is a targeted optimization for Transformer models, especially with long sequences.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the batch size to maximize parallelism and hide latency.

    Why it's wrong here

    Increasing batch size can improve utilization but does not address the root cause of many small kernels. With long sequences, memory bandwidth and kernel overhead dominate. Larger batches may increase memory pressure and not necessarily improve Tensor Core utilization if the kernels remain fragmented.

  • ✓

    Enable kernel fusion by using TensorRT-LLM's fused multi-head attention plugin.

    Why this is correct

    TensorRT-LLM provides fused multi-head attention plugins that combine multiple operations (e.g., QKV projection, attention, and output projection) into a single kernel. This reduces kernel launch overhead and increases arithmetic intensity, allowing better utilization of Tensor Cores. The fusion also minimizes memory traffic, which is critical for long sequences.

  • ✗

    Use NVIDIA Triton Inference Server with dynamic batching to improve GPU utilization.

    Why it's wrong here

    Triton with dynamic batching can improve overall throughput by grouping requests, but it does not address the internal kernel inefficiencies. The problem is within the model execution, not request handling. Dynamic batching may help but is not the direct solution for kernel fusion and Tensor Core utilization.

  • ✗

    Convert the model to FP16 precision to double the Tensor Core throughput.

    Why it's wrong here

    FP16 can improve throughput, but the model may already be in FP16. The issue is underutilization due to kernel fragmentation, not precision. Converting to FP16 alone does not fuse operations or reduce launch overhead, so Tensor Core utilization may remain low.

About these practice questions

One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.