NCP-GENL Model Optimization Practice Question
An enterprise deploying a large language model on an NVIDIA A100 GPU experiences high memory bandwidth bottlenecks during autoregressive token generation. Which optimization technique specifically addresses this memory-bound phase by merging element-wise operations and reducing global memory round-trips?
⚠ Common exam trap
Candidates frequently select generic model pruning or distillation, which reduces parameter count but fails to directly address the specific memory bandwidth bottlenecks caused by repetitive global memory access in autoregressive decoding.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enabling TensorRT-LLM custom kernel fusion for multi-head attention and activation blocks.
Kernel fusion combines multiple successive GPU operations, such as bias additions and activations, into a single CUDA kernel. This drastically reduces high-latency global memory read and write operations, directly mitigating the memory bandwidth bottleneck characteristic of autoregressive transformer decoding phases on NVIDIA hardware.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Applying static INT8 post-training quantization to all linear layers.
Why it's wrong here
Static INT8 post-training quantization reduces model footprint and computes faster integer matrix multiplications, but it introduces conversion overhead and precision loss. It does not fundamentally optimize the kernel launch overhead or fuse element-wise memory operations during the decoding phase.
- ✓
Enabling TensorRT-LLM custom kernel fusion for multi-head attention and activation blocks.
Why this is correct
TensorRT-LLM provides highly optimized, fused CUDA kernels specifically designed to eliminate redundant global memory round-trips for operations like multi-head attention, layer normalization, and activations. This directly accelerates memory-bound autoregressive text generation workloads on NVIDIA GPUs.
- ✗
Increasing the global batch size to maximize arithmetic intensity.
Why it's wrong here
Larger global batches raise arithmetic intensity during training throughput, but autoregressive token generation is inherently sequential and memory-bandwidth-bound per token, so batching cannot merge element-wise kernels or cut global memory round-trips. Kernel fusion techniques such as CUDA graphs or fused operators target that decode phase instead.
- ✗
Switching from FlashAttention to standard vanilla self-attention mechanisms.
Why it's wrong here
Vanilla self-attention materialises the full score matrix, increasing memory traffic and global round-trips, which worsens the memory-bound decode phase rather than relieving it. FlashAttention exists precisely to fuse those operations; reverting to standard attention would suit debugging or compatibility checks, not bandwidth optimisation.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.