Courseiva

NCP-GENL GPU Acceleration and Optimization Practice Question

An engineer is profiling a CUDA kernel and notices that the achieved occupancy is low, leading to underutilization of the GPU. The kernel uses a large number of registers per thread, limiting the number of resident warps. Which optimization should be attempted first to improve occupancy?

⚠ Common exam trap

The trap here is assuming that increasing block size or enabling cache automatically improves occupancy, when the real constraint is the number of registers per thread limiting resident warps.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use the __launch_bounds__ qualifier to limit registers per thread.

The low occupancy is caused by high register usage per thread. Using __launch_bounds__ instructs the compiler to limit registers, allowing more warps to be resident and improving occupancy. This directly addresses the bottleneck, whereas other options do not target register pressure and thus would not effectively improve occupancy.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use the __launch_bounds__ qualifier to limit registers per thread.

    Why this is correct

    The __launch_bounds__ qualifier allows the programmer to specify the minimum number of blocks per multiprocessor, which guides the compiler to limit register usage. By reducing registers per thread, more warps can be resident, improving occupancy. This is a direct way to address register-limited occupancy without changing the algorithm, though it may cause spilling if overused.

  • ✗

    Enable L1 cache to reduce memory latency.

    Why it's wrong here

    Enabling L1 cache can improve performance for memory-bound kernels by reducing latency, but it does not affect register usage or occupancy. The problem is that registers limit the number of resident warps, so cache optimization is orthogonal. Thus, it is not the first optimization to attempt for occupancy.

  • ✗

    Increase the block size to allow more warps per block.

    Why it's wrong here

    Increasing block size does not reduce register usage per thread and may even worsen occupancy if register pressure remains high. The limit is the number of registers per thread, which restricts the total number of warps that can reside. Therefore, simply changing block size without addressing register usage is unlikely to improve occupancy significantly.

  • ✗

    Convert the kernel to use shared memory for data reuse.

    Why it's wrong here

    Using shared memory can reduce global memory accesses and improve performance, but it does not directly reduce register pressure. While it may free some registers if global loads are replaced, the impact is indirect and not guaranteed. The primary issue is register-limited occupancy, so a more direct approach like launch bounds is preferable.

About these practice questions

One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.