Courseiva

NCP-GENL GPU Acceleration and Optimization Practice Question

Exhibit

Error: Kernel execution timed out after 5000ms.
Potential cause: Excessive occupancy or resource contention.
Action: Adjust block size or shared memory usage.

Refer to the exhibit. An engineer receives this timeout error during a CUDA kernel execution. What is the most appropriate first step to diagnose the resource contention?

⚠ Common exam trap

Candidates frequently recommend restarting the driver or increasing the OS watchdog timeout limit instead of using dedicated profiling tools to inspect kernel resource usage.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use NVIDIA Nsight Compute to profile occupancy.

CUDA kernel timeouts are typically caused by long-running kernels that exceed the GPU's watchdog timer or by excessive resource usage (registers/shared memory) that limits occupancy. Using the NVIDIA Nsight Compute profiler allows the engineer to see the exact occupancy metrics and register pressure, enabling targeted optimizations. This step is crucial for identifying if the kernel is over-provisioned for the specific GPU architecture being targeted.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the block size to 2048 threads.

    Why it's wrong here

    Increasing the block size to 2048 exceeds the hardware limit of 1024 threads per block on modern NVIDIA GPUs. This would cause a launch failure, not resolve a timeout. Furthermore, blindly increasing block sizes can often reduce occupancy due to register pressure and shared memory limits.

  • ✓

    Use NVIDIA Nsight Compute to profile occupancy.

    Why this is correct

    Nsight Compute provides detailed analysis of occupancy, register usage, and shared memory allocation. It identifies whether the kernel is struggling with resource contention, allowing the developer to adjust thread block configuration or refine memory usage to prevent the execution time from exceeding the watchdog timer.

  • ✗

    Disable the watchdog timer in the OS.

    Why it's wrong here

    Disabling the watchdog timer is a dangerous practice that can cause the entire system to freeze if a kernel enters an infinite loop. It masks the symptom rather than fixing the underlying performance bottleneck, which could lead to instability in a production inference environment.

  • ✗

    Switch the kernel to run on the CPU.

    Why it's wrong here

    Moving compute tasks to the CPU is counter-intuitive for GPU optimization. The goal of using NVIDIA GPUs is to leverage massive parallelism. If a kernel is slow, the optimization must occur within the GPU code to maximize performance, not by offloading back to the bottlenecked host.

About these practice questions

One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.