NCP-GENL GPU Acceleration and Optimization Practice Question
A team is training a large language model on a single NVIDIA H100 GPU. They observe that training throughput is significantly lower than expected, and profiling with Nsight Systems shows long periods where the GPU is idle waiting for data. The data loading pipeline uses the default PyTorch DataLoader with num_workers=0 and no pinned memory. Which change is most likely to improve GPU utilization?
⚠ Common exam trap
The trap here is assuming that GPU underutilization always means the model or kernels need optimization, rather than checking whether the input pipeline is starving the device.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable pinned memory and increase num_workers in the DataLoader.
The profiling evidence points to a data-loading bottleneck: the GPU sits idle waiting for batches. Enabling pinned memory and increasing DataLoader workers allows asynchronous, overlapped data transfer and prefetching, keeping the GPU fed. The other options target compute or launch overhead, which are not the limiting factors in this scenario.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use CUDA graphs to capture the training step and reduce kernel launch overhead.
Why it's wrong here
CUDA graphs reduce CPU launch overhead, but the symptom is GPU idle time waiting for data, not launch-bound gaps. If the GPU is starved by slow data loading, capturing the step in a graph will not feed it faster. The data pipeline must be optimized first.
- ✗
Switch from FP32 to TF32 precision for matrix multiplications.
Why it's wrong here
TF32 can accelerate compute-bound kernels, but the profiling data shows the GPU is idle due to data starvation, not compute limits. Changing precision would not reduce the idle periods; the bottleneck remains the input pipeline. It might even introduce accuracy trade-offs without addressing the observed stalls.
- ✗
Increase the batch size to better saturate the GPU.
Why it's wrong here
Increasing batch size may improve utilization only if the GPU is compute-bound, but here profiling shows idle time waiting for data. A larger batch could even worsen the problem by increasing memory pressure and requiring more data per step, potentially lengthening the stalls. The root cause is the data pipeline, not insufficient work per kernel launch.
- ✓
Enable pinned memory and increase num_workers in the DataLoader.
Why this is correct
Pinned memory allows asynchronous host-to-device copies, and multiple worker processes prefetch batches in parallel, overlapping data loading with GPU compute. This directly addresses the idle GPU time observed in profiling. With num_workers=0 and no pinning, data loading is serialized and synchronous, starving the GPU.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.