NCP-GENL GPU Acceleration and Optimization Practice Question
An engineer is training a large language model with pipeline parallelism across four NVIDIA GPUs. They observe that GPU utilization is low and training throughput is limited by idle time during pipeline bubbles. Which technique is most effective to reduce pipeline bubbles and improve utilization?
⚠ Common exam trap
Watch out — candidates often confuse memory optimization techniques with those that address pipeline scheduling inefficiencies, such as increasing micro-batches.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Increase the number of micro-batches to keep the pipeline stages busy.
Pipeline bubbles are idle periods when stages wait for data. Increasing the number of micro-batches allows more fine-grained overlap of computation across stages, keeping all GPUs busy and reducing idle time. This is a direct and effective way to improve pipeline utilization.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Increase the number of micro-batches to keep the pipeline stages busy.
Why this is correct
Pipeline bubbles occur when stages wait for data from previous stages. Increasing the number of micro-batches allows more overlapping of forward and backward passes across stages, filling the pipeline and reducing idle time. This is a standard technique in pipeline parallelism to improve utilization.
- ✗
Reduce the batch size to decrease the memory footprint and allow more concurrent kernels.
Why it's wrong here
Reducing batch size lowers memory usage but also reduces the amount of work per stage, potentially worsening pipeline utilization. It does not address the fundamental issue of pipeline bubbles, which are caused by dependencies between stages, not by memory pressure.
- ✗
Enable gradient checkpointing to trade compute for memory and increase batch size.
Why it's wrong here
Gradient checkpointing reduces activation memory, allowing larger batches, but it does not directly reduce pipeline bubbles. While larger batches can help, the primary fix for bubbles is increasing micro-batches or using interleaved scheduling, not just memory optimization.
- ✗
Switch from pipeline parallelism to data parallelism across the four GPUs.
Why it's wrong here
Data parallelism replicates the model on each GPU and splits the batch, but it requires the model to fit on a single GPU. For a large language model that needs pipeline parallelism, switching to data parallelism is often infeasible due to memory constraints and does not directly address pipeline bubbles.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.