NCP-AIO Troubleshooting and Optimization Practice Question
A data scientist reports that a PyTorch training job on an NVIDIA V100 GPU is running slower than expected. The job uses a data loader with num_workers=4. Monitoring shows GPU utilization is around 50%, and CPU usage is high. Which action should an AI operations engineer recommend to improve GPU utilization?
⚠ Common exam trap
The trap here is assuming that GPU-side optimizations like mixed precision or larger batches will help, when the real issue is insufficient data preprocessing throughput.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Increase num_workers in the DataLoader to overlap data preprocessing with GPU computation.
The combination of high CPU usage and low GPU utilization points to a data loading bottleneck where the GPU is starved for data. Increasing the number of DataLoader workers enables more parallel data preprocessing, allowing the GPU to be fed continuously. Other options do not address the CPU-bound pipeline and may not improve the situation.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Pin memory in the DataLoader to speed up host-to-device transfers.
Why it's wrong here
Pinned memory can reduce transfer overhead, but it is a secondary optimization. The high CPU usage suggests that the data loading processes are saturated, and the GPU is idle waiting for batches. Pinning memory alone would not alleviate the CPU bottleneck; increasing workers is the more impactful change.
- ✗
Enable mixed precision training to speed up computations.
Why it's wrong here
Mixed precision can accelerate GPU computations, but it does not address the data loading bottleneck. If the GPU is already underutilized due to waiting for data, making the GPU faster will not help; it may even increase the disparity. The primary issue is CPU-bound preprocessing, not compute precision.
- ✓
Increase num_workers in the DataLoader to overlap data preprocessing with GPU computation.
Why this is correct
When CPU usage is high and GPU utilization is low, the data loading pipeline cannot keep up with the GPU. Increasing num_workers allows more parallel data preprocessing processes, reducing the time the GPU waits for data. This directly addresses the bottleneck and can significantly improve GPU utilization, provided CPU resources are sufficient.
- ✗
Increase the batch size to better utilize the GPU.
Why it's wrong here
Increasing batch size can improve GPU utilization if the bottleneck is small batch size, but here CPU usage is high and GPU utilization is low, indicating a data loading bottleneck. Larger batches would increase memory usage and may not address the CPU-bound preprocessing, potentially worsening the imbalance. The root cause is likely data starvation.
About these practice questions
One of 309 original NCP-AIO practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.