Courseiva

NCP-AIO Troubleshooting and Optimization Practice Question

A team is diagnosing a training job that intermittently stalls for several seconds at the start of each epoch. The job uses a distributed data loader and an NVIDIA DGX system with local NVMe. Monitoring shows GPU utilization dropping to near zero during the stalls while host CPU utilization spikes. Which two actions should the AI operations engineer take to identify and mitigate the stall? (Choose two.)

⚠ Common exam trap

The trap here is treating the stall as a GPU or storage throughput problem and attempting to mask it with larger batches instead of addressing the serialized data loader.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use `nvidia-smi dmon` to capture per-second GPU utilization and power draw while the job runs, correlating the drop with the stall window.

The signature of host CPU spikes with idle GPUs is input pipeline starvation. Prefetching with multiple worker processes and pinned memory keeps the device fed, removing the serialized load at epoch boundaries. Concurrently, `nvidia-smi dmon` supplies the time-series evidence that ties the utilization dips to the stall windows, confirming the diagnosis before and after the fix.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the batch size until the GPU memory is nearly full to amortize the data loading cost.

    Why it's wrong here

    A larger batch may hide the stall by making each step longer, but it does not remove the data loading bottleneck and can cause out-of-memory errors. It also changes convergence behavior, which is undesirable during troubleshooting. The underlying pipeline starvation would remain and could reappear with different data.

  • ✓

    Use `nvidia-smi dmon` to capture per-second GPU utilization and power draw while the job runs, correlating the drop with the stall window.

    Why this is correct

    `nvidia-smi dmon` provides a lightweight time-series view of utilization, memory, and power. Correlating the utilization dip with the stall timestamp confirms whether the GPU is idle waiting for input rather than compute-bound. This evidence distinguishes data pipeline starvation from other causes such as thermal or power throttling.

  • ✗

    Set `CUDA_LAUNCH_BLOCKING=1` to serialize kernel launches and make the stall easier to observe.

    Why it's wrong here

    `CUDA_LAUNCH_BLOCKING=1` forces synchronous kernel execution, which distorts performance and can itself create apparent stalls. It is a debugging aid for pinpointing a failing kernel, not for profiling steady-state throughput. Using it here would confuse the measurement and degrade job performance further.

  • ✓

    Enable the PyTorch data loader with `num_workers` greater than zero and `pin_memory=True` so that batches are prefetched and staged in page-locked host memory.

    Why this is correct

    The near-zero GPU utilization combined with host CPU spikes is characteristic of the data loader starving the GPU. Using multiple worker processes parallelizes decoding and augmentation, and pinning memory allows faster asynchronous host-to-device copies. This directly removes the serialized bottleneck that causes the epoch-start stall.

  • ✗

    Move the dataset from local NVMe to a networked file system to rule out local disk contention.

    Why it's wrong here

    Moving data to a networked file system generally increases latency and reduces throughput compared to local NVMe. The stall pattern points to a serialized loader, not raw disk speed. This change would add network variability and make the diagnosis harder rather than isolating the data pipeline.

About these practice questions

Courseiva writes every NCP-AIO question from scratch — 309 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.