NCP-AIO Troubleshooting and Optimization Practice Question
A production inference service on an NVIDIA A100 GPU experiences a gradual increase in latency over several hours, eventually requiring a pod restart. GPU memory utilization climbs steadily, but the model and batch size are fixed. Which action should an AI operations engineer take first to diagnose the root cause?
⚠ Common exam trap
The trap here is assuming that nvidia-smi memory monitoring is sufficient to diagnose leaks, when it only shows aggregate usage without per-process or allocation detail.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use PyTorch's torch.cuda.memory_summary() or TensorFlow's memory profiler to capture allocation snapshots and compare over time.
The steady memory increase with fixed workload strongly suggests a memory leak in the inference application or framework. Framework-specific memory profilers provide allocation-level visibility, allowing engineers to identify unreleased tensors or cached allocations. Aggregate GPU monitoring or sanitizer tools lack the necessary granularity. Increasing limits merely postpones failure.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use PyTorch's torch.cuda.memory_summary() or TensorFlow's memory profiler to capture allocation snapshots and compare over time.
Why this is correct
Framework-level memory profilers show exactly which tensors or operations are allocating memory and whether those allocations are freed. By taking snapshots at intervals, an engineer can see if memory is retained across inference calls, indicating a leak in the application code or framework caching allocator. This directly identifies the source without disrupting the service.
- ✗
Run nvidia-smi --query-gpu=memory.used --format=csv -l 1 to log memory usage over time and correlate with request rate.
Why it's wrong here
This command only samples aggregate GPU memory usage at one-second intervals. It cannot identify which process or allocation is leaking, nor can it distinguish between model weights, activations, or framework caching. Without per-process or heap-level detail, an engineer cannot pinpoint the source of the growth, making this a weak first diagnostic step.
- ✗
Enable CUDA memory leak detection with compute-sanitizer --tool memcheck on the running inference process.
Why it's wrong here
Compute-sanitizer memcheck detects out-of-bounds and illegal memory accesses, not gradual memory leaks in a long-running process. It also requires launching the application under sanitizer control, which adds significant overhead and is impractical for a live production service. It would not reveal the steady accumulation of unreleased memory.
- ✗
Increase the GPU memory limit in the Kubernetes pod spec to prevent the pod from being killed.
Why it's wrong here
Raising the memory limit only delays the inevitable restart and does not address the underlying leak. It may also mask the problem and lead to node-level memory exhaustion affecting other workloads. The goal is to diagnose and fix the root cause, not to accommodate unbounded memory growth.
About these practice questions
This NCP-AIO question is part of Courseiva's 309-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.