NCP-AIO · domain
Troubleshooting and Optimization
This domain covers diagnosing and tuning GPU-accelerated AI workloads on NVIDIA infrastructure. Questions present exhibits such as MIG or policy JSON, CUDA and NCCL error messages, or profiling output, and ask you to identify root causes and select the correct optimization. Expect scenario-based items spanning multi-GPU scheduling, multi-node collective communication, and inference serving performance.
Focused practice
Practice Troubleshooting and Optimization questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Troubleshooting and Optimization
Be able to read NVIDIA exhibits, logs, and profiler output to pinpoint whether a failure is scheduling, device mapping, collective communication, or data-transfer bound. The single most important thing is matching the observed symptom to the correct layer before proposing a fix.
Diagnosing CUDA errors like invalid device ordinal and mismatched CUDA_VISIBLE_DEVICES or MIG device mapping
Using NCCL debug variables and topology tools such as nccl-tests and nvidia-smi topo to isolate network degradation
Interpreting GPU memory limits, MIG profiles, and scheduler policy JSON to explain resource-limit contention
Optimizing TensorRT inference by reducing host-to-device transfers with pinned memory, batching, and CUDA graphs
Watch out for
Common Troubleshooting and Optimization exam traps
- ▸Assuming sufficient total GPU memory means a job cannot hit limits, ignoring per-device, MIG slice, or policy quotas that cap the specific resource requested.
- ▸Treating CUDA invalid device ordinal as a driver bug instead of checking CUDA_VISIBLE_DEVICES, container device passthrough, and MIG UUID versus index mapping.
- ▸Chasing compute bottlenecks in slow inference when profiling clearly shows host-to-device transfer overhead, which requires pinned memory and batching fixes.
Question index
All Troubleshooting and Optimization questions (86)
Click any question to see the full explanation, or start a practice session above.
During a multi-GPU training job, you notice that one GPU consistently reports lower utilization and longer communication times compared to others. What is the most likely reason for this performance imbalance?
Medium2An AI researcher is using Nsight Systems to profile an application. They notice a large gap in the timeline where neither the CPU nor the GPU is performing significant work. What does this gap most likely represent?
Medium3What is the most accurate way to verify that a training job is utilizing Tensor Cores?
Medium4An AI operations engineer is deploying a model on an NVIDIA GPU and notices that inference latency is higher than expected. The model uses a batch size of 1, and profiling shows that the GPU is idle between kernel launches. Which optimization technique should the engineer use to reduce latency by overlapping data transfer with computation?
Easy5A production inference service using TensorRT is showing lower than expected throughput. Profiling shows that the model is spending significant time in "host-to-device" transfers. What is the best optimization?
Hard6Which TWO of the following actions should be taken to optimize GPU memory usage when encountering Out-of-Memory (OOM) errors during model training?
Hard7A team is deploying a large language model for inference using NVIDIA Triton Inference Server on a GPU. They observe that the first inference request has high latency compared to subsequent requests. What is the most likely cause and the appropriate optimization?
Medium8An AI Operations engineer is managing a multi-node training job using NVIDIA NCCL. The logs indicate frequent 'NCCL WARN' messages related to 'net_ib_init' failures. What is the most likely cause of this issue?
Medium9An AI operations team is troubleshooting a distributed training job on an NVIDIA DGX SuperPOD that uses NCCL for inter-GPU communication. The job intermittently hangs during the all-reduce phase. Which two actions should be taken to diagnose and resolve the issue? (Choose two.)
Hard10A production inference service running on NVIDIA T4 GPUs shows that GPU utilization is consistently below 20% while request latency is high. Profiling with Nsight Systems reveals that the model execution time is short but there are frequent gaps between kernels. Which optimization should be applied first to improve GPU utilization?
Hard11A data scientist reports that a Jupyter notebook running on a GPU-enabled server is extremely slow when training a small neural network, even though nvidia-smi shows the GPU is idle. The notebook uses TensorFlow. Which is the most likely cause?
Easy12When troubleshooting a NCCL collective communication timeout in a distributed training environment, which component should be the primary focus of initial investigation?
Medium13An AI operations engineer is validating a new NVIDIA-certified server before putting it into production. The job runs correctly but the team wants to confirm that the GPUs are operating at the expected clocks and not being limited by power or thermal constraints. Which command provides the most direct evidence of the current power and thermal limits and any throttling reasons?
Easy14An administrator is optimizing a large model training job to reduce checkpointing time to storage. Which strategy is most effective for minimizing the impact on training throughput?
Medium15A multi-tenant NVIDIA GPU-accelerated Kubernetes cluster utilizing NVIDIA AI Enterprise experiences intermittent out-of-memory errors on Triton Inference Server pods despite adequate node memory reservation. Which monitoring and troubleshooting action correctly isolates the root cause?
Hard16An AI administrator is tasked with monitoring GPU utilization in a multi-user cluster. Which tool provides the most granular real-time visibility into process-level GPU memory usage and compute utilization?
Medium17Refer to the exhibit. An administrator notices that 'user_a' is consistently hitting resource limits despite having sufficient total system GPU memory. Based on the policy JSON, what is the cause?
Hard18An inference model running on Triton Inference Server is reporting high latency for requests. The model uses a fixed-size batching strategy. What is the most effective way to optimize throughput while maintaining latency targets?
Hard19Refer to the exhibit. The training job fails with a CUDA OOM error. Given the memory profile, which optimization strategy provides the most immediate relief while maintaining model performance?
Medium20An administrator notices that a specific containerized training job reports high 'GPU Duty Cycle' but low 'Memory Bandwidth Utilization'. What does this pattern indicate about the workload?
Hard21An AI operations engineer is tuning a real-time inference service on NVIDIA A100 GPUs. Profiling with Nsight Systems shows that the GPU is idle for long periods while waiting for input data, and that host-to-device memory copies are frequent and small. The service uses a fixed batch size of 1 and a custom data loader. Which two changes are most likely to improve GPU utilization and reduce inference latency? (Choose two.)
Hard22An AI engineer observes that a model training job on an NVIDIA DGX system is underutilizing the GPU. The monitoring logs show high CPU wait times and low GPU duty cycles. Which action should the engineer take first to resolve the bottleneck?
Medium23Which of the following is the primary indicator of PCIe bus saturation when profiling a training job on an NVIDIA DGX system?
Hard24Refer to the exhibit. What is the most effective way to resolve this specific throttling condition?
Medium25A data scientist reports that a PyTorch training job on an NVIDIA V100 GPU is running slower than expected. The job uses a data loader with num_workers=4. Monitoring shows GPU utilization is around 50%, and CPU usage is high. Which action should an AI operations engineer recommend to improve GPU utilization?
Medium26An AI operations team is troubleshooting a distributed training job on a cluster of NVIDIA DGX A100 systems connected via InfiniBand. The job runs but achieves only 40% of expected scaling efficiency. The team suspects communication bottlenecks. Which two actions should they take to confirm and address the issue? (Choose two.)
Hard27Refer to the exhibit. The training job fails with an OOM error. Which optimization strategy will most effectively resolve this while maintaining model convergence?
Medium28Which TWO of the following actions are recommended for optimizing NVIDIA GPU utilization during a high-concurrency inference deployment?
Medium29An AI operations engineer is investigating a training job on an NVIDIA DGX system that intermittently fails with 'uncorrectable ECC error' on a GPU. The job is using NCCL for multi-GPU communication. The engineer needs to identify the appropriate immediate actions to diagnose and mitigate the issue. (Choose two.)
Hard30A team runs multi-node training with NCCL over InfiniBand on a cluster of DGX systems. Jobs scale well to four nodes but throughput drops sharply at eight nodes, and `nccl-tests` all-reduce bandwidth falls well below line rate at that size. The fabric uses a fat-tree topology with adaptive routing enabled. Which investigation is most likely to reveal the cause?
Hard31An AI engineer needs to monitor GPU utilization across a large cluster of nodes in real-time. Which NVIDIA tool is the most appropriate for this high-level observability task?
Medium32An AI operations engineer notices that a real-time inference service on an NVIDIA T4 GPU has highly variable latency, with occasional spikes to over 100 ms. The service uses TensorRT and runs in a Docker container. Which action should the engineer take to reduce latency variability?
Easy33A data scientist reports that a Jupyter notebook running on a NVIDIA GPU server is extremely slow when executing a deep learning model, even though `nvidia-smi` shows the GPU is idle. The notebook uses TensorFlow. Which action should be taken first to diagnose the issue?
Easy34An operations engineer is troubleshooting a distributed training job that uses NVIDIA Magnum IO GPUDirect Storage to read training data directly from a local NVMe SSD into GPU memory. The job reports lower than expected I/O bandwidth. `nvidia-smi` shows normal GPU utilization, and the NVMe drive's throughput is well below its peak. Which factor is most likely limiting GPUDirect Storage performance in this scenario?
Medium35A team is diagnosing a training job that intermittently stalls for several seconds at the start of each epoch. The job uses a distributed data loader and an NVIDIA DGX system with local NVMe. Monitoring shows GPU utilization dropping to near zero during the stalls while host CPU utilization spikes. Which two actions should the AI operations engineer take to identify and mitigate the stall? (Choose two.)
Hard36An AI operations engineer is optimizing a real-time inference pipeline on an NVIDIA T4 GPU. The pipeline uses TensorRT and receives requests with variable input sizes. Profiling shows that the engine recompiles for each new input shape, causing latency spikes. Which optimization should the engineer apply to eliminate recompilation while maintaining acceptable accuracy?
Medium37During deployment, an AI model experiences high variance in latency during inference. The system uses a fixed instance count. What is the most likely cause for this performance jitter?
Medium38Refer to the exhibit. An engineer is troubleshooting inconsistent training performance across two GPUs in a single DGX node. Why is one GPU reporting a lower clock speed despite being in P0 state?
Hard39A monitoring system reports that a DGX node's GPUs are running at reduced clocks during a long training job, and `nvidia-smi -q -d PERFORMANCE` shows the throttle reason as 'SW Power Cap'. The job's power draw is at the configured limit. Which action is MOST appropriate to restore higher clocks?
Easy40A production inference service on NVIDIA GPUs reports that p99 latency spikes every few minutes while p50 remains stable. Metrics show GPU memory utilization near the limit and periodic `cudaMalloc` calls in the application logs. The model and batch size are fixed. Which change is MOST likely to eliminate the latency spikes?
Hard41An AI operations engineer is troubleshooting a multi-node NCCL training job on an NVIDIA DGX SuperPOD. The job runs but scales poorly: inter-node bandwidth is roughly half of the expected 200 Gb/s per GPU, while intra-node NVLink traffic is at full rate. Running `nvidia-smi topo -m` shows that GPUs in each node are connected to the NICs through the PCIe switch, but the job sets `NCCL_SOCKET_IFNAME` to the management interface. Which action is the most appropriate to resolve the bottleneck?
Medium42An AI engineer observes that a training job on an NVIDIA DGX H100 system is experiencing significant performance degradation. The GPU utilization is high, but the throughput remains low. Which tool should be used first to identify if the bottleneck is related to data loading or PCIe bandwidth saturation?
Medium43A data scientist reports that a Jupyter notebook running on a DGX station cannot allocate GPU memory, even though other users' jobs are running fine. The notebook kernel was started before a system administrator updated the NVIDIA driver and rebooted the node. Which action should the data scientist take to resolve the issue?
Easy44An operations team is troubleshooting a multi-node NCCL training job on an NVIDIA DGX SuperPOD that stalls at initialization and never begins gradient exchange. Running `nccl-tests` with `all_reduce_perf` on the same nodes fails identically, but single-node `all_reduce_perf` succeeds. Which action should the team take first to isolate the fault?
Medium45An AI operations engineer is troubleshooting a multi-GPU training job that fails intermittently with a NCCL error: 'unhandled system error'. The job runs on a DGX-1 with eight V100 GPUs connected via NVLink. Which step should the engineer take first to resolve the issue?
Hard46An operations engineer is investigating a sudden drop in throughput for a multi-GPU training job on an NVIDIA DGX A100. The job uses PyTorch with DDP. Logs show that one GPU is consistently at 100% utilization while others are below 50%. Which tool and approach should be used to identify the bottleneck?
Hard47An AI operations team is troubleshooting a training job that crashes with a segmentation fault after several hours. The job uses multiple GPUs and NCCL for communication. System logs show no errors, but dmesg reveals repeated 'NVRM: Xid' errors. Which action should be taken first to diagnose the issue?
Easy48A machine learning engineer is optimizing a recommendation model for inference on an NVIDIA T4 GPU. The model uses dynamic input shapes, and profiling shows that kernel launch overhead is a significant contributor to latency. Which optimization technique should be applied to reduce this overhead?
Medium49A production inference service on an NVIDIA A100 GPU experiences a gradual increase in latency over several hours, eventually requiring a pod restart. GPU memory utilization climbs steadily, but the model and batch size are fixed. Which action should an AI operations engineer take first to diagnose the root cause?
Medium50An AI operations engineer is troubleshooting a model inference service deployed with NVIDIA Triton Inference Server on a GPU. The service occasionally returns incorrect predictions, and the engineer suspects that the input data is not being preprocessed correctly. The model expects input tensors in FP32 format, but the client is sending FP16 data. Which action should the engineer take to resolve the issue?
Easy51During a training job, the system reports "NCCL WARN" regarding a slow network path. What is the most likely culprit for this performance bottleneck in a multi-node InfiniBand environment?
Medium52A distributed training job on a multi-GPU node is exhibiting poor scaling efficiency: each GPU shows high utilization, but overall throughput increases by only 15% when doubling the number of GPUs. The job uses NCCL for communication. Which diagnostic step is most appropriate to identify the bottleneck?
Medium53An AI engineer is optimizing a real-time inference pipeline on an NVIDIA A100 GPU. The model uses dynamic input shapes, and profiling shows that the GPU spends significant time on memory copies between host and device. Which optimization should be implemented to reduce this overhead?
Medium54An AI operations engineer is investigating an inference service running on NVIDIA Triton Inference Server. Clients report sporadic 500 errors under peak load. The server logs show occasional 'Failed to allocate memory' messages, and `nvidia-smi` shows VRAM nearly full. The service uses dynamic batching with a maximum batch size of 64 and multiple model instances per GPU. Which change is the most appropriate first step to stabilize the service?
Medium55A production inference service running on NVIDIA GPUs exhibits periodic latency spikes every few minutes, correlating with CPU-side stalls and low GPU utilization during those intervals. Profiling with Nsight Systems shows large gaps between kernel launches and frequent cudaMalloc/cudaFree calls. Which action best addresses the root cause?
Medium56A team is profiling a distributed training job using NVIDIA NCCL for inter-GPU communication on a DGX A100 system. They observe that all-reduce operations are taking longer than expected, and the NCCL debug logs show frequent 'NVLS' (NVLink SHARP) errors. Which action should be taken to resolve the issue?
Hard57An AI operations engineer is investigating a training job that exhibits poor scaling efficiency when moving from 8 to 32 GPUs on a DGX SuperPOD. Profiling indicates that the communication time in NCCL all-reduce operations is disproportionately high. Which two actions should be taken to improve scaling efficiency? (Choose two.)
Hard58If a training job on a multi-node cluster shows a significant performance drop during checkpointing, what is the most likely bottleneck?
Medium59Which component of the NVIDIA AI Enterprise stack is primarily responsible for ensuring the long-term stability and compatibility of drivers and libraries across heterogeneous hardware configurations?
Easy60An AI researcher is running a training job using mixed precision (FP16/BF16). The loss function is diverging unexpectedly. What is the most likely culprit?
Medium61An operations team observes that a distributed training job using NVIDIA Collective Communications Library (NCCL) across eight nodes occasionally hangs during the all-reduce phase. Logs show no errors, and the hang resolves only after a node is manually restarted. Which action is MOST appropriate to diagnose the intermittent hang?
Hard62An AI operations team is using NVIDIA DCGM to monitor a cluster of GPUs. They want to set up alerts for when GPUs are running at high temperatures for extended periods. Which DCGM feature should they use?
Medium63An AI operations engineer notices that a training job on an NVIDIA A100 GPU is running slower than expected. Running nvidia-smi shows that the GPU is in 'P0' state but the 'SM Clock' is significantly lower than the maximum boost clock. The job is not memory-bound. Which action should the engineer take first to diagnose the issue?
Easy64A team runs a multi-node NCCL all-reduce training job on four DGX H100 nodes connected by an InfiniBand fabric. Scaling efficiency is poor: throughput barely improves beyond two nodes, and `nvidia-smi` shows NIC transmit counters on each GPU's assigned HCA are far below the PCIe link capacity while GPU compute utilization sits at ~55%. The fabric manager logs report all links as up with no symbol errors. Which action should the administrator take first to diagnose the interconnect bottleneck?
Hard65A system administrator is troubleshooting a 'CUDA error: invalid device ordinal' when launching a job on a multi-GPU system. What is the most likely cause?
Hard66An AI researcher is debugging a multi-node training job using NCCL. Which TWO actions should they take to diagnose potential network-related performance degradation?
Hard67An AI operations engineer is optimizing a PyTorch training job on an NVIDIA DGX A100 system. The job uses a data loader with multiple workers, but the engineer observes that GPU utilization fluctuates between 40% and 60%, and `nvidia-smi dmon` shows periods of zero GPU utilization. The engineer suspects that the data input pipeline is the bottleneck. Which TWO actions should the engineer take to improve GPU utilization? (Choose two.)
Hard68A team trains a model inside an NGC PyTorch container on a DGX H100 node. Training starts, but after a few minutes the process dies and `dmesg` shows `Xid 79: GPU has fallen off the bus` on one GPU. The team needs to determine whether the fault is hardware or software before opening an RMA. Which two actions should they take to gather useful evidence? (Choose two.)
Medium69A team is running a multi-GPU training job on an NVIDIA DGX A100 system using NCCL for inter-GPU communication. Training throughput is much lower than expected, and the NCCL logs show repeated 'NCCL WARN Call to ibv_reg_mr failed' errors. The job uses a container with host networking. Which action should the AI operations engineer take to resolve the issue?
Hard70Refer to the exhibit. During a multi-node training job, communication between nodes fails. What is the most likely cause of this error?
Medium71Refer to the exhibit. An AI administrator investigates why a GPU node is performing significantly slower than expected. Based on the output, what is the most likely cause?
Medium72An inference service runs a 70B parameter model with TensorRT-LLM on a single H100 using in-flight batching. Operators report that time-to-first-token is acceptable, but inter-token latency degrades sharply once concurrent request count exceeds a certain point, and GPU memory utilization sits near 98 percent. Which change most directly addresses the inter-token latency degradation?
Hard73An AI operations engineer is investigating intermittent failures in a long-running distributed training job. The job occasionally aborts with a collective timeout error, but no GPU errors, ECC events, or fabric link flaps appear in logs. Which action should the engineer take first to identify the root cause?
Hard74An engineer is troubleshooting a CUDA program that terminates unexpectedly. Which tool should be used to detect memory leaks and race conditions in the CUDA kernel code?
Medium75An operations team runs a multi-node NCCL all-reduce training job across four DGX nodes connected via InfiniBand. Training throughput is far below the expected linear scaling, and `nvidia-smi` shows GPU utilization oscillating between 20% and 40%. The network fabric is healthy and the GPUs are not thermally throttled. Which diagnostic step is MOST appropriate to identify the bottleneck?
Medium76An AI operations engineer is troubleshooting an inference service on an NVIDIA A100 GPU that shows intermittent stalls. The monitoring dashboard reports GPU utilization at 100%, but request throughput is far below the validated baseline. Running nvidia-smi dmon reveals the SM utilization is high while memory controller utilization is low. Which action should the engineer take first to identify the bottleneck?
Medium77An AI operations team is running a large language model inference service on NVIDIA H100 GPUs using NVIDIA Triton Inference Server. They observe that the first inference request after a period of inactivity takes significantly longer than subsequent requests. The model is loaded and ready, but the GPU shows low utilization during the first request. Which optimization should the team implement to reduce this latency spike?
Hard78Refer to the exhibit. The system has two GPUs. What is the most likely cause of the observed performance discrepancy?
Hard79A team is deploying a large language model for inference using NVIDIA TensorRT-LLM on an H100 GPU. They observe that the first inference request takes several seconds, while subsequent requests are fast. They want to reduce this initial latency. Which technique should they implement?
Hard80An administrator wants to ensure that a training process is limited to a single GPU on a multi-GPU node. Which environment variable should be set?
Hard81An AI engineer observes that a model training job on an NVIDIA DGX system is underutilizing the GPU. The training loop shows frequent "CPU bottleneck" warnings in the logs. Which action should the engineer take first to optimize throughput?
Medium82An AI researcher is deploying a multi-node training job using NCCL on an InfiniBand network. The job is suffering from intermittent latency spikes. Which TWO steps should the engineer perform to troubleshoot the network configuration?
Hard83If a GPU job consistently fails with 'Out of Memory' despite the model size being significantly smaller than the total VRAM, what is the most likely cause?
Medium84An operations engineer notices that an inference container on an A100 is intermittently returning stale predictions after a model update. The container mounts the model directory from a host path, and the update process replaces files in place. Which change most reliably prevents the stale predictions?
Easy85Refer to the exhibit. The training job is showing intermittent "thermal throttling" warnings. Which configuration change is the most appropriate adjustment?
Hard86A distributed training job using PyTorch DDP across eight GPUs on one DGX A100 node shows GPU utilization oscillating between 20 and 40 percent, while `nvidia-smi dmon` shows low SM activity but sustained high memory-controller utilization. The data loader reads from a local NVMe RAID array and applies heavy CPU augmentation. Which action best improves GPU utilization?
MediumOther domains
All NCP-AIO exam domains
Frequently asked questions
- What does the Troubleshooting and Optimization domain cover on the NCP-AIO exam?
- Be able to read NVIDIA exhibits, logs, and profiler output to pinpoint whether a failure is scheduling, device mapping, collective communication, or data-transfer bound. The single most important thing is matching the observed symptom to the correct layer before proposing a fix.
- How many questions are in this domain?
- This page lists all 86 Troubleshooting and Optimization questions in the NCP-AIO question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Troubleshooting and Optimization questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.