Courseiva
← Back to NVIDIA Certified Professional: AI Operations questions

Scenario-based practice

Troubleshooting Scenario Questions

Practise NVIDIA Certified Professional: AI Operations practice questions — original exam-style scenarios covering every exam domain, with detailed explanations, wrong-answer analysis, and common exam traps.

15
scenario questions
NCP-AIO
exam code
NVIDIA
vendor

Scenario guide

How to approach troubleshooting scenario questions

These questions describe a network symptom and ask you to identify the root cause or the correct fix. They appear across all certification exams and reward systematic thinking over memorisation. The best candidates follow a consistent troubleshooting framework even under time pressure.

Quick answer

Troubleshooting Scenario Questions questions test whether you can apply the concept in context, not just recognise a definition.

How the topic appears in realistic exam-style scenarios.

Which detail in the question changes the correct answer.

How to eliminate plausible but wrong options.

How to connect the question back to the wider exam objective.

Related practice questions

Related NCP-AIO topic practice pages

Scenario questions usually connect to one or more exam topics. Use these links to review the underlying concepts behind the scenario.

Practice set

Practice scenarios

Question 1hardmultiple choice
Full question →

A system administrator is troubleshooting a 'CUDA error: invalid device ordinal' when launching a job on a multi-GPU system. What is the most likely cause?

Question 2mediummultiple choice
Full question →

During the deployment of an AI model training workload on a multi-node cluster, the administrator notices that inter-node communication is significantly slower than expected. Which deployment aspect should be investigated first?

Question 3mediummultiple choice
Full question →

An engineer is troubleshooting a CUDA program that terminates unexpectedly. Which tool should be used to detect memory leaks and race conditions in the CUDA kernel code?

Question 4mediummultiple choice
Full question →

An AI operations engineer is troubleshooting an inference service on an NVIDIA A100 GPU that shows intermittent stalls. The monitoring dashboard reports GPU utilization at 100%, but request throughput is far below the validated baseline. Running nvidia-smi dmon reveals the SM utilization is high while memory controller utilization is low. Which action should the engineer take first to identify the bottleneck?

Question 5hardmulti select
Full question →

An AI researcher is deploying a multi-node training job using NCCL on an InfiniBand network. The job is suffering from intermittent latency spikes. Which TWO steps should the engineer perform to troubleshoot the network configuration?

Question 6mediummultiple choice
Full question →

Refer to the exhibit. An AI administrator investigates why a GPU node is performing significantly slower than expected. Based on the output, what is the most likely cause?

Exhibit

nvidia-smi -q -d PERFORMANCE
Performance State : P0
Clocks Throttle Reasons : Active
  Applications Clocks Setting : None
  SW Power Cap : Active
  HW Slowdown : Active
  HW Thermal Slowdown : Active
Question 7hardmultiple choice
Full question →

An administrator is troubleshooting a performance degradation in a multi-node NVIDIA NCCL-based training job. The job spans four DGX nodes connected via InfiniBand. The administrator suspects that NCCL is not using the optimal network path. Which action should the administrator take to verify and enforce the use of GPUDirect RDMA for inter-node communication?

Question 8mediummultiple choice
Full question →

A media company is deploying an inference service on a Kubernetes cluster with the NVIDIA GPU Operator installed. The service pods remain in Pending with the message that no nodes have the requested nvidia.com/gpu resource, even though the GPUs are healthy and the driver loads correctly on every node. Which troubleshooting step should the engineer perform first?

Question 9mediummulti select
Full question →

An AI operations engineer is troubleshooting a Kubernetes cluster where several GPU training pods fail to start with a device plugin allocation error, even though the nodes report healthy GPUs. The engineer suspects the pods are requesting more GPU resources than a single physical card can provide without a sharing mechanism. Which TWO configurations would legitimately allow multiple pods to consume a single physical GPU on these nodes? (Choose two.)

Question 10mediummultiple choice
Full question →

An AI operations engineer is troubleshooting a multi-node NCCL training job on an NVIDIA DGX SuperPOD. The job runs but scales poorly: inter-node bandwidth is roughly half of the expected 200 Gb/s per GPU, while intra-node NVLink traffic is at full rate. Running `nvidia-smi topo -m` shows that GPUs in each node are connected to the NICs through the PCIe switch, but the job sets `NCCL_SOCKET_IFNAME` to the management interface. Which action is the most appropriate to resolve the bottleneck?

Question 11mediummultiple choice
Full question →

An operations team is troubleshooting a multi-node NCCL training job on an NVIDIA DGX SuperPOD that stalls at initialization and never begins gradient exchange. Running `nccl-tests` with `all_reduce_perf` on the same nodes fails identically, but single-node `all_reduce_perf` succeeds. Which action should the team take first to isolate the fault?

Question 12hardmultiple choice
Full question →

An AI operations engineer is troubleshooting a multi-GPU training job that fails intermittently with a NCCL error: 'unhandled system error'. The job runs on a DGX-1 with eight V100 GPUs connected via NVLink. Which step should the engineer take first to resolve the issue?

Question 13hardmultiple choice
Full question →

A site reliability engineer is troubleshooting a DGX A100 node that intermittently drops out of the cluster during large NCCL all-reduce jobs. `nvidia-smi` shows all eight A100 GPUs healthy, but DCGM reports XID errors 74 and 79 on one GPU during the failures. The engineer needs to determine the most likely cause and the correct administrative action. Which combination best describes the cause and the appropriate first step?

Question 14easymultiple choice
Full question →

An AI operations team is troubleshooting a training job that crashes with a segmentation fault after several hours. The job uses multiple GPUs and NCCL for communication. System logs show no errors, but dmesg reveals repeated 'NVRM: Xid' errors. Which action should be taken first to diagnose the issue?

Question 15easymultiple choice
Full question →

An AI operations engineer is troubleshooting a model inference service deployed with NVIDIA Triton Inference Server on a GPU. The service occasionally returns incorrect predictions, and the engineer suspects that the input data is not being preprocessed correctly. The model expects input tensors in FP32 format, but the client is sending FP16 data. Which action should the engineer take to resolve the issue?

These NCP-AIO practice questions are part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style NCP-AIO questions with detailed explanations, topic-based practice, mock exams, readiness tracking, and study analytics.