NCP-AIO Administration Practice Question
A site reliability engineer is troubleshooting a DGX A100 node that intermittently drops out of the cluster during large NCCL all-reduce jobs. `nvidia-smi` shows all eight A100 GPUs healthy, but DCGM reports XID errors 74 and 79 on one GPU during the failures. The engineer needs to determine the most likely cause and the correct administrative action. Which combination best describes the cause and the appropriate first step?
⚠ Common exam trap
The trap here is treating XID 74 and 79 as generic performance or thermal warnings, when they specifically signal GPU bus loss and NVLink faults that require hardware-level diagnosis.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The errors point to a GPU falling off the bus or an NVLink error; run `nvidia-smi -q` and DCGM diagnostics on the suspect GPU, then consider reseating or replacing the GPU board.
XID 74 and 79 are driver-reported errors indicating a GPU has fallen off the bus or an NVLink error has occurred. These are hardware and link-level faults, so the right first step is to gather detailed GPU and NVLink diagnostics with `nvidia-smi -q` and DCGM, then remediate the hardware. Ignoring or masking these errors risks job failures and data corruption, while treating them as thermal or software issues misses the actual fault domain.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The errors are benign and expected during large all-reduce jobs; suppress them by setting the DCGM health check to ignore XID 74 and 79.
Why it's wrong here
XID 74 and 79 are not benign; they indicate serious GPU or NVLink faults that can cause job failures and silent data corruption. Suppressing them in DCGM health checks would hide a real hardware problem and violate operational best practice. The correct action is to investigate and remediate the fault, not to silence the monitoring that is correctly alerting on a failing component.
- ✗
The errors indicate a thermal shutdown; immediately lower the GPU clock with `nvidia-smi -lgc` and rerun the job.
Why it's wrong here
XID 74 and 79 are not thermal shutdown codes. Thermal issues typically produce XID 61 or clock-throttle events, and lowering clocks may mask a hardware fault rather than diagnose it. The engineer should first gather the full XID context and check the GPU's NVLink and PCIe health, because these XIDs point to a GPU or link-level fault, not a simple thermal condition that clock limiting would resolve.
- ✗
The errors are caused by an outdated NCCL version; upgrade NCCL to the latest release and rerun the job without further hardware checks.
Why it's wrong here
While NCCL version mismatches can cause job failures, XID 74 and 79 are GPU-level hardware and link errors reported by the driver, not software version issues. Upgrading NCCL would not address a GPU falling off the bus or an NVLink error. The engineer must first investigate the hardware state of the suspect GPU, because ignoring these XIDs can lead to data corruption and further node instability.
- ✓
The errors point to a GPU falling off the bus or an NVLink error; run `nvidia-smi -q` and DCGM diagnostics on the suspect GPU, then consider reseating or replacing the GPU board.
Why this is correct
XID 74 and 79 are associated with GPU falling off the bus and NVLink errors respectively. The correct administrative response is to collect detailed GPU and NVLink state with `nvidia-smi -q`, run DCGM diagnostics to isolate the failing GPU or link, and then perform hardware remediation such as reseating the GPU board or replacing it if diagnostics confirm a persistent fault.
About these practice questions
Courseiva writes every NCP-AIO question from scratch — 309 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.