Courseiva

NCP-AIO Troubleshooting and Optimization Practice Question

An AI operations engineer is investigating a training job on an NVIDIA DGX system that intermittently fails with 'uncorrectable ECC error' on a GPU. The job is using NCCL for multi-GPU communication. The engineer needs to identify the appropriate immediate actions to diagnose and mitigate the issue. (Choose two.)

⚠ Common exam trap

The trap here is treating an uncorrectable ECC error as a software or communication issue and attempting to fix it with job restarts or NCCL tuning.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Drain the affected GPU from the scheduler to prevent new jobs from being assigned.

The immediate actions should include checking ECC error counts with nvidia-smi -q to identify the affected GPU, and draining that GPU from the scheduler to prevent further job failures. These steps diagnose the issue and mitigate risk while preserving the ability to perform maintenance. Restarting, resetting, or adjusting NCCL timeouts do not address the underlying hardware error and may delay proper resolution.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the NCCL timeout to allow the job to recover from the ECC error.

    Why it's wrong here

    NCCL timeout settings control communication retries, not hardware ECC errors. Uncorrectable ECC errors are hardware-level faults that cannot be resolved by extending timeouts. Adjusting NCCL parameters would not prevent the job from failing due to the GPU error. The focus should be on diagnosing and isolating the faulty hardware, not on NCCL configuration.

  • ✓

    Drain the affected GPU from the scheduler to prevent new jobs from being assigned.

    Why this is correct

    Draining the GPU prevents additional jobs from being scheduled on a potentially failing device, reducing the risk of further job failures and data corruption. This is a standard operational mitigation for hardware errors. It allows the administrator to perform maintenance or replacement without impacting other workloads, and it is an appropriate immediate action after identifying the affected GPU.

  • ✗

    Restart the training job immediately to see if the error recurs.

    Why it's wrong here

    Restarting the job without diagnosing the ECC error may lead to repeated failures and wasted compute time. Uncorrectable ECC errors often indicate hardware degradation, and simply restarting does not address the root cause. It is better to first gather diagnostic information and then decide on a mitigation strategy, such as draining the GPU or replacing it.

  • ✗

    Use nvidia-smi --gpu-reset to reset the GPU and clear the error state.

    Why it's wrong here

    nvidia-smi --gpu-reset can clear transient errors, but it is not a diagnostic step and may not resolve persistent uncorrectable ECC errors. It also disrupts any running jobs on that GPU. Before resetting, it is important to check ECC counts and logs to understand the nature of the error. Resetting should be a last resort after diagnosis.

  • ✓

    Run nvidia-smi -q to check the ECC error counts and identify the affected GPU.

    Why this is correct

    nvidia-smi -q provides detailed ECC error counts, including correctable and uncorrectable errors, and identifies which GPU is affected. This is a critical first step to confirm the error and isolate the faulty GPU. It helps determine whether the error is persistent or transient, guiding subsequent mitigation steps such as draining the GPU or scheduling maintenance.

About these practice questions

This NCP-AIO question is part of Courseiva's 309-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.