Courseiva
Workload Management →mediumMultiple Choice

NCP-AIO Workload Management Practice Question

An AI operations engineer manages a Kubernetes cluster running the NVIDIA GPU Operator. A team wants its long-running inference deployment to be automatically rescheduled if the GPU on a node develops an uncorrectable error that the device plugin or health checks detect. The team also wants the node to stop accepting new GPU pods until the issue is resolved. Which combination of behaviors should the engineer rely on to meet these requirements?

⚠ Common exam trap

The trap here is expecting pod-level constructs like liveness probes or disruption budgets to handle hardware faults, when GPU error isolation is driven by device health checks and node taints.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The GPU Operator's health checks mark the GPU unhealthy, the device plugin stops advertising it, and Kubernetes taints or cordons the affected node so pods are rescheduled and no new GPU pods land there.

Meeting both requirements needs health-driven device withdrawal plus node-level scheduling exclusion. The GPU Operator's health checks detect uncorrectable errors, the device plugin stops advertising the faulty GPU, and the node is tainted or cordoned. Controllers then recreate pods on healthy nodes, and the taint keeps new GPU pods away until the fault is cleared.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    A Horizontal Pod Autoscaler scales up replicas so healthy copies absorb traffic while the faulty node remains in service.

    Why it's wrong here

    Autoscaling adds replicas but leaves the faulty GPU node in the schedulable pool, so new pods may still be placed on it and existing ones are not relocated. It masks the failure rather than resolving it and does not prevent future placement on the bad device. The node would continue advertising a GPU that cannot be trusted.

  • ✓

    The GPU Operator's health checks mark the GPU unhealthy, the device plugin stops advertising it, and Kubernetes taints or cordons the affected node so pods are rescheduled and no new GPU pods land there.

    Why this is correct

    The GPU Operator runs health checks that can detect uncorrectable GPU errors and signal the device plugin to stop advertising the faulty device. The operator can also apply taints or cordon the node. Existing pods become unschedulable and are recreated elsewhere by their controller, while new GPU pods are kept off the node, matching both stated requirements.

  • ✗

    A liveness probe on the inference container restarts the pod on the same node when the GPU error causes a request failure.

    Why it's wrong here

    A liveness probe restarts the container in place, typically on the same node with the same faulty GPU, so the pod may keep failing. It does not reschedule the workload to healthy hardware and does not stop the node from accepting new GPU pods. This addresses symptoms rather than isolating the failed device.

  • ✗

    A PodDisruptionBudget on the inference deployment forces the scheduler to migrate pods off the node when the GPU fails.

    Why it's wrong here

    A PodDisruptionBudget governs voluntary disruptions like drains; it does not detect hardware faults or trigger rescheduling on its own. Without an external actor evicting the pods, the budget does nothing. It also does not prevent new GPU pods from being placed on the failing node, so it cannot satisfy the second requirement.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

About these practice questions

One of 309 original NCP-AIO practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.