NCP-AIO Workload Management Practice Question
Which strategy is most effective for managing heterogeneous GPU clusters containing both older architectures (e.g., V100) and newer architectures (e.g., H100)?
⚠ Common exam trap
Many candidates incorrectly suggest using only automated scheduling policies like priority classes, forgetting that heterogeneous hardware requires explicit node-level constraints to prevent incompatible jobs from failing at runtime.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Using Node Selectors and Taints/Tolerations.
Using Kubernetes Node Selectors and Taints/Tolerations allows administrators to route specific workloads to compatible hardware. This is essential because newer GPUs support features like MIG or specific precision formats that older hardware lacks. By properly tagging nodes based on architecture, AI Ops prevents 'Incompatible Device' errors and ensures that high-performance training jobs are scheduled on hardware that meets their specific computational requirements, optimizing cluster-wide performance and efficiency.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enabling uniform scheduling across all nodes.
Why it's wrong here
Treating all nodes as uniform in a heterogeneous cluster leads to scheduling failures. Workloads designed for H100 features like Transformer Engine will fail on V100 nodes. Uniform scheduling fails to account for architectural differences, leading to frequent job crashes and a poor developer experience for end-user data scientists.
- ✓
Using Node Selectors and Taints/Tolerations.
Why this is correct
Node Selectors and Taints/Tolerations are the standard mechanisms for enforcing hardware compatibility in Kubernetes. They allow operators to isolate workloads to specific GPU architectures, ensuring that code requiring newer hardware features is only placed on nodes capable of supporting them, while older jobs can safely run on legacy hardware.
- ✗
Configuring the NVIDIA Device Plugin to hide older GPUs.
Why it's wrong here
Hiding hardware is not a standard practice and would result in wasted infrastructure capacity. The goal should be to utilize all available hardware by routing appropriate workloads to the correct nodes, rather than making existing resources invisible to the scheduler, which negatively impacts overall cluster utilization and ROI.
- ✗
Deploying a single unified container image for all jobs.
Why it's wrong here
While container images can be compiled to support multiple architectures, this often bloats the image size and does not solve the underlying hardware capability mismatch. Even with a universal image, the scheduling must still be architecture-aware to ensure the hardware supports the required tensor operations and memory features.
About these practice questions
One of 309 original NCP-AIO practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.