Courseiva
Workload Management →mediumMultiple Choice

NCP-AIO Workload Management Practice Question

Which strategy is most effective for managing heterogeneous GPU clusters containing both older architectures (e.g., V100) and newer architectures (e.g., H100)?

⚠ Common exam trap

Many candidates incorrectly suggest using only automated scheduling policies like priority classes, forgetting that heterogeneous hardware requires explicit node-level constraints to prevent incompatible jobs from failing at runtime.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Using Node Selectors and Taints/Tolerations.

Using Kubernetes Node Selectors and Taints/Tolerations allows administrators to route specific workloads to compatible hardware. This is essential because newer GPUs support features like MIG or specific precision formats that older hardware lacks. By properly tagging nodes based on architecture, AI Ops prevents 'Incompatible Device' errors and ensures that high-performance training jobs are scheduled on hardware that meets their specific computational requirements, optimizing cluster-wide performance and efficiency.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Enabling uniform scheduling across all nodes.

    Why it's wrong here

    Treating all nodes as uniform in a heterogeneous cluster leads to scheduling failures. Workloads designed for H100 features like Transformer Engine will fail on V100 nodes. Uniform scheduling fails to account for architectural differences, leading to frequent job crashes and a poor developer experience for end-user data scientists.

  • ✓

    Using Node Selectors and Taints/Tolerations.

    Why this is correct

    Node Selectors and Taints/Tolerations are the standard mechanisms for enforcing hardware compatibility in Kubernetes. They allow operators to isolate workloads to specific GPU architectures, ensuring that code requiring newer hardware features is only placed on nodes capable of supporting them, while older jobs can safely run on legacy hardware.

  • ✗

    Configuring the NVIDIA Device Plugin to hide older GPUs.

    Why it's wrong here

    Hiding hardware is not a standard practice and would result in wasted infrastructure capacity. The goal should be to utilize all available hardware by routing appropriate workloads to the correct nodes, rather than making existing resources invisible to the scheduler, which negatively impacts overall cluster utilization and ROI.

  • ✗

    Deploying a single unified container image for all jobs.

    Why it's wrong here

    While container images can be compiled to support multiple architectures, this often bloats the image size and does not solve the underlying hardware capability mismatch. Even with a universal image, the scheduling must still be architecture-aware to ensure the hardware supports the required tensor operations and memory features.

About these practice questions

One of 309 original NCP-AIO practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-AIO practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-AIO exam.