Courseiva

CCNA Workload Management Questions

12 of 87 questions · Page 2/2 · Workload Management topic · Answers revealed

76
MCQmedium

What is the primary function of an 'InitContainer' in an NVIDIA GPU-enabled pod deployment?

A.To serve as a secondary GPU compute thread during training.
B.To ensure prerequisites are satisfied before the main application starts.
C.To permanently store the model weights after training.
D.To bypass GPU scheduling constraints and force-load the job.
AnswerB

InitContainers provide a controlled sequence of operations. By running before the main application container, they allow for critical setup tasks like verifying library installations, checking driver compatibility, or prepping datasets, which ensures that the training job starts in a valid, predictable state, minimizing late-stage failures.

Why this answer

InitContainers are often used to ensure that environment-specific requirements—such as configuring GPU drivers, validating libraries, or mounting persistent volumes—are met before the main training application starts. In AI workloads, this is critical for ensuring that dependencies are correctly loaded or data is staged, preventing runtime failures that would occur if the main training process attempted to execute in an incomplete environment. This pattern enhances the robustness of automated AI workflows.

Exam trap

Candidates mistakenly think InitContainers handle the main application training logic or continuously monitor runtime performance throughout the pod lifecycle.

77
MCQeasy

A financial services firm must prove to auditors that an AI training job ran on hardware located only in its Frankfurt data center and that no pod could ever be scheduled onto GPUs in other regions. The cluster spans three regions with nodes labeled topology.kubernetes.io/region. Which approach most directly enforces this placement requirement?

A.Define a nodeAffinity rule in the pod spec requiring topology.kubernetes.io/region to equal eu-central-1, and mark it requiredDuringSchedulingIgnoredDuringExecution.
B.Add a preferredDuringSchedulingIgnoredDuringExecution nodeAffinity rule that favors the Frankfurt region with a high weight.
C.Set the pod's restartPolicy to Never and add a toleration for the region-specific taint on Frankfurt nodes.
D.Create a PodDisruptionBudget for the training job that limits voluntary evictions to zero during the run.
AnswerA

A required node affinity rule is a hard constraint: the kube-scheduler will only place the pod on nodes whose labels match, and it leaves the pod Pending rather than scheduling it elsewhere. Pinning the region label guarantees the training job runs exclusively in the Frankfurt nodes, which is the enforceable evidence auditors need.

Why this answer

Hard placement guarantees come from required node affinity, which makes matching labels a precondition for scheduling. Because the scheduler will not place the pod on any node whose region label differs, the job cannot leave the Frankfurt data center, giving the firm an enforceable, auditable control over GPU location.

Exam trap

The trap here is confusing soft preferred affinity, which the scheduler may ignore under pressure, with required affinity, which is a hard constraint that keeps the pod Pending instead of violating the rule.

78
MCQmedium

Which strategy is most effective for managing heterogeneous GPU clusters containing both older architectures (e.g., V100) and newer architectures (e.g., H100)?

A.Enabling uniform scheduling across all nodes.
B.Using Node Selectors and Taints/Tolerations.
C.Configuring the NVIDIA Device Plugin to hide older GPUs.
D.Deploying a single unified container image for all jobs.
AnswerB

Node Selectors and Taints/Tolerations are the standard mechanisms for enforcing hardware compatibility in Kubernetes. They allow operators to isolate workloads to specific GPU architectures, ensuring that code requiring newer hardware features is only placed on nodes capable of supporting them, while older jobs can safely run on legacy hardware.

Why this answer

Using Kubernetes Node Selectors and Taints/Tolerations allows administrators to route specific workloads to compatible hardware. This is essential because newer GPUs support features like MIG or specific precision formats that older hardware lacks. By properly tagging nodes based on architecture, AI Ops prevents 'Incompatible Device' errors and ensures that high-performance training jobs are scheduled on hardware that meets their specific computational requirements, optimizing cluster-wide performance and efficiency.

Exam trap

Many candidates incorrectly suggest using only automated scheduling policies like priority classes, forgetting that heterogeneous hardware requires explicit node-level constraints to prevent incompatible jobs from failing at runtime.

79
MCQmedium

A platform team runs an NVIDIA AI cluster with the GPU Operator deployed. Users submit jobs directly with kubectl and frequently request whole GPUs even when their notebooks only need a fraction of one. The team wants Kubernetes itself to admit and queue jobs based on GPU demand without users changing their manifests. Which component should the team deploy to meet this requirement?

A.NVIDIA MIG Manager configured for every GPU
B.NVIDIA GPU Operator with the device plugin enabled
C.NVIDIA DCGM Exporter with Prometheus alerting
D.NVIDIA KAI Scheduler
AnswerD

KAI Scheduler is NVIDIA's Kubernetes-native scheduler for AI workloads. It plugs into the cluster as a secondary scheduler and performs gang scheduling, queueing, and GPU fraction accounting, so jobs that request partial GPUs or exceed current capacity wait in a queue instead of being rejected. Because admission and queueing happen inside the scheduler, users keep submitting standard manifests and do not need to learn new tooling.

Why this answer

The requirement is for Kubernetes to make admission and queueing decisions based on GPU demand while users submit normal manifests. KAI Scheduler is the NVIDIA component that provides queue-based, gang-aware scheduling for AI workloads and understands fractional GPU requests, so jobs wait rather than fail. The GPU Operator, MIG Manager, and DCGM Exporter each address installation, partitioning, or observability rather than scheduling admission.

Exam trap

The trap here is assuming the GPU Operator itself performs scheduling or admission, when it only installs and manages the GPU software stack.

80
MCQmedium

An organization is migrating AI workloads to a private cloud. Which feature is essential for ensuring that GPU resources are dynamically reclaimed and reallocated to different departments without manual intervention?

A.Manual pod scheduling with affinity rules.
B.Static GPU reservation per user.
C.Automated cluster autoscaling and priority-based scheduling.
D.Hard-coding node IP addresses in the configuration files.
AnswerC

Automated autoscaling and priority-based scheduling allow the platform to dynamically adjust capacity based on real-time demand. High-priority workloads can preempt lower-priority tasks, ensuring critical research proceeds, while idle resources are automatically reclaimed and made available to other users, maximizing the ROI of the NVIDIA hardware investment.

Why this answer

Dynamic resource scheduling, often implemented via Kubernetes schedulers or custom job orchestrators, is essential for multi-tenant environments. By utilizing features like auto-scaling, preemption, and resource quotas, the system can automatically reclaim idle GPUs from one department and reallocate them to another. This automation maximizes hardware utility and ensures that expensive GPU infrastructure is never left sitting idle due to administrative delays.

Exam trap

Test-takers frequently suggest static allocation schemes or manual administrative interventions, missing the core requirement for automated cluster autoscaling and priority-based scheduling to dynamically reclaim resources.

81
MCQmedium

A DevOps engineer needs to monitor GPU health metrics in real-time for workload management. Which tool provides the most granular visibility into GPU utilization and power consumption for individual containers?

A.Kubernetes Horizontal Pod Autoscaler (HPA) with CPU metrics.
B.NVIDIA DCGM Exporter.
C.Standard Linux 'top' command.
D.Docker stats command.
AnswerB

The DCGM Exporter provides comprehensive, hardware-specific metrics that are essential for deep visibility into GPU performance. It enables the capture of precise utilization data per container, which allows administrators to make data-driven decisions regarding resource allocation, capacity planning, and the optimization of AI training and inference workloads.

Why this answer

The NVIDIA DCGM (Data Center GPU Manager) Exporter is the industry-standard tool for collecting fine-grained GPU telemetry. By integrating with Prometheus and Grafana, it allows administrators to track metrics like power usage, temperature, and utilization at the container level. This granular data is vital for proactive workload management, allowing for autoscaling based on actual hardware demand rather than generic CPU metrics.

Exam trap

Candidates frequently choose generic Kubernetes metrics or standard Prometheus exporters. These tools often lack the specific hooks needed to read deep GPU hardware registers like power draw and memory bandwidth.

82
MCQeasy

A data science team submits a PyTorch training job to a Kubernetes cluster where the NVIDIA GPU Operator is installed. The pod stays in Pending state, and kubectl describe shows the message '0/6 nodes are available: 6 Insufficient nvidia.com/gpu.' The administrator confirms the nodes have healthy GPUs and the device plugin pods are Running. What is the most likely cause?

A.The pod requested more nvidia.com/gpu resources than any single node can advertise
B.The container image does not include the NVIDIA CUDA base layer
C.The pod lacks a nodeSelector matching the GPU node label nvidia.com/gpu.present
D.The GPU Operator's driver container has not yet built the kernel module on the nodes
AnswerA

Extended resources such as nvidia.com/gpu are integer-quantized per node and cannot be oversubscribed across nodes. If the pod requests a count exceeding what any single node advertises, the scheduler reports insufficient nvidia.com/gpu on every node and the pod stays Pending. Reducing the request to fit one node resolves the condition.

Why this answer

Extended resources like nvidia.com/gpu are advertised per node and cannot be aggregated across nodes, so a request exceeding any single node's advertised count leaves the pod unschedulable with an insufficient-resource message. Verifying the per-node GPU count and lowering the pod's request to fit within one node restores scheduling.

Exam trap

The trap here is reading 'Insufficient nvidia.com/gpu' as a driver or image problem, when it actually means the requested GPU count exceeds what any individual node advertises.

83
MCQhard

An MLOps engineer needs to guarantee that a latency-sensitive inference Deployment always has GPU capacity available, even when a large training Job is submitted to the same namespace. The cluster uses the NVIDIA GPU Operator and nodes have four A100 GPUs each. Which approach reliably reserves GPU capacity for the inference Deployment?

A.Create a separate node pool and use nodeSelector or nodeAffinity to pin the inference Deployment to nodes tainted for inference only
B.Set a higher PriorityClass on the inference Deployment so it preempts training pods when needed
C.Enable the NVIDIA MPS control daemon and configure each inference pod with a shared memory fraction
D.Apply a ResourceQuota on the namespace that limits nvidia.com/gpu to the number of inference replicas
AnswerA

Dedicating a tainted node pool and pinning the inference Deployment with nodeSelector or nodeAffinity guarantees that training pods cannot consume those GPUs. Taints repel pods that lack the matching toleration, so the reserved capacity is genuinely protected regardless of how much training demand arrives, which is the only option that provides a hard reservation.

Why this answer

Only a dedicated, tainted node pool combined with nodeSelector or nodeAffinity creates a hard capacity reservation that training workloads cannot violate. Priority-based preemption and MPS improve scheduling behavior or sharing but do not prevent a co-located training Job from consuming the GPUs first, and a namespace ResourceQuota limits aggregate requests without reserving specific capacity for the inference Deployment.

Exam trap

The trap here is believing that a higher PriorityClass reserves GPU capacity, when it only enables eviction after a scheduling failure, not proactive reservation.

84
MCQmedium

A Kubernetes cluster running the NVIDIA GPU Operator is shared by an inference team and a research team. The research team's training pods repeatedly evict the inference pods from GPUs, causing latency spikes in production. The administrator wants to guarantee that inference pods always get GPU access first. Which Kubernetes scheduling mechanism should be configured?

A.Enable the NVIDIA MIG feature on all GPUs and dedicate a MIG instance to each inference pod.
B.Assign a higher PriorityClass to the inference pods and enable preemption on the scheduler.
C.Configure a taint on the GPU nodes and add the corresponding toleration only to the inference pods.
D.Create a PodDisruptionBudget for the inference deployment and set maxUnavailable to zero.
AnswerB

PriorityClass with preemption allows higher-priority inference pods to be scheduled and to evict lower-priority training pods when GPU resources are scarce. This directly protects production inference latency by ensuring inference workloads are admitted first, which matches the requirement to guarantee GPU access for the inference team.

Why this answer

Priority and preemption are the native Kubernetes controls that decide which pods win when GPU resources are contended. Giving inference pods a higher PriorityClass lets the scheduler admit them first and evict lower-priority training pods when necessary, which directly satisfies the requirement that production inference always obtains GPU access.

Exam trap

The trap here is assuming that taints, tolerations, or PodDisruptionBudgets control scheduler contention, when in fact only PriorityClass with preemption orders competing pods for scarce GPU resources.

85
MCQeasy

A data science team submits a PyTorch training job to a Kubernetes cluster managed by Run:ai. The job requests two GPUs but only one is allocated, and the second worker hangs waiting for a peer. Which Run:ai capability should the administrator verify is configured so the distributed job receives all requested GPUs atomically?

A.A higher priority class assigned to the training job so it preempts other workloads.
B.Gang scheduling, which ensures all pods in a distributed job are scheduled together or not at all.
C.Node affinity rules that pin each worker to a specific GPU node by hostname.
D.Enabling time-slicing on the GPU device plugin to increase the apparent GPU count.
AnswerB

Distributed training jobs require all workers to start together; partial allocation causes hangs because ranks wait for peers that never launch. Run:ai's gang scheduling treats the workload as an atomic unit, allocating all requested GPUs or leaving the job pending. Verifying this configuration addresses the symptom of one GPU allocated and a stalled peer directly.

Why this answer

Distributed training depends on every rank being present before collective operations begin; a single missing worker causes the rest to block. Run:ai's gang scheduling is the mechanism that admits the entire workload as a unit, so all requested GPUs are granted together. Confirming gang scheduling is enabled and applied to the job resolves the partial allocation that produced the hang.

Exam trap

The trap here is assuming that priority or affinity alone can prevent partial placement, when only all-or-nothing gang scheduling guarantees every rank starts together.

86
MCQhard

Refer to the exhibit. An administrator is attempting to deploy a job to a namespace with a ResourceQuota defined. What is the cause of this error?

A.The GPU driver version is incompatible with the quota controller.
B.The pod manifest is missing the required GPU resource limits.
C.The cluster is out of available GPU capacity.
D.The NVIDIA Device Plugin is not running in the namespace.
AnswerB

The error message explicitly states that limits for 'nvidia.com/gpu' must be specified. This is a common requirement in environments where quotas are implemented to ensure fair scheduling. Without these limits, the admission controller rejects the pod because it cannot account for the GPU usage against the namespace quota.

Why this answer

The error indicates that the namespace has a ResourceQuota enforcing that all pods must specify GPU limits, but the submitted pod manifest lacks a 'resources.limits.nvidia.com/gpu' entry. In multi-tenant environments, ResourceQuotas are essential for preventing a single user from consuming the entire GPU capacity. The manifest must include a valid GPU limit to satisfy the namespace policy, ensuring the cluster remains balanced across different organizational teams.

Exam trap

Candidates often assume ResourceQuota errors stem from cluster-wide node exhaustion, missing the fact that namespace-level policies explicitly require explicit GPU resource limits in the pod manifest.

87
Multi-Selecthard

A platform team operates a Kubernetes cluster where several teams submit GPU training jobs. The administrator needs to enforce per-namespace limits on the number of GPUs that can be consumed and prevent a single namespace from monopolizing all GPU capacity. Which TWO Kubernetes resources should be configured to achieve this? (Choose two.)

Select 2 answers
A.A PriorityClass that assigns a low priority value to all training pods in the namespace.
B.A PodSecurityPolicy that denies privileged containers in the namespace.
C.A LimitRange that sets a default and maximum nvidia.com/gpu value for containers in the namespace.
D.A NetworkPolicy that restricts traffic between pods in different namespaces.
E.A ResourceQuota that specifies nvidia.com/gpu in its hard limits for each namespace.
AnswersC, E

LimitRange applies defaults and bounds to individual containers. Setting a maximum nvidia.com/gpu stops a single container from requesting an excessive number of GPUs, and the default ensures pods that omit a GPU request still receive a defined value, complementing the namespace-wide cap enforced by ResourceQuota.

Why this answer

ResourceQuota enforces an aggregate ceiling on nvidia.com/gpu per namespace, while LimitRange constrains and defaults the per-container GPU request. Together they bound total namespace consumption and prevent any single container from grabbing an outsized share, which is exactly the governance the platform team needs.

Exam trap

The trap here is believing that PriorityClass or NetworkPolicy can cap GPU consumption, when only quota and limit-range objects act on resource quantities at admission time.

← PreviousPage 2 of 2 · 87 questions total

Ready to test yourself?

Try a timed practice session using only Workload Management questions.