Courseiva

CCNA Workload Management Questions

75 of 87 questions · Page 1/2 · Workload Management · Answers revealed

1
MCQhard

A research team runs a multi-node distributed training job spanning eight GPU nodes. Jobs frequently begin execution before all worker pods are running, and the collective initialization hangs until the operator manually scales the job down and up. The administrator wants the scheduler to admit the job only when all of its pods can be placed together. Which mechanism should be used?

A.Increase the kube-scheduler's default backoff period so pods that fail to schedule retry less aggressively and wait for peers to become ready.
B.Configure gang scheduling through a scheduler plugin so the job's pods are placed atomically only when the full group can be accommodated.
C.Set a high PriorityClass on every worker pod so the scheduler treats the group as important and places members as soon as resources appear.
D.Add a pod anti-affinity rule requiring each worker to run on a distinct node so that placement spreads evenly across the eight GPU nodes.
AnswerB

Gang scheduling holds the entire job group until every member can be placed on available resources, then admits them together. This directly prevents the partial-start condition that stalls collective initialization. Using a scheduler plugin that supports gang or coscheduling semantics gives the atomic placement guarantee the team needs, eliminating the manual scale-down and scale-up workaround.

Why this answer

Collective initialization requires all ranks to be present before computation proceeds, so partial startup leads to hangs. Gang scheduling solves this by treating the job group as a single scheduling unit and admitting it only when every member fits. Scheduler plugins that implement coscheduling or gang semantics provide this all-or-nothing placement, which is why they are standard in large distributed training environments.

Exam trap

The trap here is reaching for priority or anti-affinity as a fix for partial startup; those affect ordering and node distribution but never make pod admission atomic, so the collective can still begin with missing ranks.

2
MCQhard

An administrator manages a shared NVIDIA cluster where several teams run inference services. One team's pods are being evicted repeatedly, and DCGM metrics show the node's GPUs are healthy but memory on the devices is nearly exhausted. The team insists their model fits. Which action should the administrator take FIRST to identify the cause?

A.Increase the pod's memory limit in the deployment manifest and restart it.
B.Replace the affected GPUs and re-run the job on fresh hardware.
C.Lower the DCGM sampling interval so metrics capture the spike more precisely.
D.Inspect the pod specs and running processes for GPU memory held by other containers on the same node.
AnswerD

When device memory is nearly exhausted but hardware is healthy, the most likely cause is co-tenancy: other containers on the same node are holding GPU memory. Because the device plugin allocates whole GPUs by default, multiple pods can land on the same device only when sharing is explicitly enabled, so checking pod specs and active processes reveals whether a neighbor is consuming the memory the team expects to have.

Why this answer

Healthy GPUs with near-exhausted device memory in a shared cluster usually indicate that another container on the same node is holding memory on the same device. The fastest way to confirm this is to inspect pod specifications and running processes to see which workloads share the GPU. Replacing hardware, changing host memory limits, or tuning telemetry do not address the allocation conflict and delay the correct diagnosis.

Exam trap

The trap here is treating GPU memory exhaustion as a hardware fault rather than as contention between co-located workloads.

3
MCQmedium

Which scheduling strategy is recommended to maximize the efficiency of long-running training jobs on preemptible instances?

A.Avoiding preemptible instances for all training jobs.
B.Implementing granular checkpointing and automated job resumption.
C.Locking the job to a specific node using node affinity.
D.Increasing the priority of the job to the maximum level.
AnswerB

Granular checkpointing minimizes the amount of lost progress during a preemption event. Coupled with automated job resumption, the system can quickly restart the task on a new node from the last checkpoint. This allows for safe usage of low-cost preemptible instances, balancing economic efficiency with the need for persistent progress.

Why this answer

Long-running jobs on preemptible (or spot) instances require frequent, efficient checkpointing to survive node reclamation. By combining a robust checkpointing schedule with smart job resubmission logic that monitors for preemption signals, organizations can take advantage of low-cost instances while minimizing the loss of progress. This approach allows for significant cost savings in non-critical training without sacrificing the overall reliability of the research project.

Exam trap

Test-takers often rely solely on high-availability cluster setups or cheaper instance pricing without establishing application-level mechanisms to preserve state when preemption inevitably occurs.

4
Multi-Selecthard

An administrator manages a Kubernetes cluster where the NVIDIA GPU Operator has deployed the device plugin and MIG Manager. A tenant wants to run several small inference services that each need only a fraction of a GPU, isolated from other tenants' memory and fault domains. The administrator decides to use Multi-Instance GPU mode. Which TWO actions must be performed to make MIG-backed GPU resources schedulable to those pods? (Choose two.)

Select 2 answers
A.Install the NVIDIA Network Operator and enable RDMA device advertisement on the node
B.Enable MIG mode on the physical GPU and define a MIG profile/geometry via the MIG Manager configuration
C.Request the MIG device in the pod spec using the extended resource name advertised by the device plugin, such as nvidia.com/mig-1g.5gb
D.Create a RuntimeClass that points to the nvidia-container-runtime and reference it from each inference pod
E.Apply a ResourceQuota that limits nvidia.com/gpu to zero in the tenant namespace
AnswersB, C

MIG mode must be enabled on the GPU, and the MIG Manager in the GPU Operator applies a geometry configuration that carves the device into instances with dedicated memory and fault isolation. Without an applied profile, no MIG instances exist and the node advertises no MIG resources, so scheduling cannot succeed regardless of pod specification.

Why this answer

MIG-backed scheduling requires two coordinated steps: enabling MIG mode and applying a geometry through the MIG Manager so instances exist, and having pods request the specific profile resource name that the device plugin advertises, such as nvidia.com/mig-1g.5gb. Together these produce isolated GPU slices with dedicated memory and fault domains for each tenant's inference service.

Exam trap

The trap here is treating a RuntimeClass or ResourceQuota as the mechanism that creates MIG slices, when the essential pair is MIG geometry configuration plus requesting the advertised profile-specific extended resource.

5
MCQeasy

A cloud operations team is using NVIDIA AI Enterprise with Kubernetes to deploy inference workloads. They want to ensure that GPU resources are allocated to pods only when explicitly requested, and that pods without GPU requests do not consume GPU resources. Which Kubernetes feature should they use to enforce this behavior?

A.Use node selectors to schedule pods only on GPU nodes.
B.Define resource requests and limits for 'nvidia.com/gpu' in the pod specification.
C.Enable the GPU Operator's 'device plugin' to automatically inject GPU requests into all pods.
D.Use a mutating admission webhook to add GPU requests to pods based on their namespace.
AnswerB

In Kubernetes, GPU resources are requested using the 'nvidia.com/gpu' resource name in the pod's resource requests and limits. When a pod specifies this, the scheduler allocates a GPU to it; pods without this request are not allocated GPUs, even if scheduled on GPU nodes. This enforces explicit GPU allocation and prevents pods from consuming GPU resources unintentionally, satisfying the team's requirement.

Why this answer

Kubernetes requires pods to explicitly request GPU resources using the 'nvidia.com/gpu' resource name in their resource specifications. The scheduler then allocates GPUs only to those pods, ensuring that pods without such requests do not consume GPU resources even if they run on GPU nodes. This mechanism enforces the desired explicit allocation policy and is the standard way to manage GPU resources in Kubernetes.

Exam trap

The trap here is confusing node-level scheduling constraints with resource-level allocation; node selectors only control placement, not whether a GPU is actually assigned.

6
MCQmedium

In a multi-node training scenario, what is the significance of the NVIDIA Collective Communications Library (NCCL) in workload management?

A.It handles the automatic scaling of pods in the cluster.
B.It provides a mechanism to optimize inter-node data exchange.
C.It replaces the need for high-speed network cabling.
D.It monitors the temperature of the GPUs during training.
AnswerB

NCCL is purpose-built to accelerate collective operations across distributed GPUs. It intelligently utilizes hardware interconnects such as NVLink and InfiniBand to reduce communication latency, which is the primary bottleneck in distributed training, ensuring that gradient synchronization does not throttle the overall training throughput of the job.

Why this answer

NCCL is a critical communication primitive for multi-GPU, multi-node training. It optimizes the collective operations (like AllReduce) required for synchronizing gradient updates across nodes. Effective workload management requires ensuring that nodes are configured for high-bandwidth, low-latency interconnects like NVLink and InfiniBand, which NCCL utilizes to minimize synchronization overhead.

This optimization is essential for scaling training jobs to large clusters without performance degradation caused by network bottlenecks.

Exam trap

Candidates often confuse NCCL with general network protocols, failing to recognize its specific role in optimizing collective operations like AllReduce for multi-node GPU synchronization.

7
MCQmedium

Which approach is most effective for scaling an inference workload that experiences sudden, unpredictable spikes in request volume?

A.Setting a fixed number of replicas to match peak expected load.
B.Using a load balancer to redirect traffic to an idle cluster.
C.Implementing HPA with custom metrics provided by DCGM Exporter.
D.Increasing the memory allocation for each inference container.
AnswerC

HPA using custom GPU metrics allows for precise scaling triggered by actual hardware usage. As request volume spikes, GPU utilization increases, and the HPA automatically triggers the deployment of additional inference pods. This ensures that the system scales only when needed, maintaining optimal performance while minimizing resource waste during quiet periods.

Why this answer

Predictive or reactive autoscaling based on custom GPU metrics (like utilization) allows the cluster to adjust replica counts in real-time. By monitoring the GPU load and scaling the inference pods accordingly, the system maintains low latency during spikes while keeping costs low during idle periods. This responsiveness is vital for production AI services where performance SLAs are tied directly to user experience.

Exam trap

Candidates frequently choose CPU-based metrics for autoscaling, not realizing that GPU-bound workloads often saturate the GPU while CPU usage remains low, rendering standard HPA configurations ineffective for scaling.

8
MCQmedium

A production inference service experiences intermittent latency spikes. The service is deployed on shared infrastructure. Which tool would best help an AI Ops engineer identify if GPU resource contention is the cause?

A.Standard Kubernetes 'kubectl top' command.
B.NVIDIA DCGM Exporter with Prometheus/Grafana.
C.The Linux 'top' command on the worker node.
D.A simple network latency test tool like 'ping'.
AnswerB

DCGM collects granular, real-time metrics directly from the GPU hardware. By exporting these to Prometheus and visualizing them in Grafana, engineers can identify spikes in GPU activity or bandwidth saturation that correlate with the application's latency, providing a clear path to identifying the source of resource contention.

Why this answer

NVIDIA DCGM (Data Center GPU Manager) provides detailed telemetry data, including GPU utilization, memory bandwidth, and power usage per process. By correlating latency spikes with DCGM metrics, engineers can pinpoint if another process on the same GPU is competing for compute resources or memory bandwidth. This visibility is essential for performance tuning and workload placement, allowing engineers to verify if resource isolation policies are functioning as expected.

Exam trap

Candidates might suggest looking at CPU logs or standard Kubernetes pod metrics. These metrics are often blind to GPU-specific resource contention, such as memory bus saturation or shared SM usage.

9
MCQeasy

When managing workloads on NVIDIA DGX systems, what is the primary role of the NVIDIA device plugin in the Kubernetes ecosystem?

A.To compile CUDA code for specific GPU architectures during the job scheduling process.
B.To monitor the health status of GPUs and advertise them to the Kubernetes scheduler.
C.To automatically optimize neural network hyper-parameters for faster training throughput.
D.To replace the need for the NVIDIA Container Toolkit in the container runtime environment.
AnswerB

The device plugin is the bridge that allows Kubernetes to 'see' the GPUs. It queries the NVIDIA driver for device information, reports healthy devices to the API server, and manages the lifecycle of GPU assignments, ensuring the scheduler only places GPU-dependent pods on nodes that can support them.

Why this answer

The NVIDIA device plugin acts as an interface between Kubernetes and the NVIDIA drivers, enabling the scheduler to advertise and allocate GPUs as first-class resources. It is fundamental to GPU workload management because it informs the kubelet about the availability, health, and count of GPUs on the node, ensuring that pods requesting GPUs are placed on nodes with sufficient capacity.

Exam trap

Candidates often confuse the device plugin with the container runtime. The plugin only handles hardware advertisement and discovery, not the injection of CUDA libraries or driver communication into the container.

10
MCQmedium

A platform team runs an NVIDIA GPU Operator-managed Kubernetes cluster shared by two research groups. Group A's pods request `nvidia.com/gpu: 1` and are scheduled correctly, but Group B's pods that omit any GPU resource request are also landing on GPU nodes and consuming host memory and CPU, degrading Group A's jobs. The team wants Group B's non-GPU pods to stop consuming capacity on the GPU node pool without changing Group B's manifests. Which action should the administrator take?

A.Set `NVIDIA_DRIVER_CAPABILITIES=compute,utility` on the GPU Operator DaemonSet so non-GPU pods cannot attach to the device.
B.Create a `ResourceQuota` in each namespace that caps `requests.nvidia.com/gpu` at zero for Group B.
C.Enable the `nvidia` device plugin's `--fail-on-init-error=false` flag so unresolvable GPU requests fall back to CPU scheduling.
D.Apply a taint such as `nvidia.com/gpu=present:NoSchedule` to the GPU nodes and add a matching toleration to Group A's pod templates.
AnswerD

Tainting GPU nodes with a dedicated key and NoSchedule effect prevents any pod lacking a matching toleration from being placed there, so Group B's unmodified workloads are repelled while Group A's templates are explicitly admitted. This is the standard scheduler-level control for dedicating a node pool, and it requires no edits to Group B's manifests, matching the stated constraint exactly.

Why this answer

Dedicating a node pool to GPU consumers is done at the scheduler layer: a taint on the GPU nodes repels every pod that does not carry a matching toleration. Because Group B's manifests cannot be changed, repelling rather than attracting is the only workable direction, and adding tolerations to Group A's templates is a one-time change under the platform team's control. Quotas, driver flags, and plugin options all operate after or outside scheduling and cannot keep non-GPU pods off those nodes.

Exam trap

The trap here is assuming that a ResourceQuota or device-plugin setting controls node placement, when only taints, tolerations, and node affinity influence the scheduler's decisions.

11
MCQhard

An AI operations team runs mixed training and inference workloads on a Kubernetes cluster managed with the NVIDIA GPU Operator. Inference pods frequently arrive in bursts and must start within seconds, while long-running training jobs occupy most MIG-capable A100 GPUs for days. Administrators want burst inference pods to obtain GPU capacity immediately without preempting or restarting the training jobs, and they want the cluster to reclaim those resources automatically when the burst ends. Which approach best satisfies these requirements?

A.Create a node pool reserved for inference, and configure a cluster autoscaler with scale-to-zero so burst pods trigger new GPU nodes that are removed when idle.
B.Configure the inference Deployment with a topologySpreadConstraint across all GPU nodes so its replicas distribute evenly and find free capacity.
C.Enable the GPU Operator's time-slicing configuration so inference and training containers share the same physical GPUs through a shared device plugin.
D.Deploy inference pods with a higher PriorityClass and enable preemption so the scheduler evicts lower-priority training pods when capacity is exhausted.
AnswerA

A dedicated inference node pool with scale-to-zero autoscaling lets burst pods trigger immediate node provisioning while training jobs keep running untouched on their own nodes. When the burst subsides, the autoscaler drains and removes the idle inference nodes, reclaiming GPU capacity and cost automatically. This directly matches both requirements: no preemption of training and automatic reclamation after the burst.

Why this answer

The scenario requires two properties simultaneously: burst inference capacity that appears without disturbing running training jobs, and automatic reclamation when the burst ends. Only a dedicated autoscaled node pool satisfies both, because new GPU nodes are provisioned on demand for inference pods while training pods remain scheduled and running on their existing nodes. When demand drops, scale-to-zero removes the extra nodes, freeing GPU resources without any preemption or job restarts.

Exam trap

The trap here is assuming that priority and preemption are the natural answer for burst workloads, when the requirement that training jobs never restart makes preemption disqualifying.

12
MCQmedium

An organization is deploying large-scale LLM training workloads on an NVIDIA DGX SuperPOD. The data science team reports that training jobs are frequently preempted by higher-priority batch jobs, leading to significant checkpointing overhead. Which Workload Manager configuration strategy best minimizes resource fragmentation and improves overall cluster utilization while maintaining SLA requirements?

A.Increase the default job preemption grace period to allow all jobs to complete.
B.Disable multi-instance GPU (MIG) to allow all jobs to access full GPU memory.
C.Implement gang scheduling combined with strict job priority and preemption policies.
D.Manually partition the cluster into static zones for different research teams.
AnswerC

Gang scheduling ensures that all requested resources for a distributed job are acquired simultaneously, preventing deadlocks where multiple jobs wait for partial allocations. Combining this with clear priority levels allows the scheduler to preempt lower-priority tasks efficiently, maximizing cluster utility while protecting SLA-bound high-priority distributed training workloads.

Why this answer

Implementing gang scheduling with preemption thresholds ensures that resources are allocated atomically to distributed training jobs, preventing partial allocations that lead to starvation. By configuring preemption grace periods and job priorities, the scheduler can effectively balance urgent tasks against long-running training runs. This approach is critical in NVIDIA environments to ensure that high-bandwidth inter-node communication remains optimized during heavy compute cycles, preventing inefficient resource usage.

Exam trap

Candidates often suggest simple priority queues without gang scheduling. This causes 'fragmentation' where a job starts with only half its required GPUs, leading to a deadlock or inefficient training.

13
MCQeasy

Why is it important to use a persistent storage volume for model checkpoints in a distributed training job?

A.To increase the write speed of the training data.
B.To ensure model checkpoints survive pod failures.
C.To provide a high-speed cache for real-time inference.
D.To hide the model from unauthorized cluster users.
AnswerB

Pod failures are common in orchestrated environments due to node errors or preemptions. By storing checkpoints on persistent storage, the state is decoupled from the lifecycle of the individual pod, allowing a new pod instance to pick up exactly where the previous process stopped during training.

Why this answer

Distributed training jobs are prone to node failures or preemptions in shared environments. Checkpointing allows the training state to be saved periodically. By storing these checkpoints on persistent volumes, the job can resume from the last saved state rather than restarting from scratch if a failure occurs.

This minimizes wasted compute time and ensures that long-running training tasks can eventually complete, protecting the investment in expensive cluster resources.

Exam trap

Candidates often think persistent storage is for performance or speed. In reality, it is purely for fault tolerance and state persistence during inevitable pod crashes or preemptions.

14
MCQhard

Refer to the exhibit. The cluster administrator observes near-capacity memory utilization across three GPUs. What is the most likely consequence if an additional pod is scheduled to these GPUs without memory partitioning?

A.The scheduler will automatically enable MIG.
B.The job will experience OOM errors and fail.
C.The scheduler will use CPU swap memory.
D.The job will be throttled but complete successfully.
AnswerB

With memory utilization already exceeding 95% on all GPUs, there is insufficient headroom to spawn a new process. The operating system or the NVIDIA driver will trigger an Out-of-Memory (OOM) kill signal, resulting in the immediate failure of the new container and potential instability for existing tasks.

Why this answer

Scheduling an additional pod onto these GPUs will lead to Out-of-Memory (OOM) errors. Because the GPUs are already at near-maximum capacity, the driver will likely kill the incoming process or potentially destabilize existing tasks. This highlights the importance of workload monitoring and resource planning, as AI Ops must ensure that requests match actual hardware capacity to prevent catastrophic job failures in production environments.

Exam trap

Candidates might guess that the system will simply slow down or swap to system RAM, forgetting that GPU memory is non-swappable and will cause immediate process termination via OOM error.

15
MCQmedium

An AI platform team runs GPU-accelerated inference pods on a Kubernetes cluster with the NVIDIA GPU Operator. During peak load, high-priority latency-sensitive inference pods are frequently preempted by large batch training jobs that were submitted later. The team wants the scheduler to guarantee that inference pods always win placement and preemption decisions against training pods without manually cordoning nodes. Which mechanism should the administrator configure?

A.Enable the GPU Operator's time-slicing feature with a replica count greater than one so inference and training pods share each physical GPU concurrently.
B.Create a separate namespace with a ResourceQuota that limits training pods to a small number of GPUs and place all inference pods in the default namespace.
C.Assign the inference pods a higher PriorityClass value and reference it in their pod spec so the kube-scheduler preempts lower-priority training pods when resources are scarce.
D.Add a nodeSelector to the inference pods that pins them to nodes labeled with the NVIDIA GPU product name, so they only run on nodes not used by training.
AnswerC

PriorityClass is the native Kubernetes scheduling mechanism that orders pending pods; a higher integer value makes the kube-scheduler prefer those pods and preempt lower-priority workloads that already occupy GPUs. Referencing the class in the pod spec ensures inference pods win placement and preemption decisions, which is exactly what the team needs.

Why this answer

Kubernetes resolves GPU contention through the scheduler's priority and preemption logic, not through device sharing or quota limits. Giving inference pods a higher PriorityClass value makes the kube-scheduler place them first and evict lower-priority training pods when capacity is exhausted, which delivers the guaranteed preference the platform team requires.

Exam trap

The trap here is assuming that GPU sharing features such as time-slicing or MIG establish workload precedence, when only PriorityClass plus scheduler preemption actually reorders and evicts competing pods.

16
MCQhard

In an NVIDIA-accelerated Kubernetes environment, why is it critical to configure a 'RuntimeClass' for GPU-enabled pods?

A.To allow the Kubernetes scheduler to place pods on nodes based on GPU availability.
B.To define the specific NVIDIA driver version required by the application.
C.To ensure the container runtime correctly injects GPU drivers and libraries into the pod.
D.To increase the priority of GPU workloads over standard CPU-only workloads.
AnswerC

The RuntimeClass enables the node to switch to the NVIDIA-specific container runtime when needed. This runtime is responsible for the crucial task of injecting the required driver and library files, enabling the container to recognize and interact with the GPU hardware without requiring manual configuration in every single pod.

Why this answer

The RuntimeClass ensures that the container is started with the correct NVIDIA container runtime (e.g., nvidia-container-runtime). This runtime is responsible for performing the necessary hardware mapping, injecting the CUDA libraries, and setting the environment variables required for the container to actually utilize the GPU. Without this configuration, the pod might be scheduled successfully but will fail at runtime because it cannot communicate with the NVIDIA driver.

Exam trap

Candidates often assume that requesting a GPU resource in the pod spec is sufficient, forgetting that the RuntimeClass is the mechanism that actually injects the necessary NVIDIA libraries.

17
MCQeasy

An administrator must run a batch inference job that requires exactly two NVIDIA GPUs on a Kubernetes cluster managed by the NVIDIA GPU Operator. Which pod specification field should be used to request those GPUs?

A.resources.requests with the key nvidia.com/gpu set to 2 and no limit
B.annotations with the key nvidia.com/gpu.count set to 2
C.nodeSelector with the key nvidia.com/gpu.count set to 2
D.resources.limits with the key nvidia.com/gpu set to 2
AnswerD

The NVIDIA device plugin advertises GPUs as the extended resource nvidia.com/gpu, and extended resources must be requested through resources.limits. Setting the limit to 2 causes the scheduler to place the pod on a node with two allocatable GPUs and injects those devices into the container, which meets the exact requirement.

Why this answer

GPUs exposed by the NVIDIA device plugin appear to Kubernetes as the extended resource nvidia.com/gpu. Extended resources must be declared in resources.limits, and setting that limit to 2 makes the scheduler reserve two GPUs on a suitable node and pass the devices into the container.

Exam trap

The trap here is treating nvidia.com/gpu like CPU or memory and specifying it only under resources.requests, which the API server rejects for extended resources.

18
MCQeasy

A platform engineer is preparing an NVIDIA-accelerated Kubernetes cluster for a new team that will submit PyTorch training jobs. The team wants jobs to request GPUs without hardcoding device indices. Which Kubernetes resource should the engineer ensure is installed and healthy so pods can request nvidia.com/gpu resources?

A.NVIDIA Container Toolkit installed on each worker node
B.NVIDIA DCGM Exporter deployed as a DaemonSet
C.NVIDIA Network Operator with RDMA shared device plugin
D.NVIDIA GPU Operator with the device plugin component enabled
AnswerD

The NVIDIA device plugin, deployed by the GPU Operator, registers nvidia.com/gpu as an extended resource and advertises the count of available GPUs on each node. Without it, pods cannot request GPUs by resource name. Ensuring the operator and its device plugin are healthy is the correct step to let PyTorch jobs request GPUs abstractly rather than by device index.

Why this answer

Kubernetes learns about specialized hardware through device plugins that advertise extended resources. The NVIDIA device plugin, managed by the GPU Operator, publishes nvidia.com/gpu counts so the scheduler can allocate GPUs to pods. With it healthy, PyTorch jobs can declare a GPU request and receive an assigned device without the user specifying a physical index.

Exam trap

The trap here is equating GPU driver or container runtime installation with resource advertisement, when only the device plugin registers nvidia.com/gpu with the Kubernetes scheduler.

19
Multi-Selecthard

A data engineering team is deploying a distributed data processing workload. Which THREE metrics are most important to monitor in the Workload Manager to ensure optimal GPU throughput and identify potential bottlenecks?

Select 3 answers
A.GPU Duty Cycle (Active utilization percentage).
B.GPU Memory Bandwidth Utilization.
C.Container CPU usage percentage.
D.PCIe Throughput (Data transfer rates).
E.Pod restart count in the namespace.
AnswersA, B, D

The duty cycle measures the percentage of time the GPU is performing actual computation. If this number is consistently low, it indicates that the GPU is waiting for data or that the workload is CPU-bound, which is a vital indicator for troubleshooting underperforming AI training or processing jobs.

Why this answer

Monitoring these metrics allows administrators to detect inefficiencies in data loading or GPU utilization. GPU Duty Cycle tracks how often the GPU is actively processing data. Memory Bandwidth Utilization highlights if the bottleneck is in data transfer.

Finally, PCIe throughput indicates whether the bus between the CPU and GPU is saturated. Together, these metrics provide a complete picture of why a workload might be underperforming in a high-performance environment.

Exam trap

Candidates often select general CPU or memory metrics instead of specialized GPU metrics like GPU Duty Cycle, Memory Bandwidth, and PCIe Throughput when diagnosing GPU workload bottlenecks in Workload Manager.

20
MCQeasy

A data engineering team runs nightly batch inference on a Kubernetes cluster with NVIDIA GPUs. Jobs sometimes fail because two pods are scheduled onto the same physical GPU and one exhausts framebuffer memory. The team wants each pod to receive an isolated slice of a GPU with dedicated memory. Which NVIDIA feature should they enable?

A.Multi-Instance GPU (MIG)
B.NVIDIA Collective Communications Library (NCCL)
C.GPUDirect Storage
D.NVIDIA Container Toolkit
AnswerA

MIG partitions a single physical GPU into multiple independent instances, each with its own dedicated memory and compute slices. Because the memory is hardware-partitioned, one instance cannot consume another's framebuffer, which directly prevents the failure the team is seeing. The device plugin can then advertise each MIG instance as a schedulable resource so pods receive isolated slices.

Why this answer

The team needs hardware-level isolation with dedicated memory per workload. Multi-Instance GPU partitions a physical GPU into independent instances, each with its own memory and compute, and the device plugin can expose those instances as discrete resources. GPUDirect Storage, NCCL, and the Container Toolkit each address I/O paths, communication, or containerization rather than partitioning a GPU among tenants.

Exam trap

The trap here is confusing containerization or communication tooling, which makes GPUs usable, with partitioning features that actually isolate GPU memory.

21
MCQmedium

What is the primary benefit of using NVIDIA GPU Operator in a Kubernetes cluster for workload management?

A.It enables automatic model fine-tuning for specific frameworks.
B.It provides a declarative way to manage GPU drivers and software components.
C.It optimizes the neural network architecture automatically.
D.It eliminates the need for containerization in AI workflows.
AnswerB

The GPU Operator uses Kubernetes custom resources to declaratively manage NVIDIA software. This ensures that driver versions and plugin configurations are consistent across nodes, eliminating manual setup errors. This automation is crucial for scalability, as it allows administrators to manage thousands of GPUs with the same consistency as a single node.

Why this answer

The NVIDIA GPU Operator automates the lifecycle management of NVIDIA software components, including drivers, the container toolkit, and device plugins. By standardizing these installations across all cluster nodes, it ensures a consistent and stable environment for AI workloads, reducing configuration drift and manual operational overhead that often plagues large-scale distributed infrastructure deployments.

Exam trap

Test-takers often confuse the NVIDIA GPU Operator with general container runtimes or standard Kubernetes schedulers, overlooking its specific role in declaratively managing drivers and software components.

22
Multi-Selectmedium

An administrator is optimizing multi-GPU utilization. Which TWO of the following configurations allow multiple containers to share a single physical GPU on a supported NVIDIA architecture?

Select 2 answers
A.Enabling NVIDIA Time-Slicing in the GPU Operator configuration.
B.Configuring Multi-Instance GPU (MIG) profiles.
C.Increasing the CUDA_VISIBLE_DEVICES environment variable.
D.Setting a higher GPU limit in the Kubernetes manifest.
E.Deploying the NVIDIA Network Operator.
AnswersA, B

Time-Slicing allows multiple pods to share a GPU by switching context at rapid intervals. This is a software-based approach that enables oversubscription, allowing smaller workloads to execute concurrently on the same hardware, which is highly beneficial for development environments or low-throughput inference tasks that don't need full GPU power.

Why this answer

Multi-Instance GPU (MIG) and Time-Slicing are the primary methods for sharing physical GPU resources. MIG provides hardware-level isolation, while Time-Slicing provides software-based multiplexing. Understanding these options is essential for AI operations, as it allows administrators to maximize hardware ROI by supporting smaller workloads that do not require an entire A100 or H100 GPU, effectively increasing the density of the training or inference environment.

Exam trap

Candidates frequently confuse software-level multiplexing options like Time-Slicing with hypervisor features or mistake them for hardware-partitioning mechanisms like MIG, failing to recognize that both are valid sharing methods.

23
MCQhard

An administrator supports a multi-tenant cluster where several teams share GPUs. Leadership requires that each team's batch jobs receive a fair share of GPU time and that one team cannot monopolize devices by submitting thousands of low-priority pods. Jobs are submitted through a Kubernetes-native batch scheduler that supports queueing. Which approach best enforces fair-share GPU allocation across teams?

A.Use node taints and tolerations to dedicate specific GPU nodes to each team permanently
B.Assign each team a distinct Kubernetes namespace and rely on the default kube-scheduler for GPU placement
C.Define per-team queues with weighted fair-share policies and GPU-aware scheduling in the batch scheduler
D.Apply identical PriorityClass values to all team pods so the scheduler treats them equally
AnswerC

A GPU-aware batch scheduler with per-team queues and weighted fair-share policies allocates GPU capacity proportionally across tenants and prevents any single team from monopolizing devices through sheer volume. It coordinates gang scheduling and quotas at the queue level, directly delivering the fairness and anti-monopolization requirement for shared GPU clusters.

Why this answer

Weighted fair-share queues in a GPU-aware batch scheduler allocate device time proportionally among teams and enforce queue-level limits, preventing a single tenant from monopolizing GPUs by volume. The default scheduler, equal priorities, or static node partitioning cannot provide dynamic proportional sharing or gang scheduling for distributed jobs.

Exam trap

The trap here is equating namespace isolation or equal priorities with fairness, when fair share actually requires a scheduler that tracks per-queue usage and applies weighted allocation across teams.

24
MCQmedium

When deploying large-scale distributed training jobs, why is it recommended to use the NVIDIA Network Operator in conjunction with the GPU Operator?

A.To increase the number of supported GPU containers.
B.To optimize RDMA throughput for distributed training.
C.To provide automatic GPU driver updates.
D.To enable multi-tenancy on GPU nodes.
AnswerB

Distributed training relies on efficient GPU-to-GPU communication across nodes. The Network Operator manages the configuration of RDMA-enabled NICs, allowing GPUs to communicate directly with minimal CPU intervention. This dramatically increases throughput and reduces latency, which is essential for the performance of large-scale distributed machine learning training jobs.

Why this answer

The Network Operator automates the configuration of high-speed interconnects like InfiniBand or RoCE, which are crucial for distributed training across multiple nodes. By aligning network resources with GPU placement, the operator minimizes latency and maximizes throughput. This ensures that the GPU compute capacity is not bottlenecked by slow network communication, which is essential for scaling complex AI model training across large clusters effectively.

Exam trap

Candidates often assume that standard Kubernetes networking is sufficient for distributed AI training, overlooking the critical requirement for low-latency, high-throughput RDMA interconnects.

25
MCQeasy

A platform engineer manages an NVIDIA-accelerated Kubernetes cluster running the NVIDIA GPU Operator on nodes with A100 GPUs. Several data-science teams submit training jobs, and the engineer must ensure each team's pods receive a full physical GPU exclusively, with no two pods sharing the same device. Which scheduling configuration should the engineer apply to the pod specification to guarantee exclusive whole-GPU allocation?

A.Enable MIG by setting the nvidia.com/mig.config label on the node to all-balanced and request nvidia.com/mig-1g.5gb in the pod.
B.Configure the pod to mount the hostPath /dev/nvidia0 and set privileged: true so the container can access the first GPU directly.
C.Set a nodeSelector for nvidia.com/gpu.product=A100 and rely on the scheduler to place one pod per node automatically.
D.Set resources.limits to nvidia.com/gpu: 1 and resources.requests to nvidia.com/gpu: 1, letting the NVIDIA device plugin allocate an exclusive GPU.
AnswerD

Requesting and limiting nvidia.com/gpu to 1 makes the NVIDIA device plugin treat the GPU as an indivisible integer device, so the kubelet advertises whole GPUs and the scheduler binds exactly one physical A100 exclusively to the pod. Because the extended resource is integer-only, no sharing occurs, which satisfies the exclusivity requirement without extra configuration.

Why this answer

Declaring both a request and a limit for nvidia.com/gpu equal to 1 is the canonical way to obtain an exclusive whole GPU in Kubernetes. The NVIDIA device plugin advertises GPUs as integer extended resources, so the scheduler reserves one entire device for the pod. Because extended resources cannot be fractional, no co-scheduling or sharing occurs, which precisely meets the exclusive allocation requirement.

Exam trap

The trap here is assuming that a nodeSelector or hostPath device mount can enforce exclusivity, when only the nvidia.com/gpu extended resource request actually reserves a whole physical GPU through the device plugin.

26
MCQmedium

An AI researcher is running a large-scale training job on an NVIDIA DGX system using Kubernetes. They observe that GPU utilization is consistently low despite high CPU load. Which workload management configuration is most likely to resolve this bottleneck by optimizing data pipeline throughput?

A.Increase the GPU memory limit in the pod specification.
B.Enable NVIDIA Multi-Instance GPU (MIG) for this specific job.
C.Optimize the data loader by increasing worker threads and implementing buffered prefetching.
D.Lower the batch size to reduce the memory footprint on the GPU.
AnswerC

Optimizing the data loader allows the CPU to prepare future training batches while the current batch is processing on the GPU. By increasing worker threads and utilizing prefetching, the system effectively hides I/O latency, ensuring the GPU cores are saturated with data and significantly improving overall model training efficiency.

Why this answer

Low GPU utilization often stems from data starvation where the GPU waits for the CPU to preprocess or fetch data from storage. Implementing data prefetching and increasing the number of workers in the data loader ensures the GPU remains fed with batches. This is critical in AI operations to maximize compute ROI and reduce total training time for expensive cluster resources.

Exam trap

Candidates frequently try to resolve low GPU utilization by scaling up the GPU instances or adjusting cluster scheduling policies, failing to recognize that the bottleneck is actually CPU-bound data preprocessing.

27
MCQhard

An administrator is configuring a Kubernetes cluster where some nodes have A100 GPUs and others have H100 GPUs. A training job requires specific GPU memory capacity and CUDA compute capability. Which mechanism should the administrator use to ensure the job is only scheduled onto nodes with the correct GPU model?

A.Apply a taint to nodes without the required GPU and rely on the default scheduler
B.Increase the nvidia.com/gpu resource request to match the GPU memory size
C.Set a nodeSelector matching the GPU model label applied by the GPU Operator's node feature discovery
D.Use a ResourceQuota that limits GPU requests per namespace
AnswerC

Node Feature Discovery, deployed with the GPU Operator, labels nodes with GPU model information such as nvidia.com/gpu.product. A nodeSelector referencing that label constrains the scheduler to nodes with the required GPU model, ensuring the job lands only on A100 or H100 nodes as appropriate without hardcoding node names.

Why this answer

Node Feature Discovery, part of the GPU Operator stack, labels nodes with GPU product details such as nvidia.com/gpu.product. A nodeSelector or affinity rule referencing that label restricts scheduling to nodes with the required GPU model, which is the precise way to satisfy memory and compute capability requirements without manual node naming.

Exam trap

The trap here is treating nvidia.com/gpu as a model or memory selector, when it is only a device count and cannot distinguish between A100 and H100 hardware.

28
MCQeasy

A platform team is preparing a Kubernetes cluster for AI workloads and wants the GPU device plugin, driver containers, and monitoring components deployed and kept in sync automatically on every GPU node. Which component should be installed to achieve this?

A.NVIDIA Network Operator
B.NVIDIA DCGM standalone on each node
C.NVIDIA Container Toolkit installed manually on each node
D.NVIDIA GPU Operator
AnswerD

The NVIDIA GPU Operator uses the Operator pattern to deploy and manage the GPU driver, container runtime hooks, device plugin, DCGM exporter, and related components as DaemonSets. It continuously reconciles node state, so new GPU nodes are automatically provisioned with the full software stack, which directly matches the requirement for automatic, synchronized deployment.

Why this answer

The NVIDIA GPU Operator is purpose-built to manage the entire GPU software stack on Kubernetes nodes through automated reconciliation. It deploys the driver, container toolkit configuration, device plugin, and DCGM-based monitoring, so GPU resources are advertised and maintained without manual per-node work. The Network Operator addresses networking, while DCGM and the Container Toolkit are individual pieces the Operator already orchestrates.

Exam trap

The trap here is confusing the NVIDIA Container Toolkit, which only wires the runtime for GPU access, with the GPU Operator that manages the whole stack automatically.

29
MCQmedium

An AI operations engineer manages a shared Kubernetes cluster where several teams submit GPU jobs. The engineer must prevent any single namespace from consuming all GPU capacity and must also ensure that jobs from one team cannot starve others during peak periods. Which combination of Kubernetes and NVIDIA GPU Operator features should the engineer implement?

A.Define a ResourceQuota on nvidia.com/gpu per namespace and use the GPU Operator's device plugin to advertise capacity so the quota can enforce limits.
B.Enable the MIG Manager and assign a fixed mig profile to every namespace through a LimitRange.
C.Apply a NetworkPolicy that restricts pod-to-pod traffic between namespaces and rely on it to limit GPU consumption.
D.Set a node taint for nvidia.com/gpu and add tolerations only to the highest-priority team's pods.
AnswerA

ResourceQuota can cap the total nvidia.com/gpu requests within a namespace when the device plugin advertises GPUs as extended resources. This prevents one namespace from monopolizing cluster GPU capacity. Combined with the operator's device plugin, quotas become enforceable at admission time, ensuring fair sharing across teams without manual intervention.

Why this answer

ResourceQuota is the Kubernetes mechanism that caps aggregate resource consumption per namespace, and it works for GPU extended resources once the NVIDIA device plugin advertises them. By setting a quota on nvidia.com/gpu, the administrator prevents any single namespace from consuming all GPUs, ensuring other teams retain capacity. This directly addresses both the monopolization and starvation concerns.

Exam trap

The trap here is treating NetworkPolicy, taints, or LimitRange as tools for GPU capacity fairness, when only ResourceQuota enforces an aggregate per-namespace ceiling on nvidia.com/gpu.

30
MCQeasy

Which component of the NVIDIA GPU Operator is responsible for monitoring GPU health and reporting telemetry data to the Kubernetes control plane?

A.NVIDIA Device Plugin.
B.NVIDIA DCGM Exporter.
C.NVIDIA Container Toolkit.
D.NVIDIA Node Feature Discovery.
AnswerB

The DCGM Exporter collects detailed GPU metrics and exposes them in a Prometheus-compatible format. This allows administrators to monitor GPU health, performance, and power consumption. It is the core monitoring component that bridges the gap between hardware telemetry and the observability stack in a containerized AI infrastructure.

Why this answer

The DCGM Exporter is the primary tool within the NVIDIA GPU Operator ecosystem for collecting metrics. It interacts with the Data Center GPU Manager (DCGM) to gather telemetry such as utilization, temperature, and memory health. This is vital for AI Operations because it provides the observability necessary to trigger auto-scaling, identify failing hardware, and ensure that training jobs are performing optimally within the cluster environment.

Exam trap

Candidates often confuse management components like the GPU Operator or device plugin with the specialized telemetry collection agent, mixing up scheduling logic with metrics gathering.

31
Multi-Selecthard

Which TWO strategies should an administrator implement to ensure fair resource scheduling in a multi-tenant NVIDIA cluster using Kubernetes and the NVIDIA device plugin?

Select 2 answers
A.Configure Kubernetes ResourceQuotas to limit the total number of GPUs per namespace.
B.Implement static partitioning for all GPUs in the cluster.
C.Define Kubernetes PriorityClasses to influence job preemption policies.
D.Use the NVIDIA Triton Inference Server to manage all batch processing.
E.Disable the NVIDIA device plugin to allow direct node access.
AnswersA, C

ResourceQuotas provide a hard ceiling on the aggregate compute resources a namespace can consume. This prevents a single tenant from launching an excessive number of pods that might exhaust cluster GPU capacity, ensuring that other tenants maintain access to sufficient hardware for their own operational and development needs.

Why this answer

Fair scheduling in multi-tenant environments requires both hard limits to prevent resource monopolization and priority-based mechanisms to ensure critical jobs progress. By combining ResourceQuotas for capacity management and PriorityClasses for job scheduling, admins can prevent a single user from starving the cluster while maintaining high performance for latency-sensitive inference or urgent training tasks during peak usage windows.

Exam trap

Candidates often select only software scheduling flags or generic pod limits, missing the necessary combination of namespace-level capacity caps and explicit job priority classes required for multi-tenant fairness.

32
MCQhard

An administrator runs a Kubernetes cluster with the NVIDIA GPU Operator. A data science team wants to run several small inference containers that each use only a fraction of a GPU's compute and memory, but the cluster currently assigns whole GPUs per pod. Which approach allows multiple containers to share a single physical GPU with memory isolation?

A.Enable Multi-Instance GPU mode on supported GPUs and configure the device plugin to advertise MIG instances as schedulable resources.
B.Reduce the container's nvidia.com/gpu request to a fractional value such as 0.5.
C.Set the device plugin's time-slicing configuration so that each GPU is advertised multiple times.
D.Configure the MPS control daemon to allow concurrent kernel execution across containers.
AnswerA

MIG partitions a supported GPU into hardware-isolated instances with dedicated compute and memory slices. Advertising those instances through the device plugin lets each inference container receive its own MIG device, providing true memory and fault isolation while allowing multiple containers to share one physical GPU, which matches the requirement.

Why this answer

MIG is the only mechanism here that divides a physical GPU into hardware-isolated instances with separate memory and compute slices. By enabling MIG and having the device plugin advertise each instance, the scheduler can place multiple inference containers on one GPU while keeping their memory and faults isolated.

Exam trap

The trap here is conflating time-slicing or MPS, which share a GPU without memory isolation, with MIG, which provides true hardware-level memory partitioning.

33
MCQmedium

A cluster administrator notices that GPU utilization is low despite high queue volume. After analyzing the logs, they identify that many pods are failing because they cannot access the necessary CUDA libraries. What is the most likely cause, and which component should be verified?

A.Verify the NVIDIA Device Plugin status, as it manages library injection.
B.Check the NVIDIA Container Runtime configuration for proper runtime class mapping.
C.Review Kubernetes Resource Quotas, as they restrict the total amount of GPU memory.
D.Increase the GPU memory limit to ensure enough memory for loading heavy CUDA libraries.
AnswerB

The NVIDIA Container Runtime must be configured to inject the appropriate drivers and libraries into the container. If this mapping is missing, applications will fail to load CUDA libraries, leading to runtime errors. Checking the container runtime ensures that the environment is set up for correct GPU access.

Why this answer

The most likely cause is a misconfiguration of the NVIDIA Container Runtime, which is responsible for injecting the necessary NVIDIA libraries into the container namespace. If the container runtime is not properly configured, the container cannot interact with the GPU, even if the scheduler has successfully placed it. Verifying the container runtime configuration is essential to ensure that GPU-enabled containers can access host-level drivers and CUDA stacks.

Exam trap

Candidates often blame the GPU driver version first. While important, the runtime mapping is the most common configuration error that prevents the container from 'seeing' the host's GPU capabilities.

34
MCQhard

An MLOps engineer manages a Kubernetes cluster where the NVIDIA GPU Operator runs the MIG manager. Several inference pods must each receive an isolated, fixed slice of a single A100, and the team wants the slices to survive node reboots without manual reconfiguration. Which combination of settings should the engineer apply?

A.Set the MIG manager's config to 'all-disabled' and rely on the device plugin to carve profiles dynamically per pod request.
B.Disable the MIG manager entirely and create MIG instances manually with nvidia-smi on each node, then label the nodes so pods schedule there.
C.Enable time-slicing with four replicas and set the MIG manager strategy to 'single', because replicas provide the same isolation as MIG instances.
D.Configure a MIG config labeled for the MIG manager with named profiles and set the device plugin strategy to 'mixed', so whole GPUs and MIG instances are both advertised.
AnswerD

A labeled MIG config tells the MIG manager which geometry to apply and persist, so instances are recreated after reboots without manual work. The mixed strategy lets the device plugin advertise both full GPUs and MIG instances, allowing the inference pods to request a specific MIG resource such as nvidia.com/mig-1g.5gb. This yields fixed, isolated slices with hardware-level memory and fault separation.

Why this answer

Persistent MIG slices require the MIG manager to read a labeled configuration that defines the desired geometry, which it reapplies after reboots and driver reloads. Setting the device plugin strategy to mixed lets both whole GPUs and the carved MIG instances be advertised under distinct resource names. Pods can then request a specific profile such as nvidia.com/mig-1g.5gb, obtaining fixed, hardware-isolated slices that meet the isolation and persistence requirements.

Exam trap

The trap here is believing time-slicing delivers the same hardware isolation as MIG, when it only shares one engine among workloads with no memory or fault boundaries.

35
MCQhard

An AI operations team manages a shared Kubernetes cluster where a nightly batch training workload requests nvidia.com/gpu resources and occasionally consumes all GPU memory on a node, causing a co-located interactive notebook pod to fail with CUDA out-of-memory errors. The team wants the interactive notebook to be isolated from the batch workload's memory usage without adding new hardware. Which action best achieves this on supported data center GPUs?

A.Enable time-slicing with a replica count of four so the notebook and batch workloads alternate on the same GPU in round-robin fashion.
B.Set a memory limit on the batch container using the standard Kubernetes resources.limits.memory field to cap its GPU framebuffer usage.
C.Configure Multi-Instance GPU profiles so the notebook and batch workloads run on separate MIG instances with dedicated memory partitions on the same physical GPU.
D.Add a higher PriorityClass to the notebook pod so the kube-scheduler preempts the batch workload whenever memory pressure occurs.
AnswerC

MIG slices a supported GPU into hardware-isolated instances, each with its own dedicated memory and compute resources. Placing the notebook and batch workload on separate instances prevents the batch job from consuming the notebook's memory, resolving the out-of-memory failures without procuring additional GPUs.

Why this answer

The conflict is shared GPU memory, not scheduling order. MIG is the only listed mechanism that gives each workload a dedicated, hardware-isolated memory partition on the same physical device, letting the notebook and batch job coexist without the batch job starving the notebook's framebuffer, and it requires no extra hardware.

Exam trap

The trap here is assuming that time-slicing or container memory limits isolate GPU memory, when neither partitions the device framebuffer and only MIG provides hardware-level memory separation.

36
MCQeasy

A platform team runs an on-premises Kubernetes cluster for AI inference. Several teams submit pods that request the same GPU device, and the scheduler places more pods onto a node than there are available GPUs, causing OOM errors on the device. The administrator wants the Kubernetes scheduler itself to prevent overcommitting GPUs without any custom admission controller. Which action should the administrator take?

A.Install the NVIDIA device plugin so GPUs are advertised as schedulable extended resources and add a matching resource request to each pod spec.
B.Set a node taint on each GPU node and add a toleration to pods, so only pods that explicitly opt in are placed there.
C.Enable the Kubernetes Vertical Pod Autoscaler in recommendation mode on the GPU namespaces to detect device pressure and reschedule pods.
D.Create a ResourceQuota in each team namespace limiting the count of pods that may run, so fewer pods land on GPU nodes.
AnswerA

The NVIDIA k8s-device-plugin registers each GPU as an extended resource such as nvidia.com/gpu, which makes the scheduler count devices as finite node capacity. When a pod requests one GPU, the scheduler subtracts it from allocatable capacity, so no node can be overcommitted. This directly solves the reported overplacement without writing any admission logic, because scheduling math handles accounting natively.

Why this answer

Advertising GPUs as extended resources through the NVIDIA device plugin is what lets the default scheduler treat each device as countable, finite node capacity. Pods that declare a GPU request are then placed only where an unallocated device exists, and no custom admission controller is required because the scheduler performs the arithmetic itself. Quotas, taints, and autoscalers influence eligibility or CPU and memory sizing, not device-level accounting.

Exam trap

The trap here is assuming that Kubernetes natively understands GPUs as limited resources; without the device plugin exposing them as extended resources, the scheduler treats GPU nodes as ordinary nodes and will happily overcommit devices.

37
MCQhard

A research group submits a distributed PyTorch training job spanning eight GPUs across two nodes. The job completes but produces a model with accuracy far below the single-node baseline, and logs show that several ranks started training before their peers had initialized the process group. The administrator must ensure that all ranks are launched together and that a failed rank terminates the whole job. Which combination of Kubernetes mechanisms should be used?

A.Schedule the job on a single node with eight GPUs and set the CUDA_VISIBLE_DEVICES variable per container to isolate each rank.
B.Create eight separate Deployments, one per rank, and pass the rank index through an environment variable in each Deployment manifest.
C.Deploy the job as a Kubernetes Job with parallelism set to eight and completions set to one, relying on the default pod startup ordering.
D.Use a gang-scheduling mechanism such as Volcano or the Kubeflow Training Operator with PyTorchJob and configure the rendezvous endpoint so all worker replicas are admitted together.
AnswerD

Gang scheduling admits all replicas atomically, so no rank starts until every peer is schedulable, and the PyTorchJob controller injects the master address, rank, and world size into each pod while restarting or failing the group consistently. This directly fixes the initialization race and enforces all-or-nothing execution.

Why this answer

Distributed training needs collective admission and coordinated rank metadata. A gang scheduler plus the Kubeflow Training Operator's PyTorchJob admits all replicas together and injects rendezvous details, eliminating the race where early ranks initialize before their peers, while the controller fails the whole job if any replica cannot run.

Exam trap

The trap here is treating a plain Kubernetes Job with high parallelism as equivalent to gang scheduling, when it actually offers no atomic admission or rendezvous coordination across ranks.

38
Multi-Selecthard

An administrator is tuning a Kubernetes cluster that runs many small inference pods on NVIDIA GPUs. Utilization is low because each pod reserves a full GPU while using only a fraction of its memory and compute. The administrator wants to share GPUs across pods while preserving memory-level isolation between processes. Which TWO configurations achieve this? (Choose two.)

Select 2 answers
A.Enable the NVIDIA MPS control daemon through the GPU Operator and configure pods to share a GPU with memory limits
B.Configure MIG on the GPUs and expose instances as separate resource types such as nvidia.com/mig-1g.5gb
C.Set the CUDA_VISIBLE_DEVICES environment variable in each pod to a different index of the same physical GPU
D.Enable time-slicing in the device plugin so multiple pods receive the same GPU replica with no memory limit
E.Apply a LimitRange that sets a memory limit on the nvidia.com/gpu resource in the namespace
AnswersA, B

NVIDIA MPS allows multiple CUDA processes to run concurrently on one GPU with hardware-partitioned scheduling and per-client memory limits, so several inference pods can share a device safely. Enabling it via the GPU Operator's ClusterPolicy and setting memory fractions gives the required isolation while raising utilization, which satisfies both goals.

Why this answer

MPS and MIG are the two NVIDIA mechanisms that allow several pods to share a physical GPU while maintaining isolation. MPS enforces per-client memory limits and concurrent execution under a control daemon, while MIG provides hardware-level partitioning with dedicated memory and compute slices advertised as distinct resources. Time-slicing shares a GPU without memory isolation, and environment variables or LimitRanges cannot partition GPU resources.

Exam trap

The trap here is treating time-slicing as sufficient GPU sharing, when it interleaves execution without any memory isolation between processes.

39
Multi-Selecthard

Which TWO methods are effective for enforcing GPU resource isolation in a multi-tenant NVIDIA Kubernetes environment?

Select 2 answers
A.Enabling NVIDIA Multi-Instance GPU (MIG) for hardware partitioning.
B.Configuring standard Docker cgroups for GPU memory limits.
C.Applying Kubernetes Taints, Tolerations, and Node Affinity.
D.Implementing standard OS-level priority queuing via 'nice'.
E.Setting a global environment variable for GPU frequency scaling.
AnswersA, C

MIG allows a single GPU to be carved into multiple independent instances, each with its own dedicated memory, cache, and compute cores. This provides strict hardware-enforced isolation, ensuring that one workload cannot interfere with the performance or data security of another tenant running on the same physical chip.

Why this answer

Resource isolation is paramount in multi-tenant environments to prevent noisy neighbor effects where one workload consumes disproportionate GPU cycles. NVIDIA Multi-Instance GPU (MIG) provides hardware-level isolation for partitioning, while Kubernetes-native device plugins with affinity and tolerations provide software-level scheduling control. These mechanisms combined ensure predictable performance and security, preventing cross-tenant interference during intensive training or inference cycles in shared GPU infrastructure clusters.

Exam trap

Candidates tend to pick only software scheduling or only hardware partitioning, forgetting that robust multi-tenant isolation requires a combination of both MIG and Kubernetes controls.

40
MCQmedium

An administrator notices that GPU utilization on a training cluster hovers around 25 percent even though many jobs are queued. Investigation shows that each job requests a full GPU, but the models are small and alternate between short data-loading phases and brief compute bursts. The administrator wants to increase effective GPU utilization without changing model code. Which action should be taken first?

A.Enable GPU time-slicing or MIG so multiple small jobs can share a device concurrently
B.Move data loading to CPU-only nodes to eliminate the idle phases
C.Add more GPU nodes to the cluster so queued jobs start sooner
D.Increase the batch size of each job so compute bursts last longer
AnswerA

Time-slicing lets several containers share one physical GPU through rapid context switching, and MIG provides isolated slices; both allow small, bursty jobs to occupy a device concurrently instead of idling it during data-loading phases. This raises effective utilization without modifying model code, directly addressing the observed low utilization with queued work.

Why this answer

When small, bursty jobs each hold a full GPU exclusively, the device idles during data-loading windows. Enabling time-slicing or MIG allows multiple jobs to share the device concurrently, filling those idle gaps and raising effective utilization without touching model code or adding hardware.

Exam trap

The trap here is reaching for more GPUs or bigger batches when the real issue is exclusive device occupancy by jobs that spend much of their time idle.

41
MCQmedium

An administrator supports a shared inference cluster where a single A100 GPU must serve several small models concurrently. They configure the NVIDIA device plugin with a time-slicing configuration that advertises multiple replicas of the same physical device. After deployment, users report that one noisy model starves the others and latency spikes unpredictably. Which statement best explains the observed behaviour?

A.Multi-Instance GPU mode must be enabled alongside time-slicing, because the two modes are designed to be combined on the same device.
B.The device plugin failed to register replicas, so the scheduler placed all pods on one logical device instead of distributing them.
C.Time-slicing interleaves work on one physical GPU with no memory or fault isolation, so a heavy workload can dominate device time and memory bandwidth.
D.The pods lack a RuntimeClass reference, so the NVIDIA container runtime never applies the configured replica strategy at container start.
AnswerC

Time-slicing exposes several logical replicas of one physical GPU and lets the driver rotate execution among them, but the replicas share the same memory space and compute engine. There is no partitioning of memory capacity, cache, or bandwidth, so a demanding model can monopolize execution slots and evict or slow others. That matches the reported starvation and unpredictable latency exactly.

Why this answer

Time-slicing creates multiple schedulable replicas of one physical GPU, which improves packing density but provides no hardware-level isolation. All replicas contend for the same memory, caches, and execution bandwidth, so a dominant workload can starve lighter ones and produce erratic tail latency. Where strict isolation or predictable performance is required, hardware partitioning such as Multi-Instance GPU is the appropriate mechanism instead of replica multiplexing.

Exam trap

The trap here is equating multiple advertised GPU replicas with multiple isolated GPUs; time-sliced replicas share one device's memory and compute engines and therefore cannot guarantee fair or predictable service levels.

42
MCQhard

A cluster runs mixed workloads: latency-sensitive inference services and best-effort batch jobs. Administrators observe that batch jobs occasionally occupy all GPUs, causing inference requests to queue and breach service level objectives. They want inference pods to be admitted immediately while allowing batch work to use remaining capacity and be preempted when needed. Which approach should they implement?

A.Enable time-slicing on all GPUs so batch and inference pods share devices.
B.Set a ResourceQuota on the batch namespace limiting total GPU requests.
C.Apply node affinity so inference pods only run on a dedicated subset of nodes.
D.Define priority classes and a preemption policy so inference pods have higher priority than batch jobs.
AnswerD

Kubernetes priority classes let higher-priority pods preempt lower-priority ones when resources are scarce. Assigning inference a high priority and batch a low priority ensures inference is admitted immediately, while batch jobs yield their GPUs when inference arrives. This matches the requirement that batch work uses leftover capacity and is preempted rather than blocking critical services.

Why this answer

The requirement is immediate admission for inference and preemptible use of leftover capacity by batch work. Priority classes with preemption give inference pods the ability to displace lower-priority batch pods when GPUs are scarce, while batch jobs still consume idle capacity during quiet periods. Affinity, quotas, and time-slicing each constrain or share resources but cannot prioritize or preempt running workloads.

Exam trap

The trap here is equating resource sharing or quota limits with prioritization, when only priority and preemption can guarantee immediate admission.

43
MCQhard

When running multi-instance GPU (MIG) workloads, what is the main advantage of assigning specific MIG profiles to different Kubernetes namespaces?

A.It allows the GPU to switch between different CUDA versions per instance.
B.It enhances GPU memory security by isolating workload address spaces.
C.It automatically compresses the model weights for faster loading.
D.It allows the cluster to bypass standard Kubernetes scheduler policies.
AnswerB

MIG profiles provide hardware-level isolation of memory and compute resources. By assigning specific profiles to namespaces, you effectively prevent cross-talk between different workloads' address spaces. This is a critical security and operational feature for multi-tenant environments, ensuring that sensitive data in one inference pod cannot be accessed by another.

Why this answer

MIG profiles allow for fine-grained hardware isolation, ensuring that different workloads—such as small inference tasks and large training runs—do not interfere with each other. By mapping profiles to specific namespaces, administrators can enforce strict hardware separation, preventing 'noisy neighbor' issues where one workload consumes memory bandwidth or compute cycles required by another, thus improving overall cluster stability and predictability.

Exam trap

Candidates often assume MIG is primarily for performance tuning or cost optimization, missing that its fundamental architectural strength in Kubernetes is hardware-level memory isolation between distinct, potentially insecure, user workloads.

44
MCQmedium

An ML platform team runs an NVIDIA GPU Operator-managed cluster and wants to allow multiple pods to share a single A100 GPU so that small inference services can co-reside without each consuming a whole device. The team needs a time-slicing configuration that applies to all GPU nodes in the cluster. Which approach should the administrator take?

A.Add a tolerations entry for nvidia.com/gpu to each pod and increase the kubelet's --max-pods flag on GPU nodes.
B.Set the environment variable NVIDIA_VISIBLE_DEVICES=all on each pod and let the runtime divide GPU time among containers.
C.Create a ConfigMap containing the time-slicing configuration and reference it in the ClusterPolicy so the GPU Operator propagates the device plugin config across nodes.
D.Install the NVIDIA MIG Manager and set the nvidia.com/mig.config label to all-1g.5gb on each node to subdivide the GPUs.
AnswerC

The GPU Operator supports time-slicing by reading a ConfigMap referenced in the ClusterPolicy's device plugin configuration. When applied, the operator propagates the config to all GPU nodes and restarts the device plugin, causing each physical GPU to advertise a multiplied replica count so multiple pods can share one device. This is the supported cluster-wide method.

Why this answer

Time-slicing in the NVIDIA GPU Operator is driven by a device-plugin configuration that the operator distributes to every GPU node through the ClusterPolicy. The referenced ConfigMap declares a replica count, and the device plugin then advertises that many virtual GPUs per physical device, allowing the scheduler to place multiple pods on one GPU. This is the supported, cluster-wide mechanism for enabling time-sliced sharing.

Exam trap

The trap here is confusing per-pod GPU visibility variables or MIG partitioning with time-slicing, which actually requires a device-plugin ConfigMap referenced by the ClusterPolicy.

45
MCQmedium

An organization is migrating their on-premises AI training to a hybrid cloud environment. Which component is most important to maintain consistent workload management across both the on-premises DGX systems and cloud-based GPU nodes?

A.Deploying identical physical GPU hardware across all sites.
B.Using a unified Kubernetes orchestration layer with consistent GPU Operator versions.
C.Hard-coding all GPU resource requests to match the smallest cloud GPU instance.
D.Implementing separate workload managers for cloud and on-premises sites.
AnswerB

Maintaining a unified orchestration layer allows for uniform resource management and scheduling logic across environments. By standardizing the GPU Operator and the Kubernetes API, teams can move workloads seamlessly without rewriting manifests or changing operational workflows, which is the primary challenge in managing hybrid GPU infrastructure.

Why this answer

A unified Kubernetes control plane, managed by an orchestrator like NVIDIA Base Command or a managed Kubernetes service (e.g., GKE or EKS with NVIDIA GPU support), provides a consistent API. This allows developers to use the same manifest files and CI/CD pipelines regardless of whether the physical hardware is in a local datacenter or in the cloud. It ensures that GPU scheduling, resource requests, and monitoring tools remain identical, simplifying the operations lifecycle.

Exam trap

Candidates often focus on data synchronization or network latency, missing that the primary operational hurdle in hybrid environments is inconsistent software versions across the GPU Operator and drivers.

46
MCQmedium

A researcher submits a distributed training job that spans four pods, each needing one GPU, and the pods must start together or not at all. The administrator wants Kubernetes to schedule all four pods only when four GPUs are simultaneously available. Which workload management construct should be used?

A.A Kubernetes Job with parallelism set to 4 and completions set to 4.
B.A StatefulSet with podManagementPolicy set to Parallel.
C.A DaemonSet that places one training pod on each GPU node in the cluster.
D.A PodGroup managed by a scheduler that supports gang scheduling, such as Volcano or the scheduler-plugins coscheduling plugin.
AnswerD

Gang scheduling ensures that all pods in a PodGroup are scheduled together or none are, preventing partial starts that waste GPUs and stall distributed training. This directly satisfies the requirement that the four GPU pods begin only when four GPUs are simultaneously available.

Why this answer

Distributed training needs gang scheduling so that either every worker pod is placed or none is. A PodGroup interpreted by a gang-aware scheduler like Volcano or the coscheduling plugin holds the pods until the full set of GPUs is available, avoiding deadlocks and wasted accelerator time.

Exam trap

The trap here is assuming that a standard Job or StatefulSet provides atomic, all-or-nothing scheduling, when in fact only gang-scheduling constructs enforce that guarantee.

47
MCQhard

A research team submits a multi-node training job using a `Job` with eight pods, each requesting one GPU. The cluster has eight GPU nodes, each with one A100. The administrator observes that all eight pods are spread one per node and the job runs, but throughput is far below expectations and NCCL logs show repeated fallback from GPUDirect RDMA to socket transport. Which action most directly addresses the root cause?

A.Set `NCCL_P2P_DISABLE=1` in the job's environment to force peer-to-peer transfers over PCIe.
B.Deploy the NVIDIA Network Operator to install and configure the RDMA stack and GPUDirect RDMA on the GPU nodes.
C.Add a node affinity rule forcing all eight pods onto a single node to enable NVLink communication.
D.Increase the number of replicas in the Job so more GPUs participate in the collective.
AnswerB

NCCL falling back to socket transport indicates the RDMA path is unavailable, typically because the high-speed fabric drivers, RDMA devices, and GPUDirect RDMA support are not provisioned on the nodes. The Network Operator deploys and configures those components alongside the GPU Operator, restoring the RDMA transport that NCCL prefers for multi-node collectives.

Why this answer

The symptom points to NCCL abandoning RDMA and using the slower socket transport for inter-node collectives. That path depends on the high-speed fabric drivers, RDMA devices, and GPUDirect RDMA support being installed and configured on every GPU node, which is precisely what the NVIDIA Network Operator provisions in tandem with the GPU Operator. Changing rank counts, disabling peer-to-peer, or collapsing topology onto one node does not restore the missing RDMA capability.

Exam trap

The trap here is tuning NCCL environment variables or job topology when the actual gap is that the RDMA and GPUDirect components were never deployed on the nodes.

48
MCQmedium

Which mechanism does the NVIDIA Device Plugin use to communicate GPU availability to the Kubernetes Kubelet?

A.It writes to a configuration file that the Kubelet watches.
B.It uses a gRPC-based device plugin API.
C.It queries the API server directly via REST calls.
D.It relies on the container runtime to inspect hardware.
AnswerB

The Kubernetes Device Plugin framework is built on gRPC. The NVIDIA plugin implements this API to provide the Kubelet with the necessary information about GPU resources. This standard interface allows Kubernetes to treat different hardware devices consistently while maintaining the flexibility needed for NVIDIA-specific hardware management.

Why this answer

The NVIDIA Device Plugin functions as a gRPC service that registers itself with the Kubelet upon startup. It performs active monitoring of the GPU nodes and reports discovered devices, including their count and health, to the Kubelet. This information is subsequently relayed to the Kubernetes API server, allowing the scheduler to make placement decisions based on real-time GPU availability and ensuring efficient workload distribution in AI clusters.

Exam trap

Candidates often guess standard REST APIs or custom webhooks instead of the gRPC-based device plugin API used by Kubernetes Kubelet.

49
MCQmedium

When designing a workload management strategy for multi-tenant AI training, what is the most effective way to ensure isolation between different tenants using the same physical GPU nodes?

A.Set strict Kubernetes memory limits on every pod to prevent memory leakage.
B.Implement NVIDIA MIG to partition hardware at the compute and memory level.
C.Use Kubernetes namespaces to logically separate the tenants.
D.Enable GPU sharing through the NVIDIA Container Toolkit's time-slicing configuration.
AnswerB

MIG provides spatial and temporal hardware partitioning, ensuring that individual workloads have their own dedicated hardware paths. This prevents performance degradation caused by noisy neighbors and provides security by isolating memory and compute resources, which is critical for multi-tenant clusters that require predictable performance and isolation.

Why this answer

Using NVIDIA Multi-Instance GPU (MIG) is the most robust method for hardware-level isolation. MIG partitions a single physical GPU into multiple independent instances, each with its own memory and compute resources. This allows multiple tenants to run workloads simultaneously on the same hardware without interfering with each other's performance, ensuring predictable and secure resource allocation for multi-tenant AI environments.

Exam trap

Candidates often suggest software-level resource limits (like cgroups) for GPU isolation, which do not provide the same level of hardware-enforced memory and compute partitioning as NVIDIA MIG.

50
MCQhard

When debugging a workload that consistently crashes with 'Out of Memory' (OOM) errors despite sufficient GPU VRAM, what is the most likely cause related to workload management?

A.The CUDA driver version is incompatible with the installed GPU.
B.The container memory limit is lower than the data processing requirements.
C.The GPU device plugin is not correctly reporting free memory.
D.The training job is using mixed-precision training (FP16).
AnswerB

Workloads often require substantial system memory for preprocessing before moving data to the GPU. If the container memory limit is exceeded, the orchestrator will terminate the pod with an OOM error, regardless of how much GPU VRAM is available. This is a common oversight when configuring resource limits.

Why this answer

OOM errors can occur due to host-side system memory exhaustion if the workload manages large datasets in system RAM before loading them into the GPU. If the container memory limit is set too low for the data processing pipeline, the entire container will be terminated. This highlights the need to correctly balance both GPU VRAM and system memory limits in the container specification.

Exam trap

Candidates mistakenly assume that OOM errors on GPU nodes are always caused by insufficient VRAM, completely ignoring container system memory limits during data preprocessing.

51
MCQhard

An administrator manages a cluster where inference services and batch training share the same GPU nodes. During business hours, inference pods must be scheduled promptly, while training jobs can wait. The administrator wants preemption so that a pending high-priority inference pod can evict a lower-priority training pod when no GPU is free, with evicted training resuming later. Which configuration achieves this?

A.Set the training pods' terminationGracePeriodSeconds to zero and add a PodDisruptionBudget so the scheduler can remove them immediately when capacity is needed.
B.Apply node affinity rules that pin inference pods to dedicated nodes and training pods to separate nodes, then let each workload schedule only within its own pool.
C.Create a PriorityClass with a high value for inference and a lower one for training, reference them in the pod specs, and enable preemption in the scheduler so lower-priority pods are evicted when needed.
D.Define a ResourceQuota on the training namespace capping GPU requests so that free capacity always remains available for inference pods.
AnswerC

Kubernetes priority and preemption let a pending pod with a higher PriorityClass displace running pods of lower priority when resources are scarce. Assigning a high class to inference and a lower class to training makes the scheduler evict a training pod to admit the inference pod. Combined with a group-aware training controller, the evicted job can be requeued and resume from checkpoint.

Why this answer

Priority and preemption are the native scheduling features designed for exactly this pattern: a higher-priority pending pod triggers eviction of lower-priority running pods to free capacity. Defining distinct PriorityClasses for inference and training and referencing them in pod specs enables the scheduler to make that decision. Because the training workload is managed by a group-aware controller, the evicted job can be requeued and restarted from its last checkpoint.

Exam trap

The trap here is confusing admission-time controls such as quotas and affinity with runtime preemption; only priority classes cause a pending pod to evict a running one, while quotas and affinity merely shape where or whether pods are admitted.

52
Multi-Selectmedium

An administrator is tuning a Kubernetes cluster that runs GPU Operator. Users report that GPU jobs are sometimes scheduled onto nodes whose drivers are older than the CUDA version the container needs, causing runtime failures. The administrator wants to prevent incompatible placements before pods are bound. (Choose two.)

Select 2 answers
A.Rely on the device plugin to compare the container's CUDA version against the node driver and refuse to allocate the GPU when they mismatch.
B.Enable time-slicing with a replica count high enough that every node advertises spare GPU capacity, so incompatible nodes are simply bypassed.
C.Set the container image's CUDA version to match the newest driver in the cluster and rely on the container runtime to downgrade the driver automatically.
D.Advertise the driver and CUDA versions as node labels through the GPU Operator's node feature discovery, then use nodeAffinity on the pods to require a compatible version.
E.Use a validating admission policy that rejects pods requesting a CUDA version incompatible with the target node's advertised driver labels.
AnswersD, E

Node feature discovery running under the GPU Operator publishes labels such as nvidia.com/cuda.driver.major and minor versions. Pods can then express a nodeAffinity requirement matching the needed driver generation, so the scheduler only binds them to nodes whose advertised driver satisfies the container's CUDA expectation. This moves incompatibility detection to scheduling time, preventing the runtime failures users currently see.

Why this answer

Preventing incompatible placement requires exposing driver and CUDA versions as schedulable node attributes and then constraining pods against them. Node feature discovery under the GPU Operator publishes version labels, and nodeAffinity lets pods require a compatible generation. A validating admission policy adds a second layer by rejecting pods whose declared CUDA needs exceed the target node's advertised driver, catching mistakes before binding and avoiding the runtime failures users report.

Exam trap

The trap here is expecting the device plugin or container runtime to reconcile CUDA and driver versions, when that compatibility decision belongs to scheduling and admission, not device allocation.

53
MCQmedium

A platform team runs mixed training and inference workloads on a Kubernetes cluster with the NVIDIA GPU Operator. Inference pods are latency-sensitive and must not be preempted, while training pods can be interrupted and restarted. The team wants training jobs to yield GPUs to inference jobs when capacity is scarce, without manual intervention. Which Kubernetes mechanism should the team configure to achieve this behavior?

A.A PodDisruptionBudget on the training pods that guarantees a minimum number of running replicas.
B.PriorityClass with preemptionPolicy set to PreemptLowerPriority on the inference pods.
C.ResourceQuota on the training namespace limiting total GPU requests below cluster capacity.
D.A node affinity rule on the inference pods that targets nodes with the highest available GPU memory.
AnswerB

PriorityClass assigns a numeric priority, and when a high-priority pod cannot schedule, the scheduler may evict lower-priority pods on a node to make room. Setting the inference pods to a higher priority with preemption enabled lets them displace training pods, which can restart. This directly implements automatic yielding of GPUs to latency-sensitive inference workloads without manual operator action.

Why this answer

Priority and preemption are the scheduler features designed for exactly this pattern. Assigning inference pods a higher-priority class with preemption enabled lets the scheduler evict lower-priority training pods when a node lacks free GPUs. Because training pods are restartable, the disruption is acceptable, and inference keeps its latency guarantees without manual intervention.

Exam trap

The trap here is confusing PodDisruptionBudget, which limits voluntary evictions, with preemption, which actively evicts lower-priority pods to place a higher-priority one.

54
Multi-Selectmedium

An administrator is using NVIDIA Base Command Manager to manage a cluster with a mix of GPU and CPU nodes. They need to ensure that a newly added GPU node is correctly recognized and that jobs can be scheduled on it. Which TWO actions must be performed to integrate the new node into the Base Command Manager cluster? (Choose two.)

Select 2 answers
A.Install the Base Command Manager agent on the new node and register it with the head node.
B.Define the node in the Base Command Manager cluster configuration and assign it to a node group or category.
C.Add the node's hostname and IP address to the /etc/hosts file on all cluster nodes.
D.Configure a static IP address for the node and add it to the cluster's DNS server.
E.Manually install the NVIDIA GPU driver on the node without using Base Command Manager's provisioning tools.
AnswersA, B

The Base Command Manager agent is required on each node to communicate with the head node, report status, and receive commands. Installing and registering the agent allows the head node to recognize the new node, manage its resources, and include it in the cluster's scheduling pool. Without the agent, the node remains unmanaged and cannot be utilized for jobs.

Why this answer

Integrating a new node into a Base Command Manager cluster requires installing and registering the Base Command Manager agent so the head node can manage it, and defining the node in the cluster configuration with an appropriate node group assignment so it inherits policies and becomes schedulable. These two actions together ensure the node is recognized, provisioned correctly, and available for job scheduling.

Exam trap

The trap here is assuming that manual network configuration or driver installation is necessary, when Base Command Manager handles these through its own provisioning and management framework.

55
MCQeasy

Which component in the NVIDIA AI ecosystem is responsible for monitoring and reporting GPU telemetry data, such as power usage, temperature, and utilization, to Prometheus?

A.The NVIDIA Device Plugin.
B.The NVIDIA DCGM Exporter.
C.The Kubernetes Kubelet.
D.The NVIDIA Container Runtime.
AnswerB

The DCGM Exporter is purpose-built to interface with the underlying NVIDIA driver and DCGM to gather detailed telemetry. It exposes these metrics as Prometheus-formatted data, which is essential for monitoring the health and performance of GPU workloads within a cluster and making data-driven infrastructure management decisions.

Why this answer

The NVIDIA DCGM Exporter is the standard component designed to collect GPU metrics via the Data Center GPU Manager (DCGM) and export them into a format that Prometheus can scrape. This is vital for workload management because it provides the real-time observability required to trigger auto-scaling or identify jobs that are failing to utilize allocated GPU resources effectively across the cluster.

Exam trap

Candidates often confuse the exporter with the DCGM agent itself. The agent collects data, but the exporter is the specific component that translates it for Prometheus scraping.

56
Multi-Selecthard

Which THREE factors must be considered when sizing a persistent storage solution for multi-node distributed training checkpoints?

Select 3 answers
A.Aggregate write bandwidth to avoid stalling the training loop.
B.Total storage capacity to store multiple checkpoint versions.
C.The number of CPU cores on the storage controller.
D.High availability of the storage backend to prevent data loss.
E.The number of users accessing the storage simultaneously.
AnswersA, B, D

Distributed training generates massive amounts of state data simultaneously. If the storage system cannot handle the aggregate write load, the compute nodes will stall, waiting for I/O completion. Sufficient write bandwidth is therefore non-negotiable to maintain the high performance required for large-scale GPU training clusters.

Why this answer

Checkpointing large models requires significant I/O throughput to avoid blocking the training loop, sufficient capacity for multiple historical versions, and high availability to ensure data integrity. By addressing these three factors—throughput, capacity, and reliability—organizations can ensure that the training process remains robust against hardware failures without incurring unnecessary performance penalties during the frequent write cycles required for large-scale model training.

Exam trap

Candidates often overlook aggregate throughput, focusing only on total capacity. In distributed training, having enough space is useless if the write speed is too slow, causing the GPUs to idle.

57
MCQhard

An administrator manages a Kubernetes cluster where a training job repeatedly fails with an OutOfMemory error on the GPU even though the pod requests one nvidia.com/gpu. DCGM metrics show that another pod on the same node is consuming GPU memory concurrently. GPU sharing via time-slicing is enabled cluster-wide. Which action should the administrator take to prevent this cross-pod interference while preserving the ability to share GPUs among trusted inference workloads?

A.Apply a ResourceQuota limiting nvidia.com/gpu to one per namespace
B.Set the CUDA_VISIBLE_DEVICES environment variable manually in the training pod spec
C.Disable time-slicing on the node pool hosting training jobs and use MIG or dedicated GPUs for those workloads
D.Increase the pod's nvidia.com/gpu request to two GPUs
AnswerC

Time-slicing provides no memory isolation, so concurrent pods can exhaust GPU memory and cause OutOfMemory failures. Removing training jobs from time-sliced nodes and giving them dedicated GPUs or MIG instances ensures exclusive memory access while still allowing time-slicing to remain available for trusted inference workloads on separate node pools.

Why this answer

Time-slicing multiplexes a GPU across pods without memory partitioning, so one workload can starve another and trigger OutOfMemory errors. The reliable fix is to separate untrusted or memory-intensive training workloads onto dedicated GPUs or MIG instances, while leaving time-slicing for inference workloads where memory contention is acceptable and controlled.

Exam trap

The trap here is thinking that increasing GPU requests or adding quotas creates memory isolation, when only dedicated devices or MIG instances actually partition GPU memory.

58
MCQhard

An AI ops engineer notices that a specific training workload is experiencing high 'wait' times for GPU resources despite the cluster having available idle GPUs. What is the most likely cause?

A.The GPU driver is too new for the current Kubernetes version.
B.The pod has unmet affinity or toleration requirements.
C.The system clock is unsynchronized between nodes.
D.The container image is too large to pull efficiently.
AnswerB

If a pod requires a specific node label or has a taint that the available nodes do not satisfy, it will remain in a pending state. This scenario is a common cause for pods failing to schedule, even when the underlying hardware resources appear idle to the cluster administrator.

Why this answer

High wait times despite idle resources often point to scheduling constraints, such as mismatched node affinity, taints, or tolerations. The scheduler may be unable to place the pod because the available GPUs do not meet specific requirements (e.g., specific memory requirements, architecture, or interconnect features). Troubleshooting requires examining the Kubernetes scheduler logs or pod event descriptions to identify why the pending pod cannot be bound to the available nodes.

Exam trap

Test-takers frequently assume idle GPUs mean hardware failures, overlooking scheduler constraints such as unmet node affinities, taints, or tolerations.

59
MCQmedium

An AI operations engineer manages a Kubernetes cluster running the NVIDIA GPU Operator. A team wants its long-running inference deployment to be automatically rescheduled if the GPU on a node develops an uncorrectable error that the device plugin or health checks detect. The team also wants the node to stop accepting new GPU pods until the issue is resolved. Which combination of behaviors should the engineer rely on to meet these requirements?

A.A Horizontal Pod Autoscaler scales up replicas so healthy copies absorb traffic while the faulty node remains in service.
B.The GPU Operator's health checks mark the GPU unhealthy, the device plugin stops advertising it, and Kubernetes taints or cordons the affected node so pods are rescheduled and no new GPU pods land there.
C.A liveness probe on the inference container restarts the pod on the same node when the GPU error causes a request failure.
D.A PodDisruptionBudget on the inference deployment forces the scheduler to migrate pods off the node when the GPU fails.
AnswerB

The GPU Operator runs health checks that can detect uncorrectable GPU errors and signal the device plugin to stop advertising the faulty device. The operator can also apply taints or cordon the node. Existing pods become unschedulable and are recreated elsewhere by their controller, while new GPU pods are kept off the node, matching both stated requirements.

Why this answer

Meeting both requirements needs health-driven device withdrawal plus node-level scheduling exclusion. The GPU Operator's health checks detect uncorrectable errors, the device plugin stops advertising the faulty GPU, and the node is tainted or cordoned. Controllers then recreate pods on healthy nodes, and the taint keeps new GPU pods away until the fault is cleared.

Exam trap

The trap here is expecting pod-level constructs like liveness probes or disruption budgets to handle hardware faults, when GPU error isolation is driven by device health checks and node taints.

60
MCQeasy

An AI operations engineer manages a Kubernetes cluster where the NVIDIA GPU Operator's device plugin exposes GPUs as schedulable resources. A data science team submits a batch inference job that requests one GPU but does not specify a node selector or tolerations. The job stays in Pending while other GPU nodes remain idle because they carry a taint the GPU Operator applied to reserve them for a specific workload class. Which approach is the most appropriate for the engineer to make the job schedulable without disrupting the reserved nodes?

A.Add an appropriate toleration and node selector to the job so it can target the reserved GPU nodes.
B.Delete and recreate the device plugin daemonset so GPUs are re-advertised to the scheduler.
C.Increase the cluster's GPU resource quota in the namespace so the scheduler can allocate a card.
D.Remove the node taint from all GPU nodes so the job can be scheduled anywhere.
AnswerA

A taint on a node only repels pods that do not tolerate it. Adding the matching toleration and a node selector that identifies the reserved GPU nodes lets this job schedule there while every other untolerated workload stays off those nodes. The reservation remains intact and the job becomes runnable, which directly resolves the Pending state.

Why this answer

A node taint repels pods unless they carry a matching toleration. The reserved GPU nodes are intentionally tainted, so a job that neither tolerates the taint nor selects an untainted node cannot schedule. Adding the correct toleration plus a node selector places the job on the reserved nodes while preserving the reservation for other classes of work.

Exam trap

The trap here is assuming a Pending GPU pod is caused by missing GPU resources or quota, when the actual cause is an un-tolerated node taint.

61
Multi-Selectmedium

A research team is submitting many short-lived experiment jobs to an NVIDIA-accelerated Kubernetes cluster. The operations team wants to reduce GPU idle time and improve overall utilization without modifying the training code. Which TWO approaches should the operations team implement? (Choose two.)

Select 2 answers
A.Pin each experiment job to a dedicated physical GPU using node affinity
B.Use a gang-scheduling or batch scheduler with backfill so queued jobs start as soon as GPUs free up
C.Disable the NVIDIA device plugin and mount GPUs directly via hostPath
D.Increase the pod's nvidia.com/gpu request to reserve more GPU memory
E.Enable GPU sharing with time-slicing so multiple experiment pods can occupy the same GPU
AnswersB, E

Batch schedulers with backfill can start smaller jobs ahead of larger queued jobs when resources are available, keeping GPUs busy between experiment runs. This reduces idle gaps and improves throughput for many short-lived jobs, complementing GPU sharing without requiring modifications to the training scripts.

Why this answer

Short-lived experiment jobs create idle GPU gaps. Time-slicing allows multiple such pods to share a device, and a batch scheduler with backfill keeps the queue moving so GPUs are assigned as soon as capacity frees. Both measures raise utilization without touching training code, which is exactly what the operations team needs.

Exam trap

The trap here is assuming that dedicating a GPU per short job or requesting more GPUs improves utilization, when in fact both reduce concurrency and leave GPUs idle.

62
MCQhard

Refer to the exhibit. A cluster administrator notices that GPU jobs with this PriorityClass are failing to start even when empty GPUs are available. What is the most likely cause?

A.The PriorityClass value is too low to trigger preemption.
B.Insufficient nodes satisfy the 'minAvailable' requirement for the pod group.
C.The NVIDIA Device Plugin requires a restart to recognize the PriorityClass.
D.The PriorityClass name violates Kubernetes naming conventions.
AnswerB

In high-performance computing, Volcano requires that all pods within a group start simultaneously. If the requested number of GPUs cannot be satisfied across available nodes due to locality or capacity constraints, the job remains pending, even if individual nodes have free GPUs, to prevent incomplete distributed training runs.

Why this answer

When using the Volcano scheduler with NVIDIA GPUs, priority classes alone are insufficient if the Gang Scheduling policy is not met. If the job requests multiple pods that cannot be satisfied simultaneously, Volcano will leave the pods in a pending state to avoid resource fragmentation. This ensures that massive training jobs do not consume isolated resources that would otherwise result in deadlocks during collective communication synchronization.

Exam trap

Candidates often ignore the 'minAvailable' parameter in Volcano, assuming that as long as individual GPUs are free, the job will eventually start, ignoring the gang scheduling requirement.

63
MCQhard

An AI operations team runs long-running training jobs on a Kubernetes cluster with NVIDIA GPU Operator. They observe that after a node is rebooted for maintenance, some pods resume but report CUDA 'unknown error' and the device plugin shows unhealthy GPUs. Which configuration should the administrator review to ensure the driver and device plugin recover cleanly after reboot?

A.The pod's restartPolicy set to Always so containers restart automatically after node recovery.
B.The GPU Operator's driver validation and node reboot handling settings, including the driver DaemonSet and readiness gates.
C.The cluster autoscaler settings so replacement nodes are provisioned immediately after reboot.
D.The container image's CUDA version to ensure it matches the host driver version.
AnswerB

After a reboot, the GPU Operator must reload the driver and reinitialize the device plugin before workloads can use GPUs safely. Driver validation and readiness gates prevent pods from starting until the driver and plugin report healthy. Reviewing these settings addresses the CUDA error and unhealthy device plugin state observed after maintenance reboots.

Why this answer

A reboot invalidates the previously loaded driver and device plugin state, so the GPU Operator must reload the driver and re-register devices before pods can use them. Driver validation and readiness gates hold workloads until health checks pass, preventing CUDA errors from reaching applications. Reviewing these settings ensures clean recovery and avoids the unhealthy plugin state observed after maintenance.

Exam trap

The trap here is blaming the application container or its CUDA version for a fault that originates in post-reboot node initialization by the GPU Operator.

64
MCQmedium

Which of the following is the most appropriate workload management technique for a bursty AI inference workload that requires low latency but does not need full GPU power for every request?

A.Assigning one full physical GPU to every single inference request.
B.Using MIG or fractional GPU sharing to maximize hardware throughput.
C.Configuring the GPU to run in 'Maximum Performance' mode at all times.
D.Implementing a strict FIFO queue for all incoming inference requests.
AnswerB

MIG and fractional sharing allow multiple inference workloads to run concurrently on a single GPU. This effectively increases the number of available 'virtual' GPUs, maximizing utilization and ensuring that small, bursty requests can be handled with minimal latency, which is essential for cost-effective, high-scale inference services.

Why this answer

NVIDIA Multi-Instance GPU (MIG) combined with time-slicing or fractional GPU sharing is ideal for inference workloads. By partitioning GPUs, the system can handle many smaller, latency-sensitive requests efficiently. This approach balances the need for high-performance hardware with the requirement for multi-tenancy, ensuring that no single inference request monopolizes the entire GPU while maintaining the fast response times required by the application.

Exam trap

Candidates often mistake general load balancing or horizontal scaling for GPU-level resource management, failing to recognize that MIG is the specific hardware-level solution for partitioning GPUs for multi-tenant inference.

65
MCQmedium

A platform team runs a Kubernetes cluster where the NVIDIA GPU Operator is installed and time-slicing is configured with a ConfigMap that advertises four replicas per physical GPU. A data scientist submits a PyTorch training pod requesting nvidia.com/gpu: 1. The pod stays Pending indefinitely, and the scheduler event reads 'Insufficient nvidia.com/gpu'. The node's GPUs are otherwise idle and healthy. Which action most directly resolves the pending state?

A.Restart the nvidia-device-plugin pod so it re-reads the time-slicing ConfigMap and republishes the replicated resource capacity to the kubelet.
B.Add a nodeSelector that pins the pod to a specific GPU node, because the scheduler cannot match GPU requests without an explicit node affinity rule.
C.Verify the time-slicing ConfigMap is labeled for the device plugin and that the node's GPU capacity now reports the multiplied replica count, then requeue the pod.
D.Change the pod's resource request to nvidia.com/gpu.shared: 1 so it matches the resource name that time-slicing exposes.
AnswerC

Time-slicing only takes effect when the ConfigMap carries the label the device plugin watches and the plugin successfully reloads it. Confirming the node's allocatable nvidia.com/gpu reflects replicas times physical GPUs proves the configuration propagated. If capacity is still one per GPU, the ConfigMap is unlabeled or malformed, and correcting that plus requeuing the pod restores schedulability.

Why this answer

Time-slicing is enabled by a ConfigMap that the NVIDIA device plugin watches; without the expected label the plugin ignores it and keeps advertising one GPU per device. Confirming the ConfigMap is labeled and that the node's allocatable nvidia.com/gpu equals replicas multiplied by physical GPUs verifies the change propagated. Once capacity is correctly republished, the pending pod can be scheduled without altering its resource request.

Exam trap

The trap here is assuming time-slicing creates a distinct resource name such as nvidia.com/gpu.shared, when it actually multiplies the existing nvidia.com/gpu count.

66
Multi-Selectmedium

An AI operations engineer is troubleshooting a Kubernetes cluster where several GPU training pods fail to start with a device plugin allocation error, even though the nodes report healthy GPUs. The engineer suspects the pods are requesting more GPU resources than a single physical card can provide without a sharing mechanism. Which TWO configurations would legitimately allow multiple pods to consume a single physical GPU on these nodes? (Choose two.)

Select 2 answers
A.Enable NVIDIA Multi-Instance GPU (MIG) on supported GPUs and expose the MIG instances as schedulable resources.
B.Configure time-slicing in the device plugin so several pods share a GPU through interleaved execution.
C.Add a node selector that allows multiple pods to bind to the same nvidia.com/gpu resource slot.
D.Create a ResourceQuota that counts GPU requests as fractional values such as 0.5 per pod.
E.Raise the GPU Operator's driver version so the device plugin reports extra virtual GPUs per card.
AnswersA, B

MIG partitions a supported GPU into isolated instances, each with dedicated memory and compute slices. When the GPU Operator exposes these instances through the device plugin, each MIG instance is advertised as its own resource, so multiple pods can run concurrently on one physical card with hardware-level isolation. This is a supported way to share a single GPU across pods.

Why this answer

Both MIG and time-slicing change how the device plugin advertises and allocates a physical GPU, enabling concurrent consumption by more than one pod. MIG offers hardware-partitioned, isolated instances, while time-slicing interleaves workloads on the whole card. Each is a supported sharing mode configured through the GPU Operator, unlike driver upgrades, node selectors, or quota edits, which do not alter device advertisement.

Exam trap

The trap here is believing that raising a driver version or adjusting scheduling hints can create shareable GPUs, when only explicit device plugin sharing modes change resource advertisement.

67
MCQeasy

A data science team submits a PyTorch distributed training job to a Kubernetes cluster with the NVIDIA GPU Operator installed. The job's pods repeatedly fail with a CUDA initialization error, while a simple `nvidia-smi` check inside an interactive pod on the same node succeeds. The administrator confirms the node's driver is healthy and the device plugin is advertising GPUs. Which configuration should the administrator verify first?

A.That the node's `/etc/docker/daemon.json` still lists `nvidia` as the default runtime for all containers.
B.That the pods' securityContext drops the `IPC_LOCK` capability required for pinned host memory.
C.That the training pods' containers were built with a CUDA toolkit version newer than the node driver supports.
D.That each training pod includes a GPU resource request or limit so the device plugin injects the driver libraries and device nodes.
AnswerD

The GPU Operator advertises devices through the device plugin, and the kubelet only mounts the driver libraries, device nodes, and CUDA binaries into a container that requests nvidia.com/gpu. Without that request the container starts with no GPU access, so CUDA initialization fails while nvidia-smi in a GPU-requesting pod on the same node works, matching the observed behavior exactly.

Why this answer

In an operator-managed cluster, GPU access is granted during admission and kubelet setup only when a pod requests an extended GPU resource; the device plugin and admission controller then inject the driver libraries, CUDA binaries, and device nodes. A pod without that request runs with no visible device, which is precisely why CUDA initialization fails while nvidia-smi succeeds in a different pod on the same node. The other options concern version skew, a specific capability, or a legacy runtime setting that would not produce this selective symptom.

Exam trap

The trap here is chasing driver or toolkit version mismatches when the real cause is that the workload never requested a GPU, so the device was never injected into the container.

68
MCQhard

A research organization runs an NVIDIA DGX SuperPOD with a Kubernetes cluster managed by the NVIDIA GPU Operator and Network Operator. A distributed training job using PyTorch DDP across 32 nodes stalls at initialization, and the administrator suspects the collective communication library is not selecting the high-speed fabric. Which configuration should the administrator verify first to ensure NCCL uses the correct network interface and topology?

A.Increase the pod's nvidia.com/gpu limit to 8 so each node exposes all GPUs to the training process.
B.Set the CUDA_VISIBLE_DEVICES variable to list all GPUs and restart the training job.
C.Enable the NVIDIA MIG feature on all nodes so each rank gets an isolated GPU slice for communication.
D.Confirm that NCCL_IB_DISABLE is set to 0, NCCL_SOCKET_IFNAME matches the high-speed fabric interface, and NCCL_TOPO_FILE or the topology-aware plugin is loaded on each node.
AnswerD

NCCL relies on these environment variables and topology data to select InfiniBand or RoCE interfaces and to build the correct ring or tree topology across nodes. If NCCL_IB_DISABLE is set to 1, or NCCL_SOCKET_IFNAME points at the management interface, NCCL falls back to slower paths and initialization can stall. Verifying these values is the primary diagnostic step.

Why this answer

NCCL chooses transports and interfaces based on environment variables and detected topology. When NCCL_IB_DISABLE is set incorrectly or NCCL_SOCKET_IFNAME points to the wrong interface, collectives fall back to TCP over the management network or fail to connect, causing distributed jobs to hang at initialization. Verifying these variables and the topology file on every node is the correct first diagnostic step.

Exam trap

The trap here is assuming that adding GPUs per pod or adjusting CUDA_VISIBLE_DEVICES will fix a distributed hang, when the real cause is NCCL's network interface and fabric selection.

69
MCQhard

Which THREE factors should be considered when estimating GPU memory requirements for a Large Language Model (LLM) fine-tuning job?

A.The memory size of the model parameters.
B.The precision used for optimizer states.
C.The activation memory during the forward pass.
D.The total number of CPU threads used.
E.The network latency between nodes.
AnswerA, B, C

Model weights occupy a significant portion of GPU memory. For large models, this is often the baseline requirement. When determining the cluster footprint, the total size of these parameters must be considered, especially if using techniques like sharding, which distribute these weights across multiple GPUs in a cluster.

Why this answer

When sizing LLM workloads, you must account for the model weights, optimizer states, and gradient buffers. Additionally, activations consume significant memory during the forward and backward passes. Understanding these components is essential for AI Ops, as incorrect sizing leads to OOM crashes early in the training process, wasting significant compute time and delaying model delivery in production environments.

Exam trap

Candidates frequently focus only on model weights, ignoring the significant memory overhead consumed by activation buffers during the forward pass, which often causes OOM errors in large training jobs.

70
MCQmedium

An AI operations engineer manages a shared Kubernetes cluster running NVIDIA GPU Operator. Several teams report that their inference pods remain in a Pending state with the event '0/8 nodes are available: 8 Insufficient nvidia.com/gpu.' The administrator verifies that nvidia-smi on all nodes shows idle GPUs. Which action should the administrator take first to resolve the scheduling failure?

A.Reinstall the NVIDIA GPU Operator and restart the kubelet on all worker nodes.
B.Inspect the node allocatable GPU count and any taints or labels that prevent scheduling on the idle nodes.
C.Add a nodeSelector for kubernetes.io/os=linux to the pod template.
D.Increase the GPU memory limit in the pod specification so the scheduler can fit the workload.
AnswerB

The event 'Insufficient nvidia.com/gpu' means the scheduler sees zero or fewer allocatable GPUs than requested on every candidate node. This can result from missing device plugin advertisements, taints, node selectors, or nodes excluded by affinity. Checking allocatable resources and taints directly identifies why idle hardware is invisible or ineligible to the scheduler.

Why this answer

The Pending event shows the scheduler believes no node has a free nvidia.com/gpu resource, even though nvidia-smi reports idle devices. That gap is almost always caused by the device plugin not advertising GPUs, or by taints, labels, or affinity rules excluding the nodes. Inspecting allocatable counts and scheduling constraints identifies the actual blocker before any disruptive remediation.

Exam trap

The trap here is assuming that idle GPUs visible in nvidia-smi automatically become schedulable Kubernetes resources without the device plugin advertising them.

71
Multi-Selectmedium

Which TWO of the following are benefits of using containerized GPU workloads compared to bare-metal deployment?

Select 2 answers
A.Easier management of heterogeneous software dependencies.
B.Significant increase in raw GPU compute performance.
C.Improved portability across development and production environments.
D.Automatic elimination of GPU driver compatibility issues.
E.Direct access to the GPU firmware for kernel customization.
AnswersA, C

Containers allow each workload to bundle its own specific versions of libraries like CUDA and PyTorch. This avoids conflicts on the host system, where different training jobs might otherwise require incompatible driver versions or shared library dependencies, enabling more efficient sharing of the same underlying physical node.

Why this answer

Containerization provides portability, version control, and dependency isolation, which are essential for managing complex AI stacks. By packaging the environment, developers ensure that models run consistently across development, testing, and production clusters. This significantly reduces 'environment drift' and simplifies the management of different CUDA, cuDNN, and framework versions required by various research projects within the same shared infrastructure cluster.

Exam trap

Candidates often select 'increased performance' as a benefit. Containerization adds a slight abstraction layer and does not inherently increase raw GPU compute performance compared to bare-metal execution.

72
MCQhard

An AI operations team runs a shared Kubernetes cluster with the NVIDIA GPU Operator and several namespaces owned by different groups. A group reports that its training pods are stuck Pending with an event indicating insufficient nvidia.com/gpu, yet cluster-wide dashboards show many GPUs idle. Investigation reveals the idle GPUs belong to nodes labeled for another group, and the affected namespace has a node affinity rule pinning its pods to a specific GPU generation that is fully consumed. Which action best resolves the Pending pods while respecting multi-tenant boundaries?

A.Disable the device plugin on the other tenants' nodes so their GPUs become available to the affected namespace.
B.Lower the GPU request in the pod spec so the scheduler accepts a partial device from an idle node.
C.Remove the node affinity rule so the pods can schedule onto any idle GPU node in the cluster.
D.Provision additional GPUs of the required generation for that tenant or rebalance existing capacity to that node pool.
AnswerD

The pods are Pending because the specific GPU generation they require is exhausted, not because the cluster lacks GPUs entirely. Adding capacity to that node pool, or moving idle cards of the right generation into it, satisfies the affinity constraint while keeping tenants separated. This addresses the real bottleneck without weakening isolation rules.

Why this answer

The pods are constrained by a node affinity rule targeting a specific GPU generation whose nodes are fully allocated. Idle GPUs on other nodes cannot satisfy that rule. The correct fix is to add capacity of the required generation or rebalance matching cards into that pool, which resolves the shortage while preserving the tenant isolation the affinity rule enforces.

Exam trap

The trap here is treating idle cluster-wide GPUs as available capacity, when node affinity can make those GPUs ineligible for the pending pods.

73
MCQmedium

When managing GPU resources in a shared cluster, which configuration best prevents 'noisy neighbor' scenarios where one GPU task consumes all available memory bandwidth?

A.Increasing the CUDA_VISIBLE_DEVICES environment variable.
B.Implementing NVIDIA MIG and strictly defining resource limits.
C.Setting a higher priority class for inference pods.
D.Configuring the NVIDIA DCGM Exporter alerts.
AnswerB

MIG partitions the GPU at the hardware level, providing dedicated memory and compute paths. When combined with Kubernetes resource limits, it ensures that workloads are physically constrained to their slice, preventing a single process from monopolizing memory bandwidth or compute cycles, thus effectively eliminating the noisy neighbor problem.

Why this answer

Using a combination of NVIDIA MIG and Kubernetes resource quotas provides the strongest isolation. By enforcing hardware-level partitioning via MIG, you ensure that memory bandwidth is strictly bounded for each slice. This is vital for AI Ops because it prevents high-throughput training jobs from starving small inference requests, ensuring stable latency across the entire cluster environment and improving total system reliability.

Exam trap

Candidates often choose software-level container limits or basic Kubernetes namespaces alone, forgetting that true bandwidth isolation requires hardware-level partitioning like NVIDIA MIG to prevent high-throughput tasks from monopolizing shared memory paths.

74
MCQeasy

When managing large-scale model training jobs, what is the primary purpose of using a Job Scheduler like Slurm or Kubernetes Batch?

A.To increase the clock speed of individual GPU cores.
B.To automate resource allocation, job queuing, and throughput optimization.
C.To provide a direct IDE interface for writing model code.
D.To replace the need for containerization technologies.
AnswerB

Schedulers are essential for maximizing cluster utilization by managing queues and resource mapping. They automate the lifecycle of compute tasks, ensuring that jobs are placed on nodes with the required hardware specifications, thereby optimizing throughput and ensuring that expensive GPU resources are consistently kept productive.

Why this answer

Job schedulers serve as the orchestration layer to queue, prioritize, and allocate computational resources based on policy. In AI operations, they ensure that high-priority training runs receive the necessary GPU throughput while lower-priority jobs wait. By automating the allocation process, schedulers prevent resource idleness and manage contention, allowing data scientists to focus on model development rather than manual infrastructure management or resource conflict resolution during heavy cluster usage.

Exam trap

Candidates often focus on the model training code itself rather than the orchestration layer, failing to realize that job schedulers are essential for preventing resource contention in large clusters.

75
MCQeasy

Which NVIDIA technology allows for partitioning a single physical GPU into multiple independent instances, each with dedicated compute and memory resources for smaller workloads?

A.CUDA Streams
B.Multi-Instance GPU (MIG)
C.NVIDIA Collective Communications Library (NCCL)
D.NVIDIA Container Runtime
AnswerB

MIG enables the hardware-level partitioning of GPUs. By creating distinct instances with dedicated compute, memory, and cache, it provides fault isolation and guaranteed performance, which is vital for modern multi-tenant AI clusters that need to support varying workload sizes efficiently without resource contention or cross-tenant interference.

Why this answer

Multi-Instance GPU (MIG) is a feature in NVIDIA A100 and H100 architectures that enables the partitioning of a physical GPU into up to seven separate instances. This is essential for AI Ops to maximize resource utilization by running multiple inference tasks on a single GPU without interference, ensuring Quality of Service for each instance while preventing a single process from consuming all hardware resources.

Exam trap

Candidates often confuse MIG with general virtualization or container-based GPU sharing, failing to realize MIG is a hardware-level partitioning technology specific to Ampere and newer NVIDIA architectures for isolated, secure workloads.

Page 1 of 2 · 87 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Workload Management questions.