Courseiva
← Back to NVIDIA Certified Professional: AI Operations questions

Scenario-based practice

Select Two (Multi-Select) Questions

Practise NVIDIA Certified Professional: AI Operations practice questions — original exam-style scenarios covering every exam domain, with detailed explanations, wrong-answer analysis, and common exam traps.

20
scenario questions
NCP-AIO
exam code
NVIDIA
vendor

Scenario guide

How to approach select two (multi-select) questions

Multi-select questions tell you to 'Choose TWO' or 'Choose THREE'. Getting partial credit is not a thing — you must select all correct answers with no incorrect ones. The stem always states how many to choose, so trust it. These questions require precision, not best-guess elimination.

Quick answer

Select Two (Multi-Select) Questions questions test whether you can apply the concept in context, not just recognise a definition.

How the topic appears in realistic exam-style scenarios.

Which detail in the question changes the correct answer.

How to eliminate plausible but wrong options.

How to connect the question back to the wider exam objective.

Related practice questions

Related NCP-AIO topic practice pages

Scenario questions usually connect to one or more exam topics. Use these links to review the underlying concepts behind the scenario.

Practice set

Practice scenarios

Question 1mediummulti select
Full question →

An administrator is optimizing multi-GPU utilization. Which TWO of the following configurations allow multiple containers to share a single physical GPU on a supported NVIDIA architecture?

Question 2mediummulti select
Full question →

An administrator is optimizing a cluster for AI model training using NVIDIA Base Command. Which TWO tasks are critical for ensuring consistent performance across the training nodes?

Question 3hardmulti select
Full question →

Which THREE components are required for a container to successfully leverage NVIDIA GPUs on a Kubernetes cluster?

Question 4hardmulti select
Full question →

An AI researcher is debugging a multi-node training job using NCCL. Which TWO actions should they take to diagnose potential network-related performance degradation?

Question 5mediummulti select
Full question →

Which TWO of the following are benefits of using containerized GPU workloads compared to bare-metal deployment?

Question 6mediummulti select
Full question →

An administrator is deploying the NVIDIA GPU Operator on a Kubernetes cluster in an air-gapped environment. The cluster nodes have no internet access, and all container images must be pulled from a private registry. Which two actions are required to ensure a successful deployment? (Choose two.)

Question 7hardmulti select
Full question →

A research lab is deploying NVIDIA AI Enterprise on an air-gapped Kubernetes cluster. The cluster has no internet access, and all software must be installed from a local registry. The administrator plans to use the NVIDIA GPU Operator. Which two actions must be performed to ensure a successful deployment in this environment? (Choose two.)

Question 8mediummulti select
Full question →

A research team is submitting many short-lived experiment jobs to an NVIDIA-accelerated Kubernetes cluster. The operations team wants to reduce GPU idle time and improve overall utilization without modifying the training code. Which TWO approaches should the operations team implement? (Choose two.)

A cloud architect is deploying NVIDIA AI Enterprise on a vSphere cluster with multiple ESXi hosts, each containing NVIDIA A100 GPUs. The architect must enable GPU virtualization using NVIDIA vGPU. Which two components are required to support vGPU on the ESXi hosts? (Choose two.)

Question 10hardmulti select
Full question →

An AI researcher is deploying a multi-node training job using NCCL on an InfiniBand network. The job is suffering from intermittent latency spikes. Which TWO steps should the engineer perform to troubleshoot the network configuration?

Question 11mediummulti select
Full question →

An administrator is preparing to deploy an NVIDIA AI Enterprise solution on an OpenShift cluster. Which TWO steps must be completed to ensure the NVIDIA drivers are loaded correctly on the worker nodes?

Question 12mediummulti select
Full question →

Which TWO of the following actions are recommended for optimizing NVIDIA GPU utilization during a high-concurrency inference deployment?

Question 13hardmulti select
Full question →

An AI operations team is validating a new Kubernetes cluster before installing the NVIDIA GPU Operator with the driver managed by the Operator itself. The nodes run a supported Linux distribution with GPUs physically installed. Which two conditions must be satisfied for the Operator's driver container to build and load the kernel module successfully? (Choose two.)

Question 14hardmulti select
Full question →

An AI infrastructure team is deploying NVIDIA AI Enterprise on a Kubernetes cluster using the NVIDIA GPU Operator. They need to ensure that the GPU Operator can successfully manage GPUs and that workloads can consume GPU resources. Which two components does the GPU Operator deploy to enable GPU scheduling and container GPU access? (Choose two.)

Question 15hardmulti select
Full question →

A platform team operates a Kubernetes cluster where several teams submit GPU training jobs. The administrator needs to enforce per-namespace limits on the number of GPUs that can be consumed and prevent a single namespace from monopolizing all GPU capacity. Which TWO Kubernetes resources should be configured to achieve this? (Choose two.)

Question 16mediummulti select
Full question →

An administrator is using NVIDIA Base Command Manager to manage a cluster with a mix of GPU and CPU nodes. They need to ensure that a newly added GPU node is correctly recognized and that jobs can be scheduled on it. Which TWO actions must be performed to integrate the new node into the Base Command Manager cluster? (Choose two.)

Question 17mediummulti select
Full question →

A team trains a model inside an NGC PyTorch container on a DGX H100 node. Training starts, but after a few minutes the process dies and `dmesg` shows `Xid 79: GPU has fallen off the bus` on one GPU. The team needs to determine whether the fault is hardware or software before opening an RMA. Which two actions should they take to gather useful evidence? (Choose two.)

Question 18mediummulti select
Full question →

An AI operations engineer is troubleshooting a Kubernetes cluster where several GPU training pods fail to start with a device plugin allocation error, even though the nodes report healthy GPUs. The engineer suspects the pods are requesting more GPU resources than a single physical card can provide without a sharing mechanism. Which TWO configurations would legitimately allow multiple pods to consume a single physical GPU on these nodes? (Choose two.)

Question 19hardmulti select
Full question →

An administrator manages a Kubernetes cluster where the NVIDIA GPU Operator has deployed the device plugin and MIG Manager. A tenant wants to run several small inference services that each need only a fraction of a GPU, isolated from other tenants' memory and fault domains. The administrator decides to use Multi-Instance GPU mode. Which TWO actions must be performed to make MIG-backed GPU resources schedulable to those pods? (Choose two.)

Question 20hardmulti select
Full question →

An AI operations engineer is investigating a training job on an NVIDIA DGX system that intermittently fails with 'uncorrectable ECC error' on a GPU. The job is using NCCL for multi-GPU communication. The engineer needs to identify the appropriate immediate actions to diagnose and mitigate the issue. (Choose two.)

These NCP-AIO practice questions are part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style NCP-AIO questions with detailed explanations, topic-based practice, mock exams, readiness tracking, and study analytics.