NCP-AIO · domain
scenario questions
Practise NVIDIA Certified Professional: AI Operations scenario questions practice questions — original exam-style scenarios with answer choices, explanations, and analysis of common mistakes.
Focused practice
Practice scenario questions questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about scenario questions
scenario questions questions test whether you can apply the concept in context, not just recognise a definition.
How the topic appears in realistic exam-style scenarios.
Which detail in the question changes the correct answer.
How to eliminate plausible but wrong options.
How to connect the question back to the wider exam objective.
Watch out for
Common scenario questions exam traps
- ▸Answering from memory before reading the full scenario.
- ▸Missing a constraint such as cost, availability, security, scope or command context.
- ▸Choosing a broad answer when the question asks for the most specific fix.
- ▸Ignoring why the wrong options are tempting.
Question index
All scenario questions questions (309)
Click any question to see the full explanation, or start a practice session above.
During a multi-GPU training job, you notice that one GPU consistently reports lower utilization and longer communication times compared to others. What is the most likely reason for this performance imbalance?
Medium2An AI researcher is using Nsight Systems to profile an application. They notice a large gap in the timeline where neither the CPU nor the GPU is performing significant work. What does this gap most likely represent?
Medium3An administrator is preparing a cluster for a new large language model training job that will use NVIDIA Magnum IO GPUDirect Storage to stream training data directly from a parallel file system to GPU memory. The administrator must verify that the environment supports GPUDirect Storage before the job starts. (Choose two.)
Medium4A media company is deploying an inference service on a Kubernetes cluster with the NVIDIA GPU Operator installed. The service pods remain in Pending with the message that no nodes have the requested nvidia.com/gpu resource, even though the GPUs are healthy and the driver loads correctly on every node. Which troubleshooting step should the engineer perform first?
Medium5An administrator is setting up NVIDIA Base Command Manager to provision and manage a new AI cluster. They need to ensure that the cluster can automatically discover and configure new GPU nodes. Which component is responsible for node discovery and initial configuration?
Easy6A DevOps engineer is deploying the NVIDIA GPU Operator on a Kubernetes cluster that uses containerd as the container runtime. The engineer notices that the Operator's validation pod fails with an error indicating that the NVIDIA container runtime is not configured. Which action should the engineer take to resolve this?
Easy7What is the most accurate way to verify that a training job is utilizing Tensor Cores?
Medium8A research team runs a multi-node distributed training job spanning eight GPU nodes. Jobs frequently begin execution before all worker pods are running, and the collective initialization hangs until the operator manually scales the job down and up. The administrator wants the scheduler to admit the job only when all of its pods can be placed together. Which mechanism should be used?
Hard9When installing the NVIDIA Container Toolkit to enable GPU acceleration in Docker, which file must be modified or verified to ensure the container runtime can access the NVIDIA runtime?
Easy10A systems administrator is installing the NVIDIA Container Toolkit on a stand-alone server running Ubuntu 22.04 to enable Docker containers to access NVIDIA GPUs. After installation, they run a test container and find that it cannot see the GPU. Which step is most likely missing?
Easy11An AI operations engineer is deploying a model on an NVIDIA GPU and notices that inference latency is higher than expected. The model uses a batch size of 1, and profiling shows that the GPU is idle between kernel launches. Which optimization technique should the engineer use to reduce latency by overlapping data transfer with computation?
Easy12A platform engineer is preparing an Ubuntu 22.04 server that will host GPU-accelerated inference containers managed by containerd (not Docker). The team wants the NVIDIA Container Toolkit to expose GPUs to those containers. After installing the toolkit packages, which action must the engineer take so that containerd actually invokes the NVIDIA runtime for GPU workloads?
Medium13A production inference service using TensorRT is showing lower than expected throughput. Profiling shows that the model is spending significant time in "host-to-device" transfers. What is the best optimization?
Hard14Which TWO of the following actions should be taken to optimize GPU memory usage when encountering Out-of-Memory (OOM) errors during model training?
Hard15An administrator manages a shared NVIDIA cluster where several teams run inference services. One team's pods are being evicted repeatedly, and DCGM metrics show the node's GPUs are healthy but memory on the devices is nearly exhausted. The team insists their model fits. Which action should the administrator take FIRST to identify the cause?
Hard16Which scheduling strategy is recommended to maximize the efficiency of long-running training jobs on preemptible instances?
Medium17A team is deploying a large language model for inference using NVIDIA Triton Inference Server on a GPU. They observe that the first inference request has high latency compared to subsequent requests. What is the most likely cause and the appropriate optimization?
Medium18An administrator manages a Kubernetes cluster where the NVIDIA GPU Operator has deployed the device plugin and MIG Manager. A tenant wants to run several small inference services that each need only a fraction of a GPU, isolated from other tenants' memory and fault domains. The administrator decides to use Multi-Instance GPU mode. Which TWO actions must be performed to make MIG-backed GPU resources schedulable to those pods? (Choose two.)
Hard19A cloud operations team is using NVIDIA AI Enterprise with Kubernetes to deploy inference workloads. They want to ensure that GPU resources are allocated to pods only when explicitly requested, and that pods without GPU requests do not consume GPU resources. Which Kubernetes feature should they use to enforce this behavior?
Easy20An administrator is setting up an NVIDIA AI Enterprise cluster and wants to verify that the NVIDIA GPU Operator has successfully deployed all required components on a worker node. Which command should the administrator use to list the GPU Operator pods running on that node?
Easy21In a multi-node training scenario, what is the significance of the NVIDIA Collective Communications Library (NCCL) in workload management?
Medium22You are troubleshooting a node where the GPU is detected, but the application fails to utilize it. Which log source would provide the most relevant information?
Medium23An AI Operations engineer is managing a multi-node training job using NVIDIA NCCL. The logs indicate frequent 'NCCL WARN' messages related to 'net_ib_init' failures. What is the most likely cause of this issue?
Medium24In an air-gapped environment, what must an administrator do to ensure the GPU Operator correctly installs the necessary software components?
Hard25An AI operations team is troubleshooting a distributed training job on an NVIDIA DGX SuperPOD that uses NCCL for inter-GPU communication. The job intermittently hangs during the all-reduce phase. Which two actions should be taken to diagnose and resolve the issue? (Choose two.)
Hard26A production inference service running on NVIDIA T4 GPUs shows that GPU utilization is consistently below 20% while request latency is high. Profiling with Nsight Systems reveals that the model execution time is short but there are frequent gaps between kernels. Which optimization should be applied first to improve GPU utilization?
Hard27Which TWO of the following are prerequisites for installing the NVIDIA Container Toolkit on a Linux host?
Medium28An administrator is using NVIDIA Base Command Manager to provision a new GPU cluster. They need to ensure that the compute nodes are configured with the correct GPU driver and CUDA toolkit versions. Which Base Command Manager feature should they use to automate this?
Medium29A data scientist reports that a Jupyter notebook running on a GPU-enabled server is extremely slow when training a small neural network, even though nvidia-smi shows the GPU is idle. The notebook uses TensorFlow. Which is the most likely cause?
Easy30An AI operations team is deploying the NVIDIA GPU Operator on a Kubernetes cluster that uses containerd as the container runtime. The cluster nodes have NVIDIA GPUs, and the team wants to ensure that GPU workloads can request GPU resources. After installing the operator, they notice that pods requesting 'nvidia.com/gpu' remain in Pending state. Which component of the GPU Operator is most likely misconfigured or missing?
Medium31Which approach is most effective for scaling an inference workload that experiences sudden, unpredictable spikes in request volume?
Medium32A production inference service experiences intermittent latency spikes. The service is deployed on shared infrastructure. Which tool would best help an AI Ops engineer identify if GPU resource contention is the cause?
Medium33When troubleshooting an NVIDIA GPU Operator installation, which TWO locations should an administrator check to identify why the driver installation pod is failing?
Hard34Which configuration file is typically modified to enable the NVIDIA Device Plugin in a Kubernetes cluster?
Medium35When troubleshooting a NCCL collective communication timeout in a distributed training environment, which component should be the primary focus of initial investigation?
Medium36An AI operations engineer is validating a new NVIDIA-certified server before putting it into production. The job runs correctly but the team wants to confirm that the GPUs are operating at the expected clocks and not being limited by power or thermal constraints. Which command provides the most direct evidence of the current power and thermal limits and any throttling reasons?
Easy37When managing workloads on NVIDIA DGX systems, what is the primary role of the NVIDIA device plugin in the Kubernetes ecosystem?
Easy38An AI infrastructure team is preparing an air-gapped data center to install the NVIDIA GPU Operator. They have mirrored all required container images into a private registry. During installation, the Operator's pods fail with ImagePullBackOff because the components still reference images under nvcr.io. Which configuration is required to make the Operator and its managed components pull from the private registry?
Hard39A platform engineer is preparing a bare-metal Ubuntu 22.04 server with four A100 GPUs for an NVIDIA AI Enterprise deployment. The GPUs are not yet visible to the operating system tooling. Which command should the engineer run to confirm the driver loaded successfully and that all four GPUs are enumerated with their current driver version?
Easy40A platform team runs an NVIDIA GPU Operator-managed Kubernetes cluster shared by two research groups. Group A's pods request `nvidia.com/gpu: 1` and are scheduled correctly, but Group B's pods that omit any GPU resource request are also landing on GPU nodes and consuming host memory and CPU, degrading Group A's jobs. The team wants Group B's non-GPU pods to stop consuming capacity on the GPU node pool without changing Group B's manifests. Which action should the administrator take?
Medium41A financial services company is deploying NVIDIA AI Enterprise on a VMware vSphere cluster with NVIDIA A100 GPUs. The security team requires that GPU workloads be isolated at the hardware level, with separate memory and fault domains, to meet regulatory compliance. The company also wants to maximize GPU utilization by running multiple workloads concurrently. Which NVIDIA feature should be enabled to meet these requirements?
Hard42An AI operations team runs mixed training and inference workloads on a Kubernetes cluster managed with the NVIDIA GPU Operator. Inference pods frequently arrive in bursts and must start within seconds, while long-running training jobs occupy most MIG-capable A100 GPUs for days. Administrators want burst inference pods to obtain GPU capacity immediately without preempting or restarting the training jobs, and they want the cluster to reclaim those resources automatically when the burst ends. Which approach best satisfies these requirements?
Hard43Which NVIDIA tool allows you to verify that the GPU and its driver are properly installed and functioning on a Linux system?
Easy44An administrator is optimizing a large model training job to reduce checkpointing time to storage. Which strategy is most effective for minimizing the impact on training throughput?
Medium45A multi-tenant NVIDIA GPU-accelerated Kubernetes cluster utilizing NVIDIA AI Enterprise experiences intermittent out-of-memory errors on Triton Inference Server pods despite adequate node memory reservation. Which monitoring and troubleshooting action correctly isolates the root cause?
Hard46An AI administrator is tasked with monitoring GPU utilization in a multi-user cluster. Which tool provides the most granular real-time visibility into process-level GPU memory usage and compute utilization?
Medium47Which file format is commonly used to define the configuration and state for the NVIDIA GPU Operator within a Kubernetes environment?
Medium48Refer to the exhibit. An administrator notices that 'user_a' is consistently hitting resource limits despite having sufficient total system GPU memory. Based on the policy JSON, what is the cause?
Hard49An inference model running on Triton Inference Server is reporting high latency for requests. The model uses a fixed-size batching strategy. What is the most effective way to optimize throughput while maintaining latency targets?
Hard50An administrator needs to collect GPU telemetry from an NVIDIA AI Enterprise cluster and store it in a time-series database for long-term analysis. Which component should be deployed to export GPU metrics in Prometheus format?
Medium51When deploying NVIDIA AI Enterprise, why is the use of the NVIDIA NGC Catalog recommended over public container repositories?
Medium52An administrator is configuring an NVIDIA AI Enterprise cluster to run multi-tenant inference workloads on Kubernetes. The administrator must ensure that GPU resources are isolated and that tenants cannot access each other's GPU memory. Which two actions should the administrator take? (Choose two.)
Hard53An organization is deploying large-scale LLM training workloads on an NVIDIA DGX SuperPOD. The data science team reports that training jobs are frequently preempted by higher-priority batch jobs, leading to significant checkpointing overhead. Which Workload Manager configuration strategy best minimizes resource fragmentation and improves overall cluster utilization while maintaining SLA requirements?
Medium54Why is it important to use a persistent storage volume for model checkpoints in a distributed training job?
Easy55Which tool is the industry standard for monitoring and managing NVIDIA data center GPUs in a large-scale cluster deployment?
Medium56Refer to the exhibit. The cluster administrator observes near-capacity memory utilization across three GPUs. What is the most likely consequence if an additional pod is scheduled to these GPUs without memory partitioning?
Hard57An AI platform team runs GPU-accelerated inference pods on a Kubernetes cluster with the NVIDIA GPU Operator. During peak load, high-priority latency-sensitive inference pods are frequently preempted by large batch training jobs that were submitted later. The team wants the scheduler to guarantee that inference pods always win placement and preemption decisions against training pods without manually cordoning nodes. Which mechanism should the administrator configure?
Medium58In an NVIDIA-accelerated Kubernetes environment, why is it critical to configure a 'RuntimeClass' for GPU-enabled pods?
Hard59Which action is required when updating the NVIDIA driver on a node managed by the GPU Operator to ensure that running workloads are not interrupted abruptly?
Medium60An AI operations team is validating a new Kubernetes cluster before installing the NVIDIA GPU Operator with the driver managed by the Operator itself. The nodes run a supported Linux distribution with GPUs physically installed. Which two conditions must be satisfied for the Operator's driver container to build and load the kernel module successfully? (Choose two.)
Hard61Refer to the exhibit. The training job fails with a CUDA OOM error. Given the memory profile, which optimization strategy provides the most immediate relief while maintaining model performance?
Medium62What is the primary role of the NVIDIA Data Center GPU Manager (DCGM) Exporter in a cloud-native monitoring stack?
Easy63Refer to the exhibit. An administrator applies this security policy to a container runtime environment. What is the immediate effect on containerized AI applications within this scope?
Hard64An administrator must run a batch inference job that requires exactly two NVIDIA GPUs on a Kubernetes cluster managed by the NVIDIA GPU Operator. Which pod specification field should be used to request those GPUs?
Easy65An AI operations engineer is preparing a Kubernetes cluster to run GPU-accelerated workloads using the NVIDIA GPU Operator. The cluster nodes already have NVIDIA data center GPUs installed and the NVIDIA driver is pre-installed on the host. The engineer wants to use the GPU Operator to manage the container toolkit, device plugin, and monitoring components but must avoid the Operator managing or upgrading the driver. Which configuration should be applied to the GPU Operator deployment?
Medium66A platform engineer is preparing an NVIDIA-accelerated Kubernetes cluster for a new team that will submit PyTorch training jobs. The team wants jobs to request GPUs without hardcoding device indices. Which Kubernetes resource should the engineer ensure is installed and healthy so pods can request nvidia.com/gpu resources?
Easy67Which command is used to verify that the NVIDIA GPU Operator has successfully installed the necessary components on a Kubernetes node?
Easy68An administrator notices that a specific containerized training job reports high 'GPU Duty Cycle' but low 'Memory Bandwidth Utilization'. What does this pattern indicate about the workload?
Hard69A Kubernetes cluster administrator is installing the NVIDIA GPU Operator and wants to ensure that GPU workloads are scheduled only on nodes with healthy GPUs. The administrator plans to use the operator's built-in health checks. Which component is responsible for monitoring GPU health and marking nodes as unschedulable when a GPU fails?
Medium70A data engineering team is deploying a distributed data processing workload. Which THREE metrics are most important to monitor in the Workload Manager to ensure optimal GPU throughput and identify potential bottlenecks?
Hard71A data engineering team runs nightly batch inference on a Kubernetes cluster with NVIDIA GPUs. Jobs sometimes fail because two pods are scheduled onto the same physical GPU and one exhausts framebuffer memory. The team wants each pod to receive an isolated slice of a GPU with dedicated memory. Which NVIDIA feature should they enable?
Easy72What is the primary benefit of using NVIDIA GPU Operator in a Kubernetes cluster for workload management?
Medium73Which of the following describes the purpose of the NVIDIA GPU Operator's 'Driver Container'?
Medium74An administrator is responsible for maintaining a fleet of NVIDIA-certified servers running AI workloads. They need to quickly identify which servers have GPUs that are overheating and may throttle performance. Which NVIDIA tool should the administrator use to monitor GPU temperature across the fleet in real time?
Easy75When installing NVIDIA drivers via a package manager, what is the importance of the 'dkms' package?
Easy76An administrator is optimizing a cluster for AI model training using NVIDIA Base Command. Which TWO tasks are critical for ensuring consistent performance across the training nodes?
Medium77An AI operations engineer is tuning a real-time inference service on NVIDIA A100 GPUs. Profiling with Nsight Systems shows that the GPU is idle for long periods while waiting for input data, and that host-to-device memory copies are frequent and small. The service uses a fixed batch size of 1 and a custom data loader. Which two changes are most likely to improve GPU utilization and reduce inference latency? (Choose two.)
Hard78An AI engineer observes that a model training job on an NVIDIA DGX system is underutilizing the GPU. The monitoring logs show high CPU wait times and low GPU duty cycles. Which action should the engineer take first to resolve the bottleneck?
Medium79An administrator is optimizing multi-GPU utilization. Which TWO of the following configurations allow multiple containers to share a single physical GPU on a supported NVIDIA architecture?
Medium80An administrator supports a multi-tenant cluster where several teams share GPUs. Leadership requires that each team's batch jobs receive a fair share of GPU time and that one team cannot monopolize devices by submitting thousands of low-priority pods. Jobs are submitted through a Kubernetes-native batch scheduler that supports queueing. Which approach best enforces fair-share GPU allocation across teams?
Hard81When installing the NVIDIA GPU Operator, which namespace is typically used to ensure proper isolation and role-based access control?
Easy82When deploying large-scale distributed training jobs, why is it recommended to use the NVIDIA Network Operator in conjunction with the GPU Operator?
Medium83Which of the following is the primary indicator of PCIe bus saturation when profiling a training job on an NVIDIA DGX system?
Hard84Which action must be performed after updating the NVIDIA driver on a Linux host to ensure that all active GPU containers recognize the new driver version?
Medium85Refer to the exhibit. What is the most effective way to resolve this specific throttling condition?
Medium86A data scientist reports that a PyTorch training job on an NVIDIA V100 GPU is running slower than expected. The job uses a data loader with num_workers=4. Monitoring shows GPU utilization is around 50%, and CPU usage is high. Which action should an AI operations engineer recommend to improve GPU utilization?
Medium87An AI operations team is troubleshooting a distributed training job on a cluster of NVIDIA DGX A100 systems connected via InfiniBand. The job runs but achieves only 40% of expected scaling efficiency. The team suspects communication bottlenecks. Which two actions should they take to confirm and address the issue? (Choose two.)
Hard88A platform engineer manages an NVIDIA-accelerated Kubernetes cluster running the NVIDIA GPU Operator on nodes with A100 GPUs. Several data-science teams submit training jobs, and the engineer must ensure each team's pods receive a full physical GPU exclusively, with no two pods sharing the same device. Which scheduling configuration should the engineer apply to the pod specification to guarantee exclusive whole-GPU allocation?
Easy89An AI researcher is running a large-scale training job on an NVIDIA DGX system using Kubernetes. They observe that GPU utilization is consistently low despite high CPU load. Which workload management configuration is most likely to resolve this bottleneck by optimizing data pipeline throughput?
Medium90Refer to the exhibit. The training job fails with an OOM error. Which optimization strategy will most effectively resolve this while maintaining model convergence?
Medium91An administrator is configuring a Kubernetes cluster where some nodes have A100 GPUs and others have H100 GPUs. A training job requires specific GPU memory capacity and CUDA compute capability. Which mechanism should the administrator use to ensure the job is only scheduled onto nodes with the correct GPU model?
Hard92A platform team is preparing a Kubernetes cluster for AI workloads and wants the GPU device plugin, driver containers, and monitoring components deployed and kept in sync automatically on every GPU node. Which component should be installed to achieve this?
Easy93An AI operations engineer manages a shared Kubernetes cluster where several teams submit GPU jobs. The engineer must prevent any single namespace from consuming all GPU capacity and must also ensure that jobs from one team cannot starve others during peak periods. Which combination of Kubernetes and NVIDIA GPU Operator features should the engineer implement?
Medium94Which TWO of the following actions are recommended for optimizing NVIDIA GPU utilization during a high-concurrency inference deployment?
Medium95A healthcare company is deploying NVIDIA AI Enterprise on a Kubernetes cluster to run medical imaging AI models. The cluster administrator needs to verify that the NVIDIA GPU Operator is installed and functioning correctly. Which command should the administrator use to check the status of the GPU Operator pods?
Easy96An AI operations engineer is investigating a training job on an NVIDIA DGX system that intermittently fails with 'uncorrectable ECC error' on a GPU. The job is using NCCL for multi-GPU communication. The engineer needs to identify the appropriate immediate actions to diagnose and mitigate the issue. (Choose two.)
Hard97A platform engineer must validate a new NVIDIA GPU Operator deployment on a Kubernetes cluster before handing it to data scientists. Which two checks confirm that the Operator has correctly exposed GPU resources to the cluster scheduler? (Choose two.)
Hard98Which component of the NVIDIA GPU Operator is responsible for monitoring GPU health and reporting telemetry data to the Kubernetes control plane?
Easy99Which TWO strategies should an administrator implement to ensure fair resource scheduling in a multi-tenant NVIDIA cluster using Kubernetes and the NVIDIA device plugin?
Hard100A team runs multi-node training with NCCL over InfiniBand on a cluster of DGX systems. Jobs scale well to four nodes but throughput drops sharply at eight nodes, and `nccl-tests` all-reduce bandwidth falls well below line rate at that size. The fabric uses a fat-tree topology with adaptive routing enabled. Which investigation is most likely to reveal the cause?
Hard101When deploying the NVIDIA GPU Operator in a restricted-access environment (air-gapped), which THREE requirements must be addressed to ensure a successful installation?
Hard102An administrator runs a Kubernetes cluster with the NVIDIA GPU Operator. A data science team wants to run several small inference containers that each use only a fraction of a GPU's compute and memory, but the cluster currently assigns whole GPUs per pod. Which approach allows multiple containers to share a single physical GPU with memory isolation?
Hard103A cluster administrator notices that GPU utilization is low despite high queue volume. After analyzing the logs, they identify that many pods are failing because they cannot access the necessary CUDA libraries. What is the most likely cause, and which component should be verified?
Medium104An administrator manages an NVIDIA AI Enterprise cluster running multiple Kubernetes nodes, each with several A100 GPUs. After upgrading the NVIDIA GPU Operator to a newer version, the administrator notices that pods requesting GPUs remain in a Pending state, and the node's allocatable GPU count is reported as zero. Which command should the administrator run first to diagnose the issue?
Medium105Which NVIDIA technology enables the partitioning of a single physical GPU into multiple independent instances for use by different virtual machines or containers?
Medium106An AI engineer needs to monitor GPU utilization across a large cluster of nodes in real-time. Which NVIDIA tool is the most appropriate for this high-level observability task?
Medium107Which NVIDIA software component is responsible for providing the necessary CUDA libraries to containerized applications?
Easy108An AI engineer is deploying a large language model on an NVIDIA DGX system. The deployment fails with an error indicating an insufficient NVIDIA driver version for the required CUDA toolkit. Which action should the engineer take to resolve the dependency mismatch?
Medium109An AI operations engineer notices that a real-time inference service on an NVIDIA T4 GPU has highly variable latency, with occasional spikes to over 100 ms. The service uses TensorRT and runs in a Docker container. Which action should the engineer take to reduce latency variability?
Easy110A site reliability engineer is troubleshooting a DGX A100 node that intermittently drops out of the cluster during large NCCL all-reduce jobs. `nvidia-smi` shows all eight A100 GPUs healthy, but DCGM reports XID errors 74 and 79 on one GPU during the failures. The engineer needs to determine the most likely cause and the correct administrative action. Which combination best describes the cause and the appropriate first step?
Hard111An MLOps engineer manages a Kubernetes cluster where the NVIDIA GPU Operator runs the MIG manager. Several inference pods must each receive an isolated, fixed slice of a single A100, and the team wants the slices to survive node reboots without manual reconfiguration. Which combination of settings should the engineer apply?
Hard112Which THREE of the following are benefits of using Multi-Instance GPU (MIG) technology in a Kubernetes environment?
Hard113An AI operations team manages a shared Kubernetes cluster where a nightly batch training workload requests nvidia.com/gpu resources and occasionally consumes all GPU memory on a node, causing a co-located interactive notebook pod to fail with CUDA out-of-memory errors. The team wants the interactive notebook to be isolated from the batch workload's memory usage without adding new hardware. Which action best achieves this on supported data center GPUs?
Hard114An AI operations team is deploying NVIDIA AI Enterprise on a bare-metal Kubernetes cluster with DGX A100 systems. They need to enable GPUDirect Storage to accelerate data loading from a local NVMe array. Which component must be installed and configured on the DGX nodes to support GPUDirect Storage?
Hard115Which mechanism does the NVIDIA GPU Operator use to ensure that the driver installed on a worker node matches the specific architecture of the installed GPU hardware?
Medium116A platform team runs an on-premises Kubernetes cluster for AI inference. Several teams submit pods that request the same GPU device, and the scheduler places more pods onto a node than there are available GPUs, causing OOM errors on the device. The administrator wants the Kubernetes scheduler itself to prevent overcommitting GPUs without any custom admission controller. Which action should the administrator take?
Easy117An administrator is deploying NVIDIA AI Enterprise on a bare-metal cluster. Which component is mandatory to provide the necessary abstraction layer for containerized GPU resources?
Medium118A data scientist reports that a Jupyter notebook running on a NVIDIA GPU server is extremely slow when executing a deep learning model, even though `nvidia-smi` shows the GPU is idle. The notebook uses TensorFlow. Which action should be taken first to diagnose the issue?
Easy119A system administrator is installing the NVIDIA Container Toolkit on a standalone Ubuntu server to run GPU-accelerated containers. After installation, they want to verify that the toolkit is correctly configured. Which command should they run to test GPU access from a container?
Easy120An administrator is troubleshooting a performance degradation in a multi-node NVIDIA NCCL-based training job. The job spans four DGX nodes connected via InfiniBand. The administrator suspects that NCCL is not using the optimal network path. Which action should the administrator take to verify and enforce the use of GPUDirect RDMA for inter-node communication?
Hard121During the deployment of an AI model training workload on a multi-node cluster, the administrator notices that inter-node communication is significantly slower than expected. Which deployment aspect should be investigated first?
Medium122A platform engineer is preparing a bare-metal Kubernetes cluster to run GPU-accelerated AI workloads using the NVIDIA GPU Operator. The cluster nodes have NVIDIA Ampere GPUs and run Ubuntu 22.04 with containerd as the container runtime. The engineer wants to avoid installing any NVIDIA drivers or CUDA components directly on the host. Which GPU Operator configuration should be used to achieve this?
Medium123An administrator is configuring a new cluster and wants to ensure that telemetry data from GPUs is collected in a centralized manner. Which tool is best suited for this requirement?
Medium124An operations engineer is troubleshooting a distributed training job that uses NVIDIA Magnum IO GPUDirect Storage to read training data directly from a local NVMe SSD into GPU memory. The job reports lower than expected I/O bandwidth. `nvidia-smi` shows normal GPU utilization, and the NVMe drive's throughput is well below its peak. Which factor is most likely limiting GPUDirect Storage performance in this scenario?
Medium125A financial services company is deploying NVIDIA AI Enterprise in an air-gapped data center. They need to install the NVIDIA GPU Operator on their Kubernetes cluster without internet access. Which additional step must they take to ensure a successful installation?
Medium126A research group submits a distributed PyTorch training job spanning eight GPUs across two nodes. The job completes but produces a model with accuracy far below the single-node baseline, and logs show that several ranks started training before their peers had initialized the process group. The administrator must ensure that all ranks are launched together and that a failed rank terminates the whole job. Which combination of Kubernetes mechanisms should be used?
Hard127An administrator is tuning a Kubernetes cluster that runs many small inference pods on NVIDIA GPUs. Utilization is low because each pod reserves a full GPU while using only a fraction of its memory and compute. The administrator wants to share GPUs across pods while preserving memory-level isolation between processes. Which TWO configurations achieve this? (Choose two.)
Hard128A team is diagnosing a training job that intermittently stalls for several seconds at the start of each epoch. The job uses a distributed data loader and an NVIDIA DGX system with local NVMe. Monitoring shows GPU utilization dropping to near zero during the stalls while host CPU utilization spikes. Which two actions should the AI operations engineer take to identify and mitigate the stall? (Choose two.)
Hard129An administrator is deploying NVIDIA AI Enterprise on a Kubernetes cluster and must decide how GPU workloads should request accelerators. The environment has a mix of full-GPU training jobs and inference services that share a single A100. Which approach correctly allows a pod to consume a specific MIG-backed slice rather than the whole device?
Hard130Which TWO methods are effective for enforcing GPU resource isolation in a multi-tenant NVIDIA Kubernetes environment?
Hard131An AI operations team is deploying NVIDIA Base Command Manager to manage a cluster of DGX nodes. They want to ensure that only authorized users can submit jobs and that all job submissions are audited. Which combination of Base Command Manager features should the administrator configure to meet these requirements?
Medium132Which component is responsible for exposing the GPU as a schedulable resource in a Kubernetes cluster?
Medium133An AI operations team needs to monitor GPU health and utilization across a fleet of DGX nodes from a single dashboard. They want per-GPU metrics such as power, temperature, utilization, and ECC errors, and they want to retain historical data for capacity planning. Which NVIDIA tool is purpose-built to collect and expose these GPU telemetry metrics for centralized monitoring?
Easy134A research lab is deploying NVIDIA AI Enterprise on an air-gapped Kubernetes cluster. The cluster has no internet access, and all software must be installed from a local registry. The administrator plans to use the NVIDIA GPU Operator. Which two actions must be performed to ensure a successful deployment in this environment? (Choose two.)
Hard135An administrator notices that GPU utilization on a training cluster hovers around 25 percent even though many jobs are queued. Investigation shows that each job requests a full GPU, but the models are small and alternate between short data-loading phases and brief compute bursts. The administrator wants to increase effective GPU utilization without changing model code. Which action should be taken first?
Medium136An administrator supports a shared inference cluster where a single A100 GPU must serve several small models concurrently. They configure the NVIDIA device plugin with a time-slicing configuration that advertises multiple replicas of the same physical device. After deployment, users report that one noisy model starves the others and latency spikes unpredictably. Which statement best explains the observed behaviour?
Medium137A cluster runs mixed workloads: latency-sensitive inference services and best-effort batch jobs. Administrators observe that batch jobs occasionally occupy all GPUs, causing inference requests to queue and breach service level objectives. They want inference pods to be admitted immediately while allowing batch work to use remaining capacity and be preempted when needed. Which approach should they implement?
Hard138An administrator is preparing a multi-node NVIDIA DGX H100 cluster for a distributed training job using NVIDIA Base Command. The cluster nodes have InfiniBand adapters, but the job's inter-node throughput is far below expectations. The administrator runs `ibstat` and sees that the ports are in the INIT state rather than ACTIVE. Which action should the administrator take first?
Medium139An administrator is preparing a Kubernetes cluster for AI workloads and needs to ensure that the NVIDIA GPU Operator can be installed. The cluster nodes have NVIDIA GPUs, and the administrator wants to verify that the nodes are ready. Which command should the administrator run to check if the NVIDIA driver is already loaded on a node?
Easy140An administrator is configuring a multi-tenant NVIDIA AI Enterprise environment. Which mechanism is most effective for ensuring hardware-level isolation between concurrent training jobs on a single A100 GPU?
Medium141An administrator supports a shared Kubernetes cluster running NVIDIA GPU Operator. Data scientists report that their inference pods remain in Pending state, yet the GPU Operator pods and node feature discovery pods are healthy, and the GPU nodes show no hardware alarms. The administrator confirms that the cluster has a mixture of MIG-capable A100 nodes and non-MIG T4 nodes. Which immediate administrative action is most appropriate to diagnose the scheduling failure?
Medium142A system administrator is installing NVIDIA AI Enterprise on a Kubernetes cluster that will use Multi-Instance GPU (MIG) on A100 GPUs. The administrator wants to ensure that MIG instances are properly exposed as schedulable resources. Which action must be taken after enabling MIG mode on the GPUs?
Medium143An administrator is deploying the NVIDIA GPU Operator on a Kubernetes cluster using Helm. The cluster nodes have NVIDIA GPUs and the administrator wants to ensure that the GPU Operator can automatically label nodes with GPU properties and install the device plugin. Which prerequisite must be met on the cluster nodes before installing the GPU Operator?
Easy144An administrator is installing the NVIDIA GPU Operator on a Kubernetes cluster. They want to verify that the GPU Operator's components are running correctly after installation. Which command should they use to check the status of the GPU Operator pods?
Easy145A cloud operations engineer is deploying the NVIDIA GPU Operator on a managed Kubernetes service where the worker nodes already have the NVIDIA data center driver installed by the cloud provider. The team wants the Operator to manage only the device plugin, container toolkit, and monitoring components. Which Helm value should the engineer set during installation?
Easy146An AI operations engineer is optimizing a real-time inference pipeline on an NVIDIA T4 GPU. The pipeline uses TensorRT and receives requests with variable input sizes. Profiling shows that the engine recompiles for each new input shape, causing latency spikes. Which optimization should the engineer apply to eliminate recompilation while maintaining acceptable accuracy?
Medium147During deployment, an AI model experiences high variance in latency during inference. The system uses a fixed instance count. What is the most likely cause for this performance jitter?
Medium148Refer to the exhibit. An engineer is troubleshooting inconsistent training performance across two GPUs in a single DGX node. Why is one GPU reporting a lower clock speed despite being in P0 state?
Hard149When running multi-instance GPU (MIG) workloads, what is the main advantage of assigning specific MIG profiles to different Kubernetes namespaces?
Hard150An administrator is responsible for a large NVIDIA DGX SuperPOD used for multi-node training. They need to ensure that GPU telemetry and health metrics are collected centrally and can trigger alerts when GPUs exceed temperature thresholds. Which component of NVIDIA Base Command Manager (BCM) should they configure to achieve this?
Hard151A monitoring system reports that a DGX node's GPUs are running at reduced clocks during a long training job, and `nvidia-smi -q -d PERFORMANCE` shows the throttle reason as 'SW Power Cap'. The job's power draw is at the configured limit. Which action is MOST appropriate to restore higher clocks?
Easy152An ML platform team runs an NVIDIA GPU Operator-managed cluster and wants to allow multiple pods to share a single A100 GPU so that small inference services can co-reside without each consuming a whole device. The team needs a time-slicing configuration that applies to all GPU nodes in the cluster. Which approach should the administrator take?
Medium153An organization is migrating their on-premises AI training to a hybrid cloud environment. Which component is most important to maintain consistent workload management across both the on-premises DGX systems and cloud-based GPU nodes?
Medium154A production inference service on NVIDIA GPUs reports that p99 latency spikes every few minutes while p50 remains stable. Metrics show GPU memory utilization near the limit and periodic `cudaMalloc` calls in the application logs. The model and batch size are fixed. Which change is MOST likely to eliminate the latency spikes?
Hard155A researcher submits a distributed training job that spans four pods, each needing one GPU, and the pods must start together or not at all. The administrator wants Kubernetes to schedule all four pods only when four GPUs are simultaneously available. Which workload management construct should be used?
Medium156A research team submits a multi-node training job using a `Job` with eight pods, each requesting one GPU. The cluster has eight GPU nodes, each with one A100. The administrator observes that all eight pods are spread one per node and the job runs, but throughput is far below expectations and NCCL logs show repeated fallback from GPUDirect RDMA to socket transport. Which action most directly addresses the root cause?
Hard157A financial services company is deploying NVIDIA AI Enterprise on a Kubernetes cluster with strict security policies. They need to ensure that GPU workloads are isolated and that the NVIDIA GPU Operator components are deployed with least privilege. Which feature of the NVIDIA GPU Operator allows administrators to define granular permissions for its components?
Hard158An administrator manages an NVIDIA AI Enterprise deployment on Kubernetes and needs to ensure that GPU telemetry is exported to an existing Prometheus instance. The administrator deploys the NVIDIA DCGM Exporter but sees no GPU metrics in Prometheus. Which configuration should the administrator verify first?
Hard159An administrator wants to prevent unauthorized users from accessing sensitive model weights stored in GPU memory. Which security feature should be implemented to ensure hardware-level isolation of the memory space?
Medium160An AI operations engineer is troubleshooting a multi-node NCCL training job on an NVIDIA DGX SuperPOD. The job runs but scales poorly: inter-node bandwidth is roughly half of the expected 200 Gb/s per GPU, while intra-node NVLink traffic is at full rate. Running `nvidia-smi topo -m` shows that GPUs in each node are connected to the NICs through the PCIe switch, but the job sets `NCCL_SOCKET_IFNAME` to the management interface. Which action is the most appropriate to resolve the bottleneck?
Medium161An AI engineer observes that a training job on an NVIDIA DGX H100 system is experiencing significant performance degradation. The GPU utilization is high, but the throughput remains low. Which tool should be used first to identify if the bottleneck is related to data loading or PCIe bandwidth saturation?
Medium162An AI operations team is using NVIDIA DCGM (Data Center GPU Manager) to monitor a cluster of A100 GPUs. They want to set up proactive health checks to detect and mitigate GPU issues before they cause job failures. Which two DCGM features should they configure? (Choose two.)
Medium163An administrator is responsible for an NVIDIA AI Enterprise deployment on Kubernetes. The security team requires that all GPU-accelerated pods run with the least privilege necessary and that GPU device nodes are not exposed to pods that do not request them. Which combination of configurations should the administrator implement to meet these requirements?
Hard164Which mechanism does the NVIDIA Device Plugin use to communicate GPU availability to the Kubernetes Kubelet?
Medium165An administrator is deploying NVIDIA AI Enterprise on a bare-metal cluster. The workload requires full GPU isolation with minimal latency. Which configuration should the administrator select to achieve this goal?
Medium166A data scientist reports that a Jupyter notebook running on a DGX station cannot allocate GPU memory, even though other users' jobs are running fine. The notebook kernel was started before a system administrator updated the NVIDIA driver and rebooted the node. Which action should the data scientist take to resolve the issue?
Easy167Refer to the exhibit. An administrator notices poor performance in an AI training job. What is the most likely cause based on the CLI output?
Medium168An operations team is troubleshooting a multi-node NCCL training job on an NVIDIA DGX SuperPOD that stalls at initialization and never begins gradient exchange. Running `nccl-tests` with `all_reduce_perf` on the same nodes fails identically, but single-node `all_reduce_perf` succeeds. Which action should the team take first to isolate the fault?
Medium169An administrator is configuring NVIDIA GPUDirect Storage (GDS) on a cluster to accelerate data loading for AI training jobs. The cluster uses Mellanox InfiniBand adapters and NVMe storage. Which two actions are required to enable GDS and ensure optimal performance? (Choose two.)
Hard170An AI operations engineer is troubleshooting a multi-GPU training job that fails intermittently with a NCCL error: 'unhandled system error'. The job runs on a DGX-1 with eight V100 GPUs connected via NVLink. Which step should the engineer take first to resolve the issue?
Hard171When designing a workload management strategy for multi-tenant AI training, what is the most effective way to ensure isolation between different tenants using the same physical GPU nodes?
Medium172An operations engineer is investigating a sudden drop in throughput for a multi-GPU training job on an NVIDIA DGX A100. The job uses PyTorch with DDP. Logs show that one GPU is consistently at 100% utilization while others are below 50%. Which tool and approach should be used to identify the bottleneck?
Hard173During an NVIDIA AI Enterprise deployment, you are asked to configure the 'NVIDIA Container Toolkit'. What is its primary function?
Easy174Refer to the exhibit. An AI engineer observes that a model training job is running slower than expected. Based on the output, what is the primary cause of the performance degradation?
Medium175An AI operations team is troubleshooting a training job that crashes with a segmentation fault after several hours. The job uses multiple GPUs and NCCL for communication. System logs show no errors, but dmesg reveals repeated 'NVRM: Xid' errors. Which action should be taken first to diagnose the issue?
Easy176A machine learning engineer is optimizing a recommendation model for inference on an NVIDIA T4 GPU. The model uses dynamic input shapes, and profiling shows that kernel launch overhead is a significant contributor to latency. Which optimization technique should be applied to reduce this overhead?
Medium177When deploying NVIDIA containers using the NVIDIA Container Toolkit, what is the primary function of the 'nvidia-container-runtime'?
Easy178When debugging a workload that consistently crashes with 'Out of Memory' (OOM) errors despite sufficient GPU VRAM, what is the most likely cause related to workload management?
Hard179An AI operations team is installing the NVIDIA GPU Operator on a Kubernetes cluster that uses a custom containerd configuration. They need to ensure that the GPU Operator can properly manage the container runtime. Which action should they take before installing the GPU Operator?
Medium180A production inference service on an NVIDIA A100 GPU experiences a gradual increase in latency over several hours, eventually requiring a pod restart. GPU memory utilization climbs steadily, but the model and batch size are fixed. Which action should an AI operations engineer take first to diagnose the root cause?
Medium181An administrator manages a cluster where inference services and batch training share the same GPU nodes. During business hours, inference pods must be scheduled promptly, while training jobs can wait. The administrator wants preemption so that a pending high-priority inference pod can evict a lower-priority training pod when no GPU is free, with evicted training resuming later. Which configuration achieves this?
Hard182An AI operations engineer is troubleshooting a model inference service deployed with NVIDIA Triton Inference Server on a GPU. The service occasionally returns incorrect predictions, and the engineer suspects that the input data is not being preprocessed correctly. The model expects input tensors in FP32 format, but the client is sending FP16 data. Which action should the engineer take to resolve the issue?
Easy183An administrator is tuning a Kubernetes cluster that runs GPU Operator. Users report that GPU jobs are sometimes scheduled onto nodes whose drivers are older than the CUDA version the container needs, causing runtime failures. The administrator wants to prevent incompatible placements before pods are bound. (Choose two.)
Medium184When configuring the NVIDIA Device Plugin for Kubernetes, what is the purpose of the 'time-slicing' configuration?
Medium185A company runs multiple AI workloads on a shared Kubernetes cluster with NVIDIA GPUs. The administrator needs to enforce that only pods with a specific label can consume GPU resources, while other pods are denied. Which Kubernetes admission control mechanism should be used to implement this policy?
Medium186During a training job, the system reports "NCCL WARN" regarding a slow network path. What is the most likely culprit for this performance bottleneck in a multi-node InfiniBand environment?
Medium187A platform team runs mixed training and inference workloads on a Kubernetes cluster with the NVIDIA GPU Operator. Inference pods are latency-sensitive and must not be preempted, while training pods can be interrupted and restarted. The team wants training jobs to yield GPUs to inference jobs when capacity is scarce, without manual intervention. Which Kubernetes mechanism should the team configure to achieve this behavior?
Medium188An administrator is using NVIDIA Base Command Manager to manage a cluster with a mix of GPU and CPU nodes. They need to ensure that a newly added GPU node is correctly recognized and that jobs can be scheduled on it. Which TWO actions must be performed to integrate the new node into the Base Command Manager cluster? (Choose two.)
Medium189During an air-gapped installation of the NVIDIA GPU Operator, the administrator must make all required images available to the cluster. Which component is responsible for pulling the Operator's operand images from the private registry?
Medium190Which component in the NVIDIA AI ecosystem is responsible for monitoring and reporting GPU telemetry data, such as power usage, temperature, and utilization, to Prometheus?
Easy191A distributed training job on a multi-GPU node is exhibiting poor scaling efficiency: each GPU shows high utilization, but overall throughput increases by only 15% when doubling the number of GPUs. The job uses NCCL for communication. Which diagnostic step is most appropriate to identify the bottleneck?
Medium192An AI engineer is optimizing a real-time inference pipeline on an NVIDIA A100 GPU. The model uses dynamic input shapes, and profiling shows that the GPU spends significant time on memory copies between host and device. Which optimization should be implemented to reduce this overhead?
Medium193An AI platform engineer is preparing a fleet of NVIDIA DGX H100 systems for production workloads using the NVIDIA Base Command Manager (BCM). During initial bare-metal provisioning via PXE boot, the provisioning server successfully hands out IP addresses, but nodes consistently fail during the OS image deployment phase, throwing a kernel panic related to missing storage drivers. Which deployment step must be verified or corrected to ensure successful hardware-specific image deployment?
Medium194Which THREE factors must be considered when sizing a persistent storage solution for multi-node distributed training checkpoints?
Hard195An administrator is managing an NVIDIA DGX SuperPOD used for large-scale AI training. The cluster uses a Slurm workload manager. The administrator needs to ensure that jobs are scheduled only on nodes with healthy GPUs and that failed GPUs are automatically drained from the pool. Which integration should be configured to achieve this?
Medium196An AI operations engineer is investigating an inference service running on NVIDIA Triton Inference Server. Clients report sporadic 500 errors under peak load. The server logs show occasional 'Failed to allocate memory' messages, and `nvidia-smi` shows VRAM nearly full. The service uses dynamic batching with a maximum batch size of 64 and multiple model instances per GPU. Which change is the most appropriate first step to stabilize the service?
Medium197A production inference service running on NVIDIA GPUs exhibits periodic latency spikes every few minutes, correlating with CPU-side stalls and low GPU utilization during those intervals. Profiling with Nsight Systems shows large gaps between kernel launches and frequent cudaMalloc/cudaFree calls. Which action best addresses the root cause?
Medium198Which component is strictly necessary for managing NVIDIA AI Enterprise licensing across a distributed cluster of nodes?
Easy199Refer to the exhibit. The administrator has deployed the GPU Operator, but the node does not show GPU resources. What is the most likely cause?
Medium200A cloud operations team is deploying the NVIDIA GPU Operator in an environment where the Kubernetes control plane cannot reach the public internet, but worker nodes can access an internal HTTP registry that mirrors required images. The team wants to avoid manual image pulls on each node. Which two configurations should they implement to enable a successful air-gapped installation? (Choose two.)
Hard201When monitoring GPU health in an enterprise cluster, which command provides the most comprehensive snapshot of real-time power, temperature, and memory utilization?
Easy202A team is profiling a distributed training job using NVIDIA NCCL for inter-GPU communication on a DGX A100 system. They observe that all-reduce operations are taking longer than expected, and the NCCL debug logs show frequent 'NVLS' (NVLink SHARP) errors. Which action should be taken to resolve the issue?
Hard203An administrator manages a Kubernetes cluster where a training job repeatedly fails with an OutOfMemory error on the GPU even though the pod requests one nvidia.com/gpu. DCGM metrics show that another pod on the same node is consuming GPU memory concurrently. GPU sharing via time-slicing is enabled cluster-wide. Which action should the administrator take to prevent this cross-pod interference while preserving the ability to share GPUs among trusted inference workloads?
Hard204An administrator is configuring NVIDIA Base Command Manager to manage a cluster of DGX nodes. They want to ensure that when a node's GPU temperature exceeds a defined threshold, the node is automatically drained and an alert is sent to the operations team. Which combination of Base Command Manager features should the administrator configure to achieve this?
Hard205An administrator observes that despite the GPU Operator being installed, the pods cannot access the GPU. What is the most likely cause if the NVIDIA container runtime is properly configured?
Hard206An AI operations engineer is investigating a training job that exhibits poor scaling efficiency when moving from 8 to 32 GPUs on a DGX SuperPOD. Profiling indicates that the communication time in NCCL all-reduce operations is disproportionately high. Which two actions should be taken to improve scaling efficiency? (Choose two.)
Hard207An AI ops engineer notices that a specific training workload is experiencing high 'wait' times for GPU resources despite the cluster having available idle GPUs. What is the most likely cause?
Hard208An AI operations engineer manages a Kubernetes cluster running the NVIDIA GPU Operator. A team wants its long-running inference deployment to be automatically rescheduled if the GPU on a node develops an uncorrectable error that the device plugin or health checks detect. The team also wants the node to stop accepting new GPU pods until the issue is resolved. Which combination of behaviors should the engineer rely on to meet these requirements?
Medium209What is the primary role of a private container registry in an NVIDIA AI Enterprise deployment?
Medium210An AI operations engineer manages a Kubernetes cluster where the NVIDIA GPU Operator's device plugin exposes GPUs as schedulable resources. A data science team submits a batch inference job that requests one GPU but does not specify a node selector or tolerations. The job stays in Pending while other GPU nodes remain idle because they carry a taint the GPU Operator applied to reserve them for a specific workload class. Which approach is the most appropriate for the engineer to make the job schedulable without disrupting the reserved nodes?
Easy211A research team is submitting many short-lived experiment jobs to an NVIDIA-accelerated Kubernetes cluster. The operations team wants to reduce GPU idle time and improve overall utilization without modifying the training code. Which TWO approaches should the operations team implement? (Choose two.)
Medium212If a training job on a multi-node cluster shows a significant performance drop during checkpointing, what is the most likely bottleneck?
Medium213Which component of the NVIDIA AI Enterprise stack is primarily responsible for ensuring the long-term stability and compatibility of drivers and libraries across heterogeneous hardware configurations?
Easy214An AI researcher is running a training job using mixed precision (FP16/BF16). The loss function is diverging unexpectedly. What is the most likely culprit?
Medium215Refer to the exhibit. A cluster administrator notices that GPU jobs with this PriorityClass are failing to start even when empty GPUs are available. What is the most likely cause?
Hard216An operations team observes that a distributed training job using NVIDIA Collective Communications Library (NCCL) across eight nodes occasionally hangs during the all-reduce phase. Logs show no errors, and the hang resolves only after a node is manually restarted. Which action is MOST appropriate to diagnose the intermittent hang?
Hard217An AI operations team runs long-running training jobs on a Kubernetes cluster with NVIDIA GPU Operator. They observe that after a node is rebooted for maintenance, some pods resume but report CUDA 'unknown error' and the device plugin shows unhealthy GPUs. Which configuration should the administrator review to ensure the driver and device plugin recover cleanly after reboot?
Hard218An AI operations team is using NVIDIA DCGM to monitor a cluster of GPUs. They want to set up alerts for when GPUs are running at high temperatures for extended periods. Which DCGM feature should they use?
Medium219Which of the following is the most appropriate workload management technique for a bursty AI inference workload that requires low latency but does not need full GPU power for every request?
Medium220An administrator is deploying NVIDIA AI Enterprise on a Kubernetes cluster and wants to verify that the GPU Operator has successfully installed all required components. Which command should the administrator use to check the status of the GPU Operator pods?
Easy221An AI operations engineer notices that a training job on an NVIDIA A100 GPU is running slower than expected. Running nvidia-smi shows that the GPU is in 'P0' state but the 'SM Clock' is significantly lower than the maximum boost clock. The job is not memory-bound. Which action should the engineer take first to diagnose the issue?
Easy222An AI operations engineer is preparing a DGX H100 system for a multi-node training workload. The engineer runs `nvidia-smi topo -m` and notices that GPU4 and GPU5 report a connection type of SYS, while all other GPU pairs show NV18. What is the most likely cause of this topology anomaly?
Medium223A team runs a multi-node NCCL all-reduce training job on four DGX H100 nodes connected by an InfiniBand fabric. Scaling efficiency is poor: throughput barely improves beyond two nodes, and `nvidia-smi` shows NIC transmit counters on each GPU's assigned HCA are far below the PCIe link capacity while GPU compute utilization sits at ~55%. The fabric manager logs report all links as up with no symbol errors. Which action should the administrator take first to diagnose the interconnect bottleneck?
Hard224Which NVIDIA technology enables a GPU to be shared among multiple virtual machines or containers while maintaining strict hardware isolation?
Easy225An administrator is troubleshooting a multi-node NVIDIA GPU training job that intermittently hangs during the all-reduce phase. The job uses NCCL over InfiniBand. Logs show that some ranks time out while others complete. The administrator suspects a network fabric issue. Which action should the administrator take first to isolate whether the problem is in the InfiniBand fabric or in the NCCL configuration?
Medium226A platform team runs a Kubernetes cluster where the NVIDIA GPU Operator is installed and time-slicing is configured with a ConfigMap that advertises four replicas per physical GPU. A data scientist submits a PyTorch training pod requesting nvidia.com/gpu: 1. The pod stays Pending indefinitely, and the scheduler event reads 'Insufficient nvidia.com/gpu'. The node's GPUs are otherwise idle and healthy. Which action most directly resolves the pending state?
Medium227A system administrator is troubleshooting a 'CUDA error: invalid device ordinal' when launching a job on a multi-GPU system. What is the most likely cause?
Hard228Which THREE components are required for a container to successfully leverage NVIDIA GPUs on a Kubernetes cluster?
Hard229A company is deploying NVIDIA AI Enterprise on a vSphere cluster with multiple ESXi hosts, each containing NVIDIA A100 GPUs. They need to enable vGPU functionality for virtual machines running AI workloads. Which configuration step is required on the ESXi host to allow vGPU assignment to VMs?
Hard230An AI operations engineer is troubleshooting a Kubernetes cluster where several GPU training pods fail to start with a device plugin allocation error, even though the nodes report healthy GPUs. The engineer suspects the pods are requesting more GPU resources than a single physical card can provide without a sharing mechanism. Which TWO configurations would legitimately allow multiple pods to consume a single physical GPU on these nodes? (Choose two.)
Medium231An administrator is managing a Kubernetes cluster with NVIDIA GPU Operator. A new policy requires that all GPU workloads run with specific environment variables set, such as NVIDIA_VISIBLE_DEVICES and NVIDIA_DRIVER_CAPABILITIES. The administrator wants to enforce these variables automatically for any pod that requests a GPU, without modifying each pod specification manually. Which approach should the administrator use?
Medium232A platform team is deploying the NVIDIA GPU Operator on a Kubernetes cluster to manage GPU nodes. They want the Operator to automatically install the NVIDIA driver, the container toolkit, and the device plugin on each GPU node. Which component of the GPU Operator is responsible for installing the NVIDIA driver on the host?
Medium233An AI researcher is debugging a multi-node training job using NCCL. Which TWO actions should they take to diagnose potential network-related performance degradation?
Hard234Which utility is primarily used to monitor and manage NVIDIA GPU power, temperature, and usage statistics in real-time on a Linux-based deployment?
Easy235Refer to the exhibit. An administrator attempts to deploy a GPU-based pod, but it remains in the 'Pending' state. What is the most likely cause based on the error log?
Medium236An AI operations engineer is optimizing a PyTorch training job on an NVIDIA DGX A100 system. The job uses a data loader with multiple workers, but the engineer observes that GPU utilization fluctuates between 40% and 60%, and `nvidia-smi dmon` shows periods of zero GPU utilization. The engineer suspects that the data input pipeline is the bottleneck. Which TWO actions should the engineer take to improve GPU utilization? (Choose two.)
Hard237A team trains a model inside an NGC PyTorch container on a DGX H100 node. Training starts, but after a few minutes the process dies and `dmesg` shows `Xid 79: GPU has fallen off the bus` on one GPU. The team needs to determine whether the fault is hardware or software before opening an RMA. Which two actions should they take to gather useful evidence? (Choose two.)
Medium238An administrator is preparing to install the NVIDIA GPU Operator on a new Kubernetes cluster. The cluster uses containerd as the container runtime. Which prerequisite must be satisfied on each GPU node before the operator can successfully deploy the driver container?
Easy239A team is running a multi-GPU training job on an NVIDIA DGX A100 system using NCCL for inter-GPU communication. Training throughput is much lower than expected, and the NCCL logs show repeated 'NCCL WARN Call to ibv_reg_mr failed' errors. The job uses a container with host networking. Which action should the AI operations engineer take to resolve the issue?
Hard240A data science team submits a PyTorch distributed training job to a Kubernetes cluster with the NVIDIA GPU Operator installed. The job's pods repeatedly fail with a CUDA initialization error, while a simple `nvidia-smi` check inside an interactive pod on the same node succeeds. The administrator confirms the node's driver is healthy and the device plugin is advertising GPUs. Which configuration should the administrator verify first?
Easy241A research organization runs an NVIDIA DGX SuperPOD with a Kubernetes cluster managed by the NVIDIA GPU Operator and Network Operator. A distributed training job using PyTorch DDP across 32 nodes stalls at initialization, and the administrator suspects the collective communication library is not selecting the high-speed fabric. Which configuration should the administrator verify first to ensure NCCL uses the correct network interface and topology?
Hard242Which administrative practice ensures that a cluster is prepared for the arrival of new NVIDIA GPU hardware with minimal downtime?
Medium243An administrator is preparing to deploy an NVIDIA AI Enterprise solution on an OpenShift cluster. Which TWO steps must be completed to ensure the NVIDIA drivers are loaded correctly on the worker nodes?
Medium244An administrator is responsible for maintaining an NVIDIA AI Enterprise cluster and needs to ensure high availability of GPU resources for critical inference workloads. Which two practices should the administrator implement? (Choose two.)
Medium245Which THREE factors should be considered when estimating GPU memory requirements for a Large Language Model (LLM) fine-tuning job?
Hard246Refer to the exhibit. During a multi-node training job, communication between nodes fails. What is the most likely cause of this error?
Medium247Refer to the exhibit. An AI administrator investigates why a GPU node is performing significantly slower than expected. Based on the output, what is the most likely cause?
Medium248An AI operations engineer manages a shared Kubernetes cluster running NVIDIA GPU Operator. Several teams report that their inference pods remain in a Pending state with the event '0/8 nodes are available: 8 Insufficient nvidia.com/gpu.' The administrator verifies that nvidia-smi on all nodes shows idle GPUs. Which action should the administrator take first to resolve the scheduling failure?
Medium249An AI operations engineer is preparing a Kubernetes cluster to run GPU-accelerated inference workloads using the NVIDIA GPU Operator. The cluster nodes already have NVIDIA data center GPUs installed, and the engineer wants to avoid installing the driver manually on each node. Which component of the GPU Operator is responsible for automatically deploying the NVIDIA driver on worker nodes?
Easy250An administrator manages an NVIDIA AI Enterprise cluster using NVIDIA Run:ai. A data science team complains that their submitted training job has been stuck in a Pending state for over an hour, even though the Run:ai scheduler shows free GPUs in the cluster. The administrator verifies that the job requests 2 GPUs and the node pool has 4 idle GPUs. Which Run:ai administrative configuration is the most likely cause of the job remaining Pending?
Medium251A platform engineer is preparing a Kubernetes cluster to run AI training jobs that require GPU access. The cluster nodes have NVIDIA GPUs, and the engineer wants the GPU Operator to manage the driver lifecycle. Which component must be installed on the host nodes to allow the GPU Operator to load kernel modules and create device nodes?
Medium252An administrator is planning to monitor GPU utilization across a large cluster. Which component should be deployed to collect metrics that are compatible with Prometheus?
Medium253What is the primary benefit of using an NVIDIA-Certified System for AI Enterprise deployments?
Medium254When deploying NVIDIA AI Enterprise, why is the selection of the correct CUDA version in the container image critical during the installation phase?
Hard255Which TWO of the following are benefits of using containerized GPU workloads compared to bare-metal deployment?
Medium256An inference service runs a 70B parameter model with TensorRT-LLM on a single H100 using in-flight batching. Operators report that time-to-first-token is acceptable, but inter-token latency degrades sharply once concurrent request count exceeds a certain point, and GPU memory utilization sits near 98 percent. Which change most directly addresses the inter-token latency degradation?
Hard257An AI operations team runs a shared Kubernetes cluster with the NVIDIA GPU Operator and several namespaces owned by different groups. A group reports that its training pods are stuck Pending with an event indicating insufficient nvidia.com/gpu, yet cluster-wide dashboards show many GPUs idle. Investigation reveals the idle GPUs belong to nodes labeled for another group, and the affected namespace has a node affinity rule pinning its pods to a specific GPU generation that is fully consumed. Which action best resolves the Pending pods while respecting multi-tenant boundaries?
Hard258An administrator is responsible for monitoring a large-scale AI cluster with hundreds of NVIDIA GPUs. They need to collect telemetry data such as GPU utilization, temperature, and power consumption from all nodes and store it centrally for analysis and alerting. Which NVIDIA tool should they use to collect and export GPU metrics to a monitoring system like Prometheus?
Easy259An AI operations engineer is deploying a multi-node Kubernetes cluster with NVIDIA A100 GPUs for distributed training. The engineer wants to ensure that GPUs are correctly discovered and that workloads can request GPU resources. After installing the NVIDIA GPU Operator, the engineer notices that the GPU nodes are not advertising any 'nvidia.com/gpu' resources. Which component should the engineer verify first to resolve this issue?
Hard260Which TWO of the following steps are essential when deploying the NVIDIA GPU Operator on a Kubernetes cluster to ensure that GPU resources are discoverable by the scheduler?
Medium261An administrator manages an NVIDIA AI Enterprise cluster and needs to enforce GPU resource quotas across multiple Kubernetes namespaces. Which NVIDIA component should be configured to enforce these quotas?
Medium262An AI operations engineer is investigating intermittent failures in a long-running distributed training job. The job occasionally aborts with a collective timeout error, but no GPU errors, ECC events, or fabric link flaps appear in logs. Which action should the engineer take first to identify the root cause?
Hard263When managing GPU resources in a shared cluster, which configuration best prevents 'noisy neighbor' scenarios where one GPU task consumes all available memory bandwidth?
Medium264When managing large-scale model training jobs, what is the primary purpose of using a Job Scheduler like Slurm or Kubernetes Batch?
Easy265An engineer is troubleshooting a CUDA program that terminates unexpectedly. Which tool should be used to detect memory leaks and race conditions in the CUDA kernel code?
Medium266Which NVIDIA technology allows for partitioning a single physical GPU into multiple independent instances, each with dedicated compute and memory resources for smaller workloads?
Easy267An operations team runs a multi-node NCCL all-reduce training job across four DGX nodes connected via InfiniBand. Training throughput is far below the expected linear scaling, and `nvidia-smi` shows GPU utilization oscillating between 20% and 40%. The network fabric is healthy and the GPUs are not thermally throttled. Which diagnostic step is MOST appropriate to identify the bottleneck?
Medium268Which security configuration is necessary when deploying NVIDIA GPUs in a multi-tenant environment to prevent unauthorized access between containers?
Medium269What is the primary function of an 'InitContainer' in an NVIDIA GPU-enabled pod deployment?
Medium270An administrator is preparing a bare-metal GPU server for AI workloads and needs to verify that the NVIDIA driver and CUDA toolkit are properly installed. The server has an NVIDIA A100 GPU. Which command should the administrator run to display the GPU model, driver version, and CUDA version?
Medium271An AI operations engineer is troubleshooting an inference service on an NVIDIA A100 GPU that shows intermittent stalls. The monitoring dashboard reports GPU utilization at 100%, but request throughput is far below the validated baseline. Running nvidia-smi dmon reveals the SM utilization is high while memory controller utilization is low. Which action should the engineer take first to identify the bottleneck?
Medium272What is the primary function of the NVIDIA Persistence Daemon in an AI deployment?
Medium273A financial services firm must prove to auditors that an AI training job ran on hardware located only in its Frankfurt data center and that no pod could ever be scheduled onto GPUs in other regions. The cluster spans three regions with nodes labeled topology.kubernetes.io/region. Which approach most directly enforces this placement requirement?
Easy274An AI operations team is deploying NVIDIA AI Enterprise on a Kubernetes cluster using the NVIDIA GPU Operator. They need to ensure that GPU metrics such as utilization, memory usage, and temperature are collected and exposed to Prometheus for monitoring. Which component of the GPU Operator is responsible for this?
Medium275An administrator needs to ensure that all GPU drivers are updated across a heterogeneous cluster without causing downtime. What is the best strategy?
Medium276An administrator is deploying the NVIDIA GPU Operator on a Kubernetes cluster in an air-gapped environment. The cluster nodes have no internet access, and all container images must be pulled from a private registry. Which two actions are required to ensure a successful deployment? (Choose two.)
Medium277Refer to the exhibit. An administrator observes this in the logs during a multi-GPU training job. What is the performance implication of this setting?
Hard278When configuring a node for NVIDIA AI Enterprise in a Kubernetes environment, what is the primary function of the NVIDIA Container Toolkit?
Medium279Which strategy is most effective for managing heterogeneous GPU clusters containing both older architectures (e.g., V100) and newer architectures (e.g., H100)?
Medium280A platform team is preparing a bare-metal Kubernetes cluster to run AI workloads. They want the GPU Operator to install and manage the NVIDIA driver automatically on each node. Which prerequisite must be satisfied on every worker node before the GPU Operator can succeed?
Easy281A platform team runs an NVIDIA AI cluster with the GPU Operator deployed. Users submit jobs directly with kubectl and frequently request whole GPUs even when their notebooks only need a fraction of one. The team wants Kubernetes itself to admit and queue jobs based on GPU demand without users changing their manifests. Which component should the team deploy to meet this requirement?
Medium282An organization is migrating AI workloads to a private cloud. Which feature is essential for ensuring that GPU resources are dynamically reclaimed and reallocated to different departments without manual intervention?
Medium283A cloud architect is deploying NVIDIA AI Enterprise on a vSphere cluster with multiple ESXi hosts, each containing NVIDIA A100 GPUs. The architect must enable GPU virtualization using NVIDIA vGPU. Which two components are required to support vGPU on the ESXi hosts? (Choose two.)
Hard284A DevOps engineer needs to monitor GPU health metrics in real-time for workload management. Which tool provides the most granular visibility into GPU utilization and power consumption for individual containers?
Medium285When configuring a multi-tenant environment on an NVIDIA DGX system, how are MIG (Multi-Instance GPU) instances best provisioned?
Hard286A platform team is installing the NVIDIA GPU Operator on a Kubernetes cluster that runs a mix of GPU and non-GPU nodes. They want the operator to manage the driver lifecycle only on nodes that actually have NVIDIA GPUs, without requiring manual taints on non-GPU nodes. Which configuration should they apply to the ClusterPolicy to achieve this?
Medium287A data science team submits a PyTorch training job to a Kubernetes cluster where the NVIDIA GPU Operator is installed. The pod stays in Pending state, and kubectl describe shows the message '0/6 nodes are available: 6 Insufficient nvidia.com/gpu.' The administrator confirms the nodes have healthy GPUs and the device plugin pods are Running. What is the most likely cause?
Easy288An AI operations team is running a large language model inference service on NVIDIA H100 GPUs using NVIDIA Triton Inference Server. They observe that the first inference request after a period of inactivity takes significantly longer than subsequent requests. The model is loaded and ready, but the GPU shows low utilization during the first request. Which optimization should the team implement to reduce this latency spike?
Hard289An administrator is tasked with deploying a multi-node training job using the NVIDIA GPU Operator. Which configuration must be present to ensure that pods are scheduled on nodes with identical GPU architectures to prevent performance degradation?
Hard290An MLOps engineer needs to guarantee that a latency-sensitive inference Deployment always has GPU capacity available, even when a large training Job is submitted to the same namespace. The cluster uses the NVIDIA GPU Operator and nodes have four A100 GPUs each. Which approach reliably reserves GPU capacity for the inference Deployment?
Hard291A Kubernetes cluster running the NVIDIA GPU Operator is shared by an inference team and a research team. The research team's training pods repeatedly evict the inference pods from GPUs, causing latency spikes in production. The administrator wants to guarantee that inference pods always get GPU access first. Which Kubernetes scheduling mechanism should be configured?
Medium292Refer to the exhibit. The system has two GPUs. What is the most likely cause of the observed performance discrepancy?
Hard293An AI infrastructure team is deploying NVIDIA AI Enterprise on a Kubernetes cluster using the NVIDIA GPU Operator. They need to ensure that the GPU Operator can successfully manage GPUs and that workloads can consume GPU resources. Which two components does the GPU Operator deploy to enable GPU scheduling and container GPU access? (Choose two.)
Hard294A team is deploying a large language model for inference using NVIDIA TensorRT-LLM on an H100 GPU. They observe that the first inference request takes several seconds, while subsequent requests are fast. They want to reduce this initial latency. Which technique should they implement?
Hard295A data science team submits a PyTorch training job to a Kubernetes cluster managed by Run:ai. The job requests two GPUs but only one is allocated, and the second worker hangs waiting for a peer. Which Run:ai capability should the administrator verify is configured so the distributed job receives all requested GPUs atomically?
Easy296An administrator wants to ensure that a training process is limited to a single GPU on a multi-GPU node. Which environment variable should be set?
Hard297An AI engineer observes that a model training job on an NVIDIA DGX system is underutilizing the GPU. The training loop shows frequent "CPU bottleneck" warnings in the logs. Which action should the engineer take first to optimize throughput?
Medium298An AI researcher is deploying a multi-node training job using NCCL on an InfiniBand network. The job is suffering from intermittent latency spikes. Which TWO steps should the engineer perform to troubleshoot the network configuration?
Hard299An administrator notices that GPU utilization is high, but throughput in an AI training job remains low. What is the most likely bottleneck?
Medium300An administrator is deploying NVIDIA AI Enterprise on a vSphere cluster with multiple ESXi hosts, each containing several NVIDIA A100 GPUs. They need to enable GPU sharing across multiple virtual machines to maximize utilization. Which vSphere feature should they configure?
Easy301If a GPU job consistently fails with 'Out of Memory' despite the model size being significantly smaller than the total VRAM, what is the most likely cause?
Medium302An administrator is deploying the NVIDIA GPU Operator into an existing Kubernetes cluster where the NVIDIA driver is already installed and maintained by the node image. The team wants the Operator to manage only the container runtime, device plugin, and monitoring components. Which configuration should be applied to the GPU Operator's ClusterPolicy?
Medium303An administrator is deploying NVIDIA AI Enterprise on a bare-metal cluster. Which component must be installed first to ensure proper communication between the Kubernetes scheduler and the underlying GPU hardware?
Medium304An operations engineer notices that an inference container on an A100 is intermittently returning stale predictions after a model update. The container mounts the model directory from a host path, and the update process replaces files in place. Which change most reliably prevents the stale predictions?
Easy305Refer to the exhibit. An administrator is attempting to deploy a job to a namespace with a ResourceQuota defined. What is the cause of this error?
Hard306Which component in the NVIDIA AI Enterprise stack is responsible for providing the necessary user-space libraries and binaries to run GPU-accelerated applications inside containers?
Easy307A platform team operates a Kubernetes cluster where several teams submit GPU training jobs. The administrator needs to enforce per-namespace limits on the number of GPUs that can be consumed and prevent a single namespace from monopolizing all GPU capacity. Which TWO Kubernetes resources should be configured to achieve this? (Choose two.)
Hard308Refer to the exhibit. The training job is showing intermittent "thermal throttling" warnings. Which configuration change is the most appropriate adjustment?
Hard309A distributed training job using PyTorch DDP across eight GPUs on one DGX A100 node shows GPU utilization oscillating between 20 and 40 percent, while `nvidia-smi dmon` shows low SM activity but sustained high memory-controller utilization. The data loader reads from a local NVMe RAID array and applies heavy CPU augmentation. Which action best improves GPU utilization?
MediumOther domains
All NCP-AIO exam domains
Frequently asked questions
- What does the scenario questions domain cover on the NCP-AIO exam?
- scenario questions questions test whether you can apply the concept in context, not just recognise a definition.
- How many questions are in this domain?
- This page lists all 309 scenario questions questions in the NCP-AIO question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only scenario questions questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.