Sample questions
NVIDIA Certified Professional: AI Operations practice questions
Refer to the exhibit. An administrator notices that 'user_a' is consistently hitting resource limits despite having sufficient total system GPU memory. Based on the policy JSON, wh…
Refer to the exhibit. An administrator applies this security policy to a container runtime environment. What is the immediate effect on containerized AI applications within this sc…
Which NVIDIA technology enables a GPU to be shared among multiple virtual machines or containers while maintaining strict hardware isolation?
An administrator is optimizing multi-GPU utilization. Which TWO of the following configurations allow multiple containers to share a single physical GPU on a supported NVIDIA archi…
An administrator is optimizing a cluster for AI model training using NVIDIA Base Command. Which TWO tasks are critical for ensuring consistent performance across the training nodes…
An AI researcher is running a large-scale training job on an NVIDIA DGX system using Kubernetes. They observe that GPU utilization is consistently low despite high CPU load. Which…
Which THREE components are required for a container to successfully leverage NVIDIA GPUs on a Kubernetes cluster?
A system administrator is troubleshooting a 'CUDA error: invalid device ordinal' when launching a job on a multi-GPU system. What is the most likely cause?
During the deployment of an AI model training workload on a multi-node cluster, the administrator notices that inter-node communication is significantly slower than expected. Which…
A platform team is deploying the NVIDIA GPU Operator on a Kubernetes cluster to manage GPU nodes. They want the Operator to automatically install the NVIDIA driver, the container t…
An AI researcher is debugging a multi-node training job using NCCL. Which TWO actions should they take to diagnose potential network-related performance degradation?
Which utility is primarily used to monitor and manage NVIDIA GPU power, temperature, and usage statistics in real-time on a Linux-based deployment?
Refer to the exhibit. An administrator attempts to deploy a GPU-based pod, but it remains in the 'Pending' state. What is the most likely cause based on the error log?
An administrator wants to prevent unauthorized users from accessing sensitive model weights stored in GPU memory. Which security feature should be implemented to ensure hardware-le…
When debugging a workload that consistently crashes with 'Out of Memory' (OOM) errors despite sufficient GPU VRAM, what is the most likely cause related to workload management?
Refer to the exhibit. A cluster administrator notices that GPU jobs with this PriorityClass are failing to start even when empty GPUs are available. What is the most likely cause?
Which TWO of the following steps are essential when deploying the NVIDIA GPU Operator on a Kubernetes cluster to ensure that GPU resources are discoverable by the scheduler?
Which administrative practice ensures that a cluster is prepared for the arrival of new NVIDIA GPU hardware with minimal downtime?
Which THREE factors should be considered when estimating GPU memory requirements for a Large Language Model (LLM) fine-tuning job?
An administrator is planning to monitor GPU utilization across a large cluster. Which component should be deployed to collect metrics that are compatible with Prometheus?
A production inference service using TensorRT is showing lower than expected throughput. Profiling shows that the model is spending significant time in "host-to-device" transfers.…
Which TWO of the following are benefits of using containerized GPU workloads compared to bare-metal deployment?
When managing large-scale model training jobs, what is the primary purpose of using a Job Scheduler like Slurm or Kubernetes Batch?
An engineer is troubleshooting a CUDA program that terminates unexpectedly. Which tool should be used to detect memory leaks and race conditions in the CUDA kernel code?