Courseiva
← Back to NVIDIA Certified Professional: AI Operations questions

Scenario-based practice

Refer to the Exhibit Practice Questions

Practise NVIDIA Certified Professional: AI Operations practice questions — original exam-style scenarios covering every exam domain, with detailed explanations, wrong-answer analysis, and common exam traps.

15
scenario questions
NCP-AIO
exam code
NVIDIA
vendor

Scenario guide

How to approach refer to the exhibit practice questions

Practise exhibit-style questions that ask you to read a topology, table, command output or diagram before choosing the best answer.

Quick answer

Exhibit-style questions test whether you can read a topology, command output, diagram or table before choosing the best answer.

How to extract the relevant detail from an exhibit.

How topology, command output or routing information affects the answer.

How to avoid answering from memory before reading the evidence.

How to map the exhibit back to the exam objective.

Related practice questions

Related NCP-AIO topic practice pages

Scenario questions usually connect to one or more exam topics. Use these links to review the underlying concepts behind the scenario.

Practice set

Practice scenarios

Question 1hardmultiple choice
Full question →

Refer to the exhibit. An administrator notices that 'user_a' is consistently hitting resource limits despite having sufficient total system GPU memory. Based on the policy JSON, what is the cause?

Exhibit

{
  "policy": "restrict_gpu_access",
  "targets": ["user_a"],
  "max_concurrent_jobs": 2,
  "resource_quota": {
    "gpu_memory": "16GB"
  }
}
Question 2hardmultiple choice
Full question →

Refer to the exhibit. An administrator applies this security policy to a container runtime environment. What is the immediate effect on containerized AI applications within this scope?

Exhibit

{ "policy": "deny", "resources": ["/dev/nvidia*"], "action": "restrict_access" }
Question 3mediummultiple choice
Full question →

Refer to the exhibit. An administrator attempts to deploy a GPU-based pod, but it remains in the 'Pending' state. What is the most likely cause based on the error log?

Exhibit

NAME: gpu-pod
STATUS: Pending
EVENTS:
  Warning: FailedScheduling: 0/3 nodes are available: 3 Insufficient nvidia.com/gpu.
Question 4hardmultiple choice
Full question →

Refer to the exhibit. A cluster administrator notices that GPU jobs with this PriorityClass are failing to start even when empty GPUs are available. What is the most likely cause?

Exhibit

apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
  name: high-priority-gpu
value: 1000000
globalDefault: false
description: "High priority for distributed training"
Question 5easymultiple choice
Full question →

A financial services firm must prove to auditors that an AI training job ran on hardware located only in its Frankfurt data center and that no pod could ever be scheduled onto GPUs in other regions. The cluster spans three regions with nodes labeled topology.kubernetes.io/region. Which approach most directly enforces this placement requirement?

Question 6hardmultiple choice
Full question →

Refer to the exhibit. The system has two GPUs. What is the most likely cause of the observed performance discrepancy?

Network Topology
nvidia-smiquery-gpu=utilization.gpuformat=csvutilization.gpu [%], memory.used [MiB]98 %, 15400 MiB25 %, 15200 MiB
Question 7hardmultiple choice
Full question →

Refer to the exhibit. The training job is showing intermittent "thermal throttling" warnings. Which configuration change is the most appropriate adjustment?

Exhibit

JSON Policy Configuration:
{
  "persistence": "enabled",
  "ecc_mode": "enabled",
  "compute_mode": "default",
  "power_limit_watts": 250
}
Question 8mediummultiple choice
Full question →

Refer to the exhibit. An AI engineer observes that a model training job is running slower than expected. Based on the output, what is the primary cause of the performance degradation?

Exhibit

nvidia-smi -q -d PERFORMANCE

Performance State : P12
Clocks Throttle Reasons : SW Power Cap
Question 9hardmultiple choice
Full question →

Refer to the exhibit. An administrator observes this in the logs during a multi-GPU training job. What is the performance implication of this setting?

Exhibit

Error Log: [NCCL] NCCL_P2P_DISABLE=1 detected. Using network-based communication instead of NVLink.
Question 10hardmultiple choice
Full question →

Refer to the exhibit. An administrator is attempting to deploy a job to a namespace with a ResourceQuota defined. What is the cause of this error?

Exhibit

Error: Failed to create pod: pods "gpu-job" is forbidden: 
failed quota: gpu-quota: must specify limits for nvidia.com/gpu
Question 11mediummultiple choice
Full question →

Refer to the exhibit. An administrator notices poor performance in an AI training job. What is the most likely cause based on the CLI output?

Exhibit

nvidia-smi -q -d PERFORMANCE

Performance State : P0
Clocks Throttle Reasons : Active
   Clocks Throttle Reason Sw Power Cap : Active
   Clocks Throttle Reason HW Slowdown : Not Active
Question 12mediummultiple choice
Full question →

Refer to the exhibit. The administrator has deployed the GPU Operator, but the node does not show GPU resources. What is the most likely cause?

Exhibit

kubectl get nodes
NAME    STATUS   ROLES   AGE   VERSION
node-1  Ready    <none>  10d   v1.28.0

kubectl describe node node-1 | grep nvidia
<no output>
Question 13mediummultiple choice
Full question →

Refer to the exhibit. An AI administrator investigates why a GPU node is performing significantly slower than expected. Based on the output, what is the most likely cause?

Exhibit

nvidia-smi -q -d PERFORMANCE
Performance State : P0
Clocks Throttle Reasons : Active
  Applications Clocks Setting : None
  SW Power Cap : Active
  HW Slowdown : Active
  HW Thermal Slowdown : Active
Question 14mediummultiple choice
Full question →

Refer to the exhibit. During a multi-node training job, communication between nodes fails. What is the most likely cause of this error?

Exhibit

Error Log: [NCCL WARN] NET/Socket : Connection refused. [NCCL WARN] Call to connect() failed. Rank 0: GPU 0: Peer 1 is unreachable via NCCL_SOCKET_IFNAME.
Question 15hardmultiple choice
Full question →

A research organization runs an NVIDIA DGX SuperPOD with a Kubernetes cluster managed by the NVIDIA GPU Operator and Network Operator. A distributed training job using PyTorch DDP across 32 nodes stalls at initialization, and the administrator suspects the collective communication library is not selecting the high-speed fabric. Which configuration should the administrator verify first to ensure NCCL uses the correct network interface and topology?

These NCP-AIO practice questions are part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style NCP-AIO questions with detailed explanations, topic-based practice, mock exams, readiness tracking, and study analytics.