Courseiva

CCNA Installation and Deployment Questions

75 of 92 questions · Page 1/2 · Installation and Deployment · Answers revealed

1
MCQmedium

A media company is deploying an inference service on a Kubernetes cluster with the NVIDIA GPU Operator installed. The service pods remain in Pending with the message that no nodes have the requested nvidia.com/gpu resource, even though the GPUs are healthy and the driver loads correctly on every node. Which troubleshooting step should the engineer perform first?

A.Reinstall the NVIDIA driver on every node to refresh the kernel module.
B.Increase the kubelet's pod density limit so the plugin can be scheduled.
C.Check whether the NVIDIA device plugin pods are running and have registered the GPUs with kubelet.
D.Enable Multi-Instance GPU mode on all GPUs so each partition advertises capacity.
AnswerC

The scheduler only sees nvidia.com/gpu capacity after the device plugin registers each GPU with kubelet through the plugin socket. If the device plugin pods are crashing or not scheduled, nodes show no GPU resource and pods stay Pending. Verifying plugin pod status and logs is the most direct first step when the driver itself is healthy.

Why this answer

Kubernetes learns about GPUs through the device plugin framework: the NVIDIA device plugin advertises nvidia.com/gpu for each visible GPU by registering with kubelet. When the driver is healthy but pods report no such resource, the plugin is the missing link. Checking its pod status and logs quickly reveals scheduling failures, crashes, or socket registration errors before deeper investigation.

Exam trap

The trap here is jumping to driver or hardware remediation when the symptom, a healthy driver with no advertised nvidia.com/gpu resource, points squarely at the device plugin registration path.

2
MCQeasy

A DevOps engineer is deploying the NVIDIA GPU Operator on a Kubernetes cluster that uses containerd as the container runtime. The engineer notices that the Operator's validation pod fails with an error indicating that the NVIDIA container runtime is not configured. Which action should the engineer take to resolve this?

A.Manually install the NVIDIA Container Toolkit on all nodes and set the default runtime to nvidia.
B.Switch the cluster to use Docker as the container runtime, as the GPU Operator only supports Docker.
C.Ensure that the GPU Operator's container-toolkit daemonset is enabled and has the necessary permissions to modify the container runtime configuration.
D.Disable the validation pod in the GPU Operator's Helm chart to suppress the error.
AnswerC

The GPU Operator deploys a container-toolkit daemonset that installs and configures the NVIDIA Container Toolkit. If this daemonset is disabled or lacks permissions (e.g., privileged access), it cannot modify the containerd configuration. Enabling it and ensuring proper RBAC and security context allows the Operator to set up the runtime, resolving the validation error.

Why this answer

The GPU Operator uses a container-toolkit daemonset to install and configure the NVIDIA Container Toolkit on each node, including modifying the containerd configuration. If this daemonset is disabled or lacks privileges, the runtime won't be set up, causing validation failures. Ensuring the daemonset is enabled and has proper permissions resolves the issue.

Exam trap

The trap here is thinking manual installation or switching runtimes is needed, when the GPU Operator is designed to handle runtime configuration automatically.

3
MCQeasy

When installing the NVIDIA Container Toolkit to enable GPU acceleration in Docker, which file must be modified or verified to ensure the container runtime can access the NVIDIA runtime?

A./etc/fstab
B./etc/docker/daemon.json
C./etc/environment
D./boot/grub/grub.cfg
AnswerB

This is the primary configuration file for the Docker daemon. Including the nvidia-container-runtime in this JSON configuration is mandatory to register the runtime with Docker. This allows the engine to recognize the --gpus flag, enabling GPU hardware pass-through for accelerated workloads inside containers.

Why this answer

The daemon configuration file, typically located at /etc/docker/daemon.json, must be updated to include the NVIDIA runtime. This integration is the foundational step for AI containerization, as it allows the Docker engine to map host GPU resources into the container namespace. Without this configuration, containers will fail to detect GPUs, rendering them incapable of running accelerated AI applications.

Exam trap

Candidates often look for environment variables or shell scripts, missing that the persistent configuration for the Docker runtime resides in the /etc/docker/daemon.json file.

4
MCQeasy

A systems administrator is installing the NVIDIA Container Toolkit on a stand-alone server running Ubuntu 22.04 to enable Docker containers to access NVIDIA GPUs. After installation, they run a test container and find that it cannot see the GPU. Which step is most likely missing?

A.Rebooting the server after installing the toolkit.
B.Adding the user to the 'docker' group to run containers with GPU support.
C.Configuring the Docker daemon to use the NVIDIA container runtime as the default runtime.
D.Installing the NVIDIA GPU driver on the host.
AnswerC

The NVIDIA Container Toolkit requires the Docker daemon to be configured to use the 'nvidia' runtime. This is typically done by adding 'nvidia' to the 'runtimes' section in '/etc/docker/daemon.json' and optionally setting it as the default. Without this configuration, Docker will not use the NVIDIA runtime, and containers will not have GPU access, even if the toolkit is installed.

Why this answer

After installing the NVIDIA Container Toolkit, the Docker daemon must be configured to use the NVIDIA container runtime. This involves editing '/etc/docker/daemon.json' to add the runtime and restarting Docker. Without this configuration, containers will not have GPU access, even though the toolkit is installed.

Exam trap

The trap here is thinking that installing the toolkit alone is sufficient, when Docker must also be told to use the NVIDIA runtime.

5
MCQmedium

A platform engineer is preparing an Ubuntu 22.04 server that will host GPU-accelerated inference containers managed by containerd (not Docker). The team wants the NVIDIA Container Toolkit to expose GPUs to those containers. After installing the toolkit packages, which action must the engineer take so that containerd actually invokes the NVIDIA runtime for GPU workloads?

A.Install the nvidia-container-runtime package and symlink it as /usr/bin/runc on the host.
B.Add the user to the video group and grant read-write access to /dev/nvidiactl.
C.Set the environment variable NVIDIA_VISIBLE_DEVICES=all in the host shell profile and reboot the node.
D.Run nvidia-ctk runtime configure --runtime=containerd and restart the containerd service.
AnswerD

The nvidia-ctk runtime configure command edits the containerd configuration (typically /etc/containerd/config.toml) to register the NVIDIA runtime and set it as the default, and containerd must then be restarted to load the change. Without this registration, containerd keeps using its stock runc runtime and GPU devices never appear inside containers, even though the toolkit binaries are installed.

Why this answer

Registering the NVIDIA runtime with the container engine is the essential post-install step. The nvidia-ctk runtime configure command writes the runtime entry and default-runtime setting into containerd's config.toml, and containerd must be restarted to apply it. Merely installing packages or setting container environment variables does not change which runtime containerd uses, so GPU devices remain invisible to containers until the engine configuration is updated and reloaded.

Exam trap

The trap here is assuming that installing the NVIDIA Container Toolkit packages is sufficient and that containers automatically gain GPU access without reconfiguring the container engine's runtime.

6
MCQmedium

You are troubleshooting a node where the GPU is detected, but the application fails to utilize it. Which log source would provide the most relevant information?

A.The BIOS system event log.
B.The NVIDIA container runtime logs.
C.The cluster's physical network switch logs.
D.The local NTP synchronization logs.
AnswerB

The container runtime logs show the interaction between the runtime and the GPU drivers during container instantiation. If the runtime fails to inject the necessary libraries or access the GPU device, these logs will capture the error, which is the most likely cause when a GPU is physically detected.

Why this answer

The NVIDIA container runtime logs and the application-level logs are the most important sources. If the GPU is visible to the system but not the application, the issue is likely a driver/runtime mismatch or a library path configuration. Checking these logs allows an administrator to isolate whether the fault is in the container orchestration layer or the application's software environment, which is vital for rapid resolution in production environments.

Exam trap

Candidates often suggest checking the kernel logs or application code itself, overlooking that the NVIDIA container runtime is the specific layer responsible for bridging the GPU to the containerized application.

7
MCQhard

In an air-gapped environment, what must an administrator do to ensure the GPU Operator correctly installs the necessary software components?

A.Enable the 'offline-mode' flag in the Kubernetes API.
B.Configure the GPU Operator to use a local image registry.
C.Manually copy the VIB files to every node.
D.Disable the validation of container image signatures.
AnswerB

To succeed in air-gapped environments, the GPU Operator must be instructed to pull its images from an internally accessible registry. This involves mapping image paths and ensuring that all dependencies are hosted locally, which allows the deployment to proceed without needing external network egress to public servers.

Why this answer

In an air-gapped environment, the cluster cannot reach external repositories (e.g., NGC or Docker Hub). The administrator must pre-populate a local container registry with all required images and update the Operator configuration to point to this local registry. This is a common enterprise task for secure environments that prevents deployment failures caused by connection timeouts or authentication issues when accessing public NVIDIA software mirrors.

Exam trap

Candidates mistakenly think they can rely on standard internet-based NGC pulls by adjusting firewall rules, ignoring the true definition of an air-gapped network.

8
MCQmedium

Which TWO of the following are prerequisites for installing the NVIDIA Container Toolkit on a Linux host?

A.A compatible NVIDIA driver already installed on the host.
B.The latest version of the CUDA Toolkit installed in every container.
C.A container runtime such as Docker or containerd.
D.An active subscription to NVIDIA AI Enterprise.
E.A pre-configured Kubernetes cluster with Helm.
AnswerA, C

The NVIDIA driver acts as the kernel-mode component that communicates with the hardware. The container toolkit is essentially a wrapper that relies on this driver to provide GPU access to containers. Without a working driver, the toolkit has no hardware interface to pass through to containers.

Why this answer

To successfully deploy the NVIDIA Container Toolkit, the host must have a functional NVIDIA driver and a container runtime like Docker or containerd. These prerequisites ensure that the runtime has a target to interface with and can correctly map host-side GPU resources into the container namespace. Without these core components, the toolkit cannot bridge the gap between physical hardware and isolated containerized processes.

Exam trap

Candidates select guest OS configurations or specific AI frameworks, forgetting that container toolkits strictly depend on low-level host drivers and runtimes.

9
MCQmedium

An AI operations team is deploying the NVIDIA GPU Operator on a Kubernetes cluster that uses containerd as the container runtime. The cluster nodes have NVIDIA GPUs, and the team wants to ensure that GPU workloads can request GPU resources. After installing the operator, they notice that pods requesting 'nvidia.com/gpu' remain in Pending state. Which component of the GPU Operator is most likely misconfigured or missing?

A.The NVIDIA device plugin, which advertises GPU resources to the Kubernetes API server.
B.The Kubernetes scheduler configuration, which may not be aware of GPU resources.
C.The NVIDIA Container Toolkit, which enables containers to access GPUs.
D.The GPU Operator's driver container, which loads the NVIDIA kernel modules.
AnswerA

The NVIDIA device plugin is responsible for discovering GPUs on each node and advertising them as schedulable resources like 'nvidia.com/gpu'. If it is not running or misconfigured, the Kubernetes scheduler will not see any GPU resources, causing pods that request them to remain Pending. This is the most direct cause for the described symptom.

Why this answer

The NVIDIA device plugin is a DaemonSet that runs on each node and registers GPUs as extended resources. When it is missing or misconfigured, the Kubernetes API server has no knowledge of available GPUs, so any pod requesting 'nvidia.com/gpu' cannot be scheduled. Ensuring the device plugin is healthy and running is essential for GPU scheduling.

Exam trap

The trap here is confusing the role of the NVIDIA Container Toolkit with that of the device plugin; the toolkit enables GPU access at runtime, but the device plugin is what makes GPUs visible to the scheduler.

10
Multi-Selecthard

When troubleshooting an NVIDIA GPU Operator installation, which TWO locations should an administrator check to identify why the driver installation pod is failing?

Select 2 answers
A.The output of 'kubectl describe pod <driver-pod-name>'.
B.The contents of the /etc/kubernetes/manifests folder.
C.The logs of the driver pod using 'kubectl logs'.
D.The system-wide /var/log/syslog file on the control plane.
E.The NVIDIA license server status page.
AnswersA, C

Describing the pod reveals critical event information such as scheduling errors, image pull failures, or readiness probe failures. This is the first step in diagnosing why a pod failed to reach a 'Running' state, providing clues about potential resource constraints or registry authentication issues preventing deployment.

Why this answer

Checking the pod's logs and describe output is the standard diagnostic path. The 'describe' output shows events like ImagePullBackOff or scheduling failures, while the logs provide the specific error message from the driver installation script itself, such as kernel header mismatches or network timeouts. These two sources provide the necessary visibility to pinpoint whether the failure is infrastructure-related, configuration-based, or due to a missing environmental dependency.

Exam trap

Candidates frequently choose 'kubectl get events' or 'dmesg' on the host. While helpful, the question specifically asks for the two standard Kubernetes-native diagnostic locations for a failing pod.

11
MCQmedium

Which configuration file is typically modified to enable the NVIDIA Device Plugin in a Kubernetes cluster?

A./etc/docker/daemon.json
B.A Kubernetes DaemonSet manifest (YAML).
C./etc/nvidia/nvidia-container-runtime.json
D./etc/kubernetes/kubelet.conf
AnswerB

The NVIDIA Device Plugin is deployed as a DaemonSet in Kubernetes. The configuration for the plugin, including resource settings and feature flags, is defined within the manifest YAML file, which is then applied via kubectl to the cluster to manage the discovery and allocation of GPUs.

Why this answer

The NVIDIA Device Plugin is typically deployed as a DaemonSet using a YAML manifest. Modifying this manifest allows administrators to customize settings like the MIG strategy, time-slicing, and resource allocation policies. Correct configuration of this file is essential for ensuring that the Kubernetes scheduler correctly recognizes and assigns GPU resources to pods, which is the cornerstone of effective AI cluster management in modern cloud-native environments.

Exam trap

Candidates often confuse the NVIDIA Device Plugin with the NVIDIA Container Runtime, mistakenly believing configuration changes happen in a Docker daemon file instead of a Kubernetes DaemonSet manifest.

12
MCQhard

An AI infrastructure team is preparing an air-gapped data center to install the NVIDIA GPU Operator. They have mirrored all required container images into a private registry. During installation, the Operator's pods fail with ImagePullBackOff because the components still reference images under nvcr.io. Which configuration is required to make the Operator and its managed components pull from the private registry?

A.Enable the Operator's auto-mirroring feature so it copies images from nvcr.io at runtime.
B.Set the chart values for the image repository and tag for each component, and provide an image pull secret in the Operator namespace.
C.Configure a DNS override so nvcr.io resolves to the private registry's IP address.
D.Add nvcr.io to the cluster's imagePullPolicy as Never.
AnswerB

In an air-gapped installation, every component image reference must be overridden to point at the private registry, including the Operator itself and the driver, toolkit, device plugin, and DCGM exporter images. The pull secret must exist in the namespace so kubelet can authenticate. Without these overrides the manifests still name nvcr.io and pulls fail.

Why this answer

Air-gapped installs require pre-mirroring every image and rewriting all image references in the Helm values to the private registry, including the Operator and each managed component. Authentication to the private registry is supplied through an image pull secret placed in the Operator namespace so the kubelet can pull. DNS tricks, pull policy changes, or imagined auto-mirroring do not redirect image references and leave the components failing.

Exam trap

The trap here is believing that a network-level redirection such as DNS or a pull policy change can substitute for rewriting the actual image repository references and supplying registry credentials.

13
MCQeasy

A platform engineer is preparing a bare-metal Ubuntu 22.04 server with four A100 GPUs for an NVIDIA AI Enterprise deployment. The GPUs are not yet visible to the operating system tooling. Which command should the engineer run to confirm the driver loaded successfully and that all four GPUs are enumerated with their current driver version?

A.dcgmi discovery -l
B.nvidia-container-cli --info
C.nvidia-smi
D.lspci -d 10de:
AnswerC

nvidia-smi queries the loaded NVIDIA kernel driver and prints every enumerated GPU together with the driver version and CUDA version. On a freshly prepared AI Enterprise node it is the canonical first check that the driver bound to all four A100 devices before any container runtime or GPU Operator work begins.

Why this answer

The NVIDIA kernel driver is the foundation of every AI Enterprise workload, and nvidia-smi is the standard utility that both triggers a driver query and renders the full GPU inventory with driver and CUDA versions. Administrative toolkits such as DCGM or the container toolkit sit above the driver and only report meaningful data once the driver has bound successfully to all adapters.

Exam trap

The trap here is assuming that seeing the GPUs in a PCI listing proves the driver is installed, when PCI enumeration and driver binding are independent states.

14
MCQhard

A financial services company is deploying NVIDIA AI Enterprise on a VMware vSphere cluster with NVIDIA A100 GPUs. The security team requires that GPU workloads be isolated at the hardware level, with separate memory and fault domains, to meet regulatory compliance. The company also wants to maximize GPU utilization by running multiple workloads concurrently. Which NVIDIA feature should be enabled to meet these requirements?

A.NVIDIA NVLink with SHARP
B.NVIDIA vGPU with time-sliced scheduling
C.NVIDIA GPUDirect Storage
D.NVIDIA Multi-Instance GPU (MIG)
AnswerD

MIG partitions an A100 GPU into up to seven independent instances, each with dedicated memory, cache, and compute resources. This provides hardware-level isolation and fault domain separation, satisfying regulatory requirements for workload isolation. It also allows multiple workloads to run concurrently on a single GPU, maximizing utilization. MIG is the correct choice for this scenario because it uniquely combines isolation and concurrency on A100 hardware.

Why this answer

NVIDIA Multi-Instance GPU (MIG) is the only feature that provides hardware-level partitioning of an A100 GPU into isolated instances with dedicated memory, cache, and compute resources. This meets the need for fault domain separation and regulatory compliance while enabling concurrent execution of multiple workloads, thereby maximizing utilization. Other options either share resources without isolation or address different concerns like data transfer or inter-GPU communication.

Exam trap

The trap here is confusing time-sliced vGPU or other GPU sharing technologies with true hardware partitioning, which only MIG provides on A100 GPUs.

15
MCQeasy

Which NVIDIA tool allows you to verify that the GPU and its driver are properly installed and functioning on a Linux system?

A.nvcc --version
B.nvidia-smi
C.lspci | grep nvidia
D.docker run --gpus all
AnswerB

This command is the primary tool for verifying that the NVIDIA driver is loaded and communicating correctly with the GPU. It provides essential diagnostic information, including device names, driver versions, and current memory usage, which are the fundamental metrics for confirming that a GPU installation was successful.

Why this answer

The 'nvidia-smi' (System Management Interface) tool is the standard utility for interacting with the NVIDIA driver. It provides a real-time status of GPU utilization, temperature, memory usage, and driver versions. Being proficient with this tool is essential for an AI Ops professional to quickly validate hardware health, confirm driver installation, and identify if a GPU is accessible by the host OS after a fresh installation or reboot.

Exam trap

Candidates often confuse nvidia-smi with higher-level management tools like NVIDIA AI Enterprise or DCGM. They overlook that nvidia-smi is the fundamental, low-level command for basic driver and hardware verification.

16
MCQmedium

Which file format is commonly used to define the configuration and state for the NVIDIA GPU Operator within a Kubernetes environment?

A.JSON
B.YAML
C.TOML
D.XML
AnswerB

YAML is the native format for Kubernetes configurations. The GPU Operator relies on YAML files to define the ClusterPolicy and other custom resources. This structure allows administrators to define the desired state of their GPU infrastructure, enabling automated reconciliation and lifecycle management by the operator.

Why this answer

Kubernetes uses YAML to define declarative states for resources, including Custom Resource Definitions (CRDs) used by operators. The GPU Operator leverages YAML manifests to configure settings such as driver version, toolkit installation, and monitoring agents. Understanding this format is vital for AI Ops professionals, as it allows for version-controlled infrastructure-as-code practices, ensuring deployments are reproducible and auditable across various environments.

Exam trap

Candidates might assume binary or proprietary configuration formats are used, forgetting that Kubernetes operators rely heavily on standard declarative YAML files for managing Custom Resource Definitions.

17
MCQmedium

When deploying NVIDIA AI Enterprise, why is the use of the NVIDIA NGC Catalog recommended over public container repositories?

A.NGC images are the only way to bypass the need for a valid NVIDIA license.
B.NGC images are pre-configured to automatically perform hardware diagnostics.
C.NGC images are curated and optimized for NVIDIA hardware compatibility.
D.NGC images can be deployed without any need for the NVIDIA Container Toolkit.
AnswerC

NGC images are specifically curated to ensure that all required CUDA libraries and drivers are compatible. This optimization guarantees that the software stack works as intended with NVIDIA GPUs, preventing the 'DLL hell' or library version conflicts often encountered when assembling container environments from generic public images and manual installs.

Why this answer

The NGC Catalog provides verified, performance-optimized, and security-scanned container images specifically tailored for the NVIDIA AI Enterprise stack. These images contain all necessary dependencies and are guaranteed to be compatible with supported GPU drivers. Using these images reduces the risk of deployment failures caused by library version mismatches, missing dependencies, or unoptimized software versions that are common in generic, public repositories, thereby ensuring a reliable production-grade AI environment.

Exam trap

Candidates mistakenly believe that public repositories contain identical images to NGC. They ignore the performance tuning and security scanning inherent in curated NGC images that prevent runtime bottlenecks.

18
MCQmedium

Which tool is the industry standard for monitoring and managing NVIDIA data center GPUs in a large-scale cluster deployment?

A.nvidia-smi
B.NVIDIA DCGM
C.CUDA Debugger
D.NVIDIA Triton Inference Server
AnswerB

DCGM is the enterprise-grade solution for managing and monitoring GPUs. It supports cluster-wide health monitoring, performance profiling, and configurable diagnostic tests. It is the core component for NVIDIA AI operations, providing the telemetry needed for dashboarding and automated resource management at scale.

Why this answer

The NVIDIA Data Center GPU Manager (DCGM) provides a comprehensive set of APIs and tools for managing and monitoring GPUs in production environments. It is essential for AI operations because it allows for real-time health checks, diagnostic testing, and policy-based management of GPU resources across a distributed cluster, ensuring high availability and identifying performance bottlenecks before they impact critical training or inference workflows.

Exam trap

Candidates often select generic monitoring tools like Prometheus or Grafana. While these visualize data, they are not the primary NVIDIA-specific engine used to manage and diagnose GPU hardware health.

19
MCQmedium

Which action is required when updating the NVIDIA driver on a node managed by the GPU Operator to ensure that running workloads are not interrupted abruptly?

A.Manually stop all running pods.
B.Use the operator to perform a rolling update.
C.Reboot the entire cluster simultaneously.
D.Delete the node object from the cluster.
AnswerB

The GPU Operator automates the rolling update process, which includes cordoning and draining nodes to move workloads before applying driver upgrades. This ensures that the maintenance happens without unexpected service outages, fulfilling the requirement to manage infrastructure updates safely while maintaining high availability for the dependent AI workloads.

Why this answer

The GPU Operator supports seamless driver upgrades by cordoning and draining nodes. When an update is initiated, the operator gracefully moves workloads to other available nodes before updating the driver. This process prevents application crashes, ensures data integrity, and maintains cluster stability during maintenance windows, which is a key responsibility for AI operations professionals managing production-grade, long-running AI training or inference tasks on shared GPU resources.

Exam trap

Candidates often suggest manually stopping pods or deleting the node, forgetting that the GPU Operator provides automated rolling update capabilities that handle pod draining and cordoning safely.

20
Multi-Selecthard

An AI operations team is validating a new Kubernetes cluster before installing the NVIDIA GPU Operator with the driver managed by the Operator itself. The nodes run a supported Linux distribution with GPUs physically installed. Which two conditions must be satisfied for the Operator's driver container to build and load the kernel module successfully? (Choose two.)

Select 2 answers
A.Each node must have at least one Multi-Instance GPU profile pre-created.
B.The kubelet must be configured with the NVIDIA device plugin endpoint flag.
C.Secure Boot must be disabled or the module must be signed with an enrolled key.
D.The cluster must have the NVIDIA Network Operator installed first.
E.The node's kernel headers for the running kernel must be available to the driver container.
AnswersC, E

When UEFI Secure Boot is enabled, the kernel refuses to load unsigned modules. The NVIDIA driver container either needs Secure Boot turned off or must sign the generated module with a Machine Owner Key that is enrolled in the firmware. Without one of these, module insertion is rejected and the GPU stays unusable even though the build completed.

Why this answer

For the Operator to manage drivers, the driver container must compile the NVIDIA kernel module against the exact running kernel, which requires matching kernel headers and build dependencies on the node. In addition, UEFI Secure Boot blocks unsigned modules, so it must be disabled or the module signed with an enrolled key. These two conditions gate whether the module can be built and inserted; without them the node cannot advertise GPU capacity to the scheduler.

Exam trap

The trap here is treating unrelated cluster add-ons, such as the Network Operator or MIG profiles, as prerequisites for the driver container when the real gating factors are kernel headers and module signing.

21
MCQeasy

What is the primary role of the NVIDIA Data Center GPU Manager (DCGM) Exporter in a cloud-native monitoring stack?

A.To act as a load balancer for GPU-accelerated traffic.
B.To provide a Prometheus-compatible endpoint for GPU metrics.
C.To manage the deployment of containerized AI models.
D.To encrypt data transmission between the GPU and the CPU.
AnswerB

The DCGM Exporter collects raw telemetry data from the DCGM API and exposes it as a scrapeable endpoint for Prometheus. This enables administrators to visualize GPU health, track usage trends, and set alerts for thresholds like high temperature or memory usage within a Grafana dashboard.

Why this answer

The DCGM Exporter is the bridge between NVIDIA hardware metrics and monitoring systems like Prometheus. It gathers real-time telemetry from DCGM and exposes it in a format that Prometheus can scrape. This integration is vital for AI operations because it provides observability into GPU utilization, memory usage, and thermal health, enabling automated alerts and performance dashboards that are essential for maintaining a production-grade AI cluster.

Exam trap

Candidates confuse the DCGM Exporter with the underlying metrics collector tool itself, missing its specific role in formatting data for Prometheus scraping.

22
MCQmedium

An AI operations engineer is preparing a Kubernetes cluster to run GPU-accelerated workloads using the NVIDIA GPU Operator. The cluster nodes already have NVIDIA data center GPUs installed and the NVIDIA driver is pre-installed on the host. The engineer wants to use the GPU Operator to manage the container toolkit, device plugin, and monitoring components but must avoid the Operator managing or upgrading the driver. Which configuration should be applied to the GPU Operator deployment?

A.Use the NVIDIA GPU Operator with the --set toolkit.enabled=false option to prevent driver installation.
B.Set the driver.enabled parameter to false in the GPU Operator's Helm chart values.
C.Install the GPU Operator with the --set operator.driverVersion=latest flag to pin the driver version.
D.Deploy the GPU Operator but remove the nvidia-driver-daemonset after installation.
AnswerB

Setting driver.enabled=false instructs the GPU Operator to skip deploying the driver container and instead rely on the pre-installed host driver. This is the supported method for clusters where the driver is managed externally, such as via the node's package manager. The Operator will still deploy the container toolkit, device plugin, and DCGM exporter, allowing full GPU scheduling and monitoring without touching the driver.

Why this answer

When the NVIDIA driver is already present on host nodes and should remain externally managed, the GPU Operator must be configured to not deploy its own driver container. The driver.enabled=false Helm value achieves this by skipping the driver daemonset while still deploying other components like the container toolkit, device plugin, and DCGM exporter. This preserves the existing driver and avoids conflicts.

Exam trap

The trap here is assuming that specifying a driver version or disabling the toolkit will prevent driver installation, when only the driver.enabled flag controls driver management.

23
MCQeasy

Which command is used to verify that the NVIDIA GPU Operator has successfully installed the necessary components on a Kubernetes node?

A.nvidia-smi check-components
B.kubectl get pods -n gpu-operator
C.docker inspect nvidia-gpu-operator
D.kube-config verify --gpu
AnswerB

The GPU Operator runs in a specific namespace. Listing the pods in this namespace allows an administrator to see the status of the daemonsets, such as the driver installer and device plugin. All pods being in a 'Running' or 'Completed' state confirms a successful deployment of the operator components.

Why this answer

The 'kubectl get pods -n gpu-operator' command is the primary method to check the status of the operator and its managed components, such as the device plugin and driver daemonsets. This step is essential because it confirms that the operator's control loop has successfully completed the deployment, ensuring that the GPU software stack is active and ready to handle incoming AI workload scheduling requests.

Exam trap

Candidates often confuse cluster-wide resource inspection commands like 'kubectl get nodes' or generic pod queries with operator-specific status checks, forgetting to target the dedicated namespace where the GPU Operator components reside.

24
MCQmedium

A Kubernetes cluster administrator is installing the NVIDIA GPU Operator and wants to ensure that GPU workloads are scheduled only on nodes with healthy GPUs. The administrator plans to use the operator's built-in health checks. Which component is responsible for monitoring GPU health and marking nodes as unschedulable when a GPU fails?

A.NVIDIA DCGM Exporter
B.NVIDIA GPU Feature Discovery
C.NVIDIA Device Plugin
D.NVIDIA GPU Operator's health check component (part of the operator's node validation)
AnswerD

The GPU Operator includes a health check mechanism that periodically validates GPU health. When a GPU is found unhealthy, the operator applies a taint to the node, preventing new GPU workloads from being scheduled there. This integrates with Kubernetes scheduling to avoid placing workloads on failing hardware.

Why this answer

The GPU Operator's health check component continuously monitors GPU status and taints nodes with unhealthy GPUs. This prevents Kubernetes from scheduling new GPU workloads on those nodes, improving reliability. The device plugin handles resource advertisement, while DCGM Exporter provides metrics; neither automatically taints nodes based on health.

Exam trap

The trap here is attributing health-based node tainting to the device plugin or DCGM Exporter, when it is actually the operator's dedicated health check component that performs this action.

25
MCQmedium

Which of the following describes the purpose of the NVIDIA GPU Operator's 'Driver Container'?

A.To store the persistent state of trained neural networks.
B.To provide a platform-agnostic way to deploy NVIDIA drivers.
C.To manage the licensing of the AI Enterprise suite.
D.To act as a gateway for remote GPU access.
AnswerB

The Driver Container encapsulates the driver installation logic, making it consistent across different node operating systems. It handles the nuances of kernel headers and source code compilation, allowing the GPU Operator to manage drivers as standard Kubernetes workloads rather than requiring manual installation on each node.

Why this answer

The Driver Container is a critical component that builds or pulls the correct driver for the specific host OS and kernel. It automates the complex process of driver installation, ensuring compatibility across heterogeneous node environments. By containerizing the driver, the Operator simplifies maintenance and upgrades, reducing the risk of configuration drift and ensuring that nodes always have a functional driver compatible with the latest AI software releases.

Exam trap

Candidates often think the driver container installs the driver directly onto the host OS. In reality, it packages the driver to be portable and compatible across different kernel versions without manual compilation.

26
MCQeasy

When installing NVIDIA drivers via a package manager, what is the importance of the 'dkms' package?

A.It provides the graphical user interface for managing GPU settings.
B.It ensures the driver module is rebuilt automatically during kernel upgrades.
C.It improves the performance of the CUDA compiler during the build process.
D.It manages the firmware updates for the GPU hardware.
AnswerB

DKMS is designed to maintain kernel modules across kernel updates. By automatically recompiling the NVIDIA kernel module against the new kernel headers, it ensures that the GPU driver remains compatible and loaded correctly after the OS performs a system update or upgrade.

Why this answer

DKMS (Dynamic Kernel Module Support) is critical because it automates the recompilation of the NVIDIA kernel module whenever a new Linux kernel is installed. In AI production environments, kernel updates are common for security and stability. DKMS ensures that the GPU remains functional after an update without requiring manual driver intervention, preventing unplanned downtime for AI workloads and simplifying system administration tasks.

Exam trap

Candidates often confuse DKMS with standard package managers or assume manual recompilation is sufficient, forgetting that system updates happen automatically and break modules without DKMS.

27
MCQeasy

When installing the NVIDIA GPU Operator, which namespace is typically used to ensure proper isolation and role-based access control?

A.default
B.gpu-operator
C.kube-system
D.public-apps
AnswerB

Creating a dedicated namespace for the GPU Operator is standard industry practice. It provides logical isolation and allows administrators to apply specific RBAC policies to the operator's components, ensuring that the critical system-level software is segregated from general tenant workloads and other cluster services for better security posture.

Why this answer

Using a dedicated namespace like 'gpu-operator' is a best practice in Kubernetes. It isolates the operator's resources, permissions, and lifecycle from other cluster services. This separation allows for granular security policies, making it easier to manage access and ensuring that only authorized personnel can modify the operator's configuration, which is vital for maintaining the security and integrity of the GPU-enabled infrastructure in a production environment.

Exam trap

Candidates often install the Operator in the default namespace. This creates security risks and makes it difficult to apply specific RBAC policies or manage the lifecycle of the operator independently of applications.

28
MCQmedium

Which action must be performed after updating the NVIDIA driver on a Linux host to ensure that all active GPU containers recognize the new driver version?

A.Rebuild all container images with the latest CUDA toolkit.
B.Restart the containerized applications.
C.Upgrade the Docker daemon version.
D.Re-run the nvidia-container-toolkit installation script.
AnswerB

Active containers hold references to the old driver files that were mapped during their initialization. To ensure that the containers are using the updated driver libraries from the host, the processes must be terminated and relaunched so the container runtime can remount the latest versions.

Why this answer

Containers utilize the host-level NVIDIA drivers injected via the container toolkit. When the host driver is updated, active containers will still be referencing the old driver libraries in their environment until they are restarted. Restarting the containers is the mandatory step to force them to bind against the newly installed libraries, ensuring compatibility and preventing runtime errors associated with outdated driver interfaces.

Exam trap

Candidates assume updating the host driver is instantly inherited by running containers, forgetting that active containers retain bindings to old libraries until restarted.

29
MCQeasy

A healthcare company is deploying NVIDIA AI Enterprise on a Kubernetes cluster to run medical imaging AI models. The cluster administrator needs to verify that the NVIDIA GPU Operator is installed and functioning correctly. Which command should the administrator use to check the status of the GPU Operator pods?

A.kubectl describe daemonset nvidia-device-plugin -n kube-system
B.kubectl logs -n gpu-operator nvidia-driver-daemonset
C.kubectl get nodes -o wide
D.kubectl get pods -n gpu-operator
AnswerD

The GPU Operator is typically deployed in the gpu-operator namespace. Running kubectl get pods -n gpu-operator lists all pods managed by the Operator, allowing the administrator to verify that components like the driver, container toolkit, and device plugin are running.

Why this answer

The NVIDIA GPU Operator deploys its components into the gpu-operator namespace. To verify that the Operator is installed and functioning, the administrator should list the pods in that namespace. This provides a quick overview of all related pods, such as the driver, container toolkit, and device plugin, and their current status.

Exam trap

The trap here is assuming that GPU Operator components reside in kube-system or that checking nodes alone is sufficient; the Operator uses its own namespace, and pod status is the definitive check.

30
Multi-Selecthard

A platform engineer must validate a new NVIDIA GPU Operator deployment on a Kubernetes cluster before handing it to data scientists. Which two checks confirm that the Operator has correctly exposed GPU resources to the cluster scheduler? (Choose two.)

Select 2 answers
A.Confirm that the GPU Operator's driver DaemonSet pods are scheduled on all nodes, including CPU-only nodes.
B.Run a CUDA-enabled pod that requests nvidia.com/gpu and verify it reaches Running state and reports the expected device.
C.Confirm that nodes advertise the nvidia.com/gpu resource and that its allocatable count matches the physical GPU count.
D.Check that the NVIDIA driver container image tag matches the CUDA toolkit version installed in the workload image.
E.Verify that the container runtime on each node has been switched from containerd to Docker with the nvidia runtime as default.
AnswersB, C

A scheduled pod that requests the GPU resource exercises the full path: scheduler admission, device plugin allocation, container runtime injection, and driver access inside the container. If the pod runs and nvidia-smi or a CUDA sample reports the device, the end-to-end GPU enablement is verified, not just the resource advertisement.

Why this answer

GPU exposure is confirmed by two complementary signals: the node advertises the nvidia.com/gpu extended resource with an allocatable count matching the physical GPUs, and a test pod requesting that resource schedules successfully and can access the device. Together they validate device plugin registration and end-to-end runtime injection. Driver image tags, node runtime swaps, and DaemonSet placement on CPU-only nodes are irrelevant to this validation.

Exam trap

The trap here is treating driver or runtime version matching as proof of GPU exposure instead of checking the advertised extended resource and an actual scheduled GPU pod.

31
Multi-Selecthard

When deploying the NVIDIA GPU Operator in a restricted-access environment (air-gapped), which THREE requirements must be addressed to ensure a successful installation?

Select 3 answers
A.Mirror all required images to a private registry.
B.Upgrade the host BIOS to the latest version.
C.Configure the operator to use the private registry.
D.Provide local access to kernel headers.
E.Install the NVIDIA Triton Inference Server first.
AnswersA, C, D

Without internet access, the operator cannot fetch images from public sources like NGC. Mirroring is mandatory to ensure the container runtime can pull the images locally. The operator deployment manifest must also be updated to reference the internal registry URL instead of the default public registry paths.

Why this answer

In air-gapped environments, the inability to reach public registries is the primary failure point. Administrators must mirror all required container images to a local private registry, configure the operator to point to these local locations, and ensure that all necessary kernel headers are available locally for the driver build process. These steps ensure that the GPU Operator can complete its setup without external internet dependencies.

Exam trap

Candidates often forget the requirement for local kernel headers. Even with mirrored images, the driver build process will fail if it cannot access the necessary kernel headers locally.

32
MCQmedium

Which NVIDIA technology enables the partitioning of a single physical GPU into multiple independent instances for use by different virtual machines or containers?

A.NVIDIA GPUDirect Storage.
B.NVIDIA Multi-Instance GPU (MIG).
C.NVIDIA NVLink.
D.NVIDIA vGPU Profiles.
AnswerB

MIG enables hardware-level partitioning of the GPU, allowing each instance to have its own compute cores, memory, and cache. This provides robust isolation and performance guarantees for multiple applications or users, ensuring that one workload does not adversely impact the performance of another co-located on the same device.

Why this answer

NVIDIA Multi-Instance GPU (MIG) technology is the correct answer. It allows a single GPU to be securely partitioned at the hardware level, providing guaranteed QoS and isolation for different workloads. This is essential for maximizing GPU utilization in enterprise AI, as it enables the co-location of small inference tasks alongside larger training workloads on a single piece of high-end hardware.

Exam trap

Candidates often confuse MIG with vGPU or Time-Slicing, failing to specify that MIG is the unique hardware-level partitioning technology for NVIDIA GPUs.

33
MCQeasy

Which NVIDIA software component is responsible for providing the necessary CUDA libraries to containerized applications?

A.NVIDIA Driver
B.NVIDIA Container Toolkit
C.NVIDIA vGPU Manager
D.NVIDIA Triton
AnswerB

The NVIDIA Container Toolkit provides the runtime hooks and libraries that allow containers to access the GPU and CUDA acceleration. It enables the container engine to identify, map, and utilize the host's NVIDIA hardware, ensuring that deep learning frameworks can call CUDA primitives directly from within the container.

Why this answer

The NVIDIA Container Toolkit is the critical bridge. It provides the necessary libraries and container runtime hooks that allow processes inside a container to access the host's GPU and CUDA environment. This setup is fundamental for AI Enterprise, as it enables portability of AI applications while maintaining high-performance access to physical GPU hardware, regardless of the underlying host OS distribution or container runtime used.

Exam trap

Candidates often confuse the Container Toolkit with the NVIDIA driver itself, failing to recognize that the Toolkit provides the bridge for containers to consume host-side drivers.

34
MCQmedium

An AI engineer is deploying a large language model on an NVIDIA DGX system. The deployment fails with an error indicating an insufficient NVIDIA driver version for the required CUDA toolkit. Which action should the engineer take to resolve the dependency mismatch?

A.Reinstall the CUDA toolkit using a generic installer without checking the NVIDIA driver version compatibility matrix.
B.Downgrade the OS kernel to a legacy version to force compatibility with an older CUDA toolkit.
C.Upgrade the NVIDIA driver to a version verified as compatible with the required CUDA toolkit.
D.Modify the LD_LIBRARY_PATH environment variable to prioritize older CUDA libraries found on the system.
AnswerC

Matching the driver version to the CUDA toolkit requirements is the standard procedure for fixing driver-level mismatches. This ensures that the GPU hardware can successfully communicate with the user-space libraries, allowing the AI software stack to utilize the full range of CUDA capabilities.

Why this answer

Drivers and CUDA versions maintain a strict compatibility matrix. Installing the latest driver version supported by the specific DGX OS is essential to ensure the CUDA runtime can interface correctly with the GPU hardware. This task is critical in production environments because mismatched drivers lead to kernel panics or silent performance degradation, directly impacting the availability of AI workloads running on the cluster infrastructure.

Exam trap

Candidates often suggest reinstalling the CUDA toolkit or the entire container runtime. The root issue is the driver-to-CUDA compatibility matrix, which requires updating the driver to match the toolkit.

35
MCQhard

Which THREE of the following are benefits of using Multi-Instance GPU (MIG) technology in a Kubernetes environment?

A.Hardware-level isolation between GPU workloads.
B.Increased total GPU count available to the OS kernel.
C.Deterministic Quality of Service (QoS) for different workloads.
D.Automatic translation of CUDA code to work on different GPU architectures.
E.Optimized utilization of expensive GPU hardware resources.
AnswerA, C, E

MIG provides true hardware-level partitioning, which prevents processes in one instance from accessing the memory or compute resources of another instance. This is far more robust than software-based time-slicing, ensuring that different tenants or workloads remain entirely isolated from one another.

Why this answer

MIG allows for hardware-level isolation, ensuring that one workload's memory usage or compute spikes do not impact others on the same physical GPU. This provides deterministic Quality of Service (QoS), which is critical for multi-tenant AI environments. By partitioning resources, administrators can optimize GPU allocation, allowing smaller, less intensive tasks to run on a fraction of the GPU, thereby significantly increasing overall cluster efficiency.

Exam trap

Candidates often mistake MIG for a software-based virtualization tool rather than a hardware-level partitioning technology, leading them to select incorrect benefits related to dynamic software-defined resource sharing instead of isolation.

36
MCQhard

An AI operations team is deploying NVIDIA AI Enterprise on a bare-metal Kubernetes cluster with DGX A100 systems. They need to enable GPUDirect Storage to accelerate data loading from a local NVMe array. Which component must be installed and configured on the DGX nodes to support GPUDirect Storage?

A.NVIDIA Peer-to-Peer (P2P) over PCIe with IOMMU disabled
B.NVIDIA GPU Operator with the RDMA shared device plugin enabled
C.NVIDIA Container Toolkit with the 'nvidia-container-runtime' configured for privileged mode
D.NVIDIA Magnum IO GPUDirect Storage kernel module and user-space libraries
AnswerD

GPUDirect Storage requires the nvidia-fs kernel module and CUDA libraries that enable direct memory access between storage and GPU memory. These are part of Magnum IO GPUDirect Storage. Installing and configuring them on the DGX nodes allows applications to bypass the CPU and system memory, reducing latency and increasing throughput for data-intensive AI workloads.

Why this answer

GPUDirect Storage is enabled by installing the NVIDIA Magnum IO GPUDirect Storage components, which include the nvidia-fs kernel module and CUDA libraries. These allow direct DMA transfers between NVMe storage and GPU memory, bypassing the CPU. On DGX systems, these components are typically part of the DGX software stack and must be properly configured.

Exam trap

The trap here is confusing GPUDirect Storage with other NVIDIA technologies like RDMA or P2P, which address different data paths.

37
MCQmedium

Which mechanism does the NVIDIA GPU Operator use to ensure that the driver installed on a worker node matches the specific architecture of the installed GPU hardware?

A.NVIDIA System Management Interface (nvidia-smi).
B.Node Feature Discovery (NFD).
C.The container runtime's default configuration.
D.Hardcoding the driver version in the deployment manifest.
AnswerB

NFD labels nodes based on hardware attributes like GPU model and architecture. The GPU Operator uses these labels to match the node to the appropriate driver image. This automated mechanism is essential for scaling deployments across diverse hardware configurations without requiring manual node configuration for every single machine.

Why this answer

The GPU Operator uses Node Feature Discovery (NFD) to identify the specific GPU hardware present on each node. By labeling nodes with these hardware characteristics, the operator can ensure that the correct driver and kernel modules are selected and deployed. This mapping is vital in heterogeneous clusters, where different nodes might require different driver builds, preventing installation errors and ensuring optimal performance across the entire fleet.

Exam trap

Candidates often guess that the GPU Operator performs the hardware detection itself. It actually relies on the Node Feature Discovery (NFD) to identify and label hardware for the operator.

38
MCQmedium

An administrator is deploying NVIDIA AI Enterprise on a bare-metal cluster. Which component is mandatory to provide the necessary abstraction layer for containerized GPU resources?

A.NVIDIA Triton Inference Server
B.NVIDIA GPU Operator
C.NVIDIA Base Command
D.NVIDIA CUDA Toolkit
AnswerB

The GPU Operator utilizes the Kubernetes operator pattern to automate the installation and management of NVIDIA drivers, the NVIDIA Container Toolkit, and device plugins. This provides the critical abstraction layer required to expose physical GPUs as schedulable resources within a container orchestrator environment seamlessly.

Why this answer

The NVIDIA GPU Operator is essential for automating the management of all NVIDIA software components needed to provision Kubernetes with GPU support. It manages the lifecycle of drivers, container runtimes, and monitoring tools. By leveraging the Operator, administrators ensure consistent configuration across the cluster, preventing drift and ensuring that CUDA workloads have the correct dependencies to execute reliably on bare-metal infrastructure.

Exam trap

Candidates often name individual components like the driver or toolkit. The GPU Operator is the mandatory overarching framework that automates the deployment and management of these individual components.

39
MCQeasy

A system administrator is installing the NVIDIA Container Toolkit on a standalone Ubuntu server to run GPU-accelerated containers. After installation, they want to verify that the toolkit is correctly configured. Which command should they run to test GPU access from a container?

A.docker run --rm --gpus all nvidia/cuda:11.0-base nvidia-smi
B.nvidia-smi
C.systemctl status nvidia-container-runtime
D.nvidia-container-cli --version
AnswerA

This command runs a CUDA container with all GPUs exposed and executes nvidia-smi inside the container. If the toolkit is configured correctly, the container will have access to the GPUs and nvidia-smi will display them. This directly tests the container runtime's ability to pass through GPU devices. It is the standard method to verify NVIDIA Container Toolkit functionality.

Why this answer

To verify that the NVIDIA Container Toolkit is correctly configured, the administrator should run a container with GPU access and execute nvidia-smi inside it. The command docker run --rm --gpus all nvidia/cuda:11.0-base nvidia-smi uses the --gpus all flag to expose all GPUs, and the nvidia-smi output confirms that the container can see and use the GPUs. This tests the entire stack from Docker to the toolkit.

Other commands only check installation or host-level GPU status.

Exam trap

The trap here is assuming that running nvidia-smi on the host or checking the toolkit version is sufficient to verify container GPU access, when a container-based test is required.

40
MCQmedium

During the deployment of an AI model training workload on a multi-node cluster, the administrator notices that inter-node communication is significantly slower than expected. Which deployment aspect should be investigated first?

A.Verify that the GPU memory usage is below 50% on all nodes.
B.Check the NCCL_DEBUG and network interface configuration for the training job.
C.Restart the Kubernetes API server to refresh the node connection states.
D.Update the NVIDIA driver to the latest gaming-optimized release.
AnswerB

NCCL communication relies heavily on the correct identification of high-speed network interfaces. If the job is defaulting to an Ethernet interface instead of InfiniBand, training will be severely throttled. Setting NCCL_DEBUG allows administrators to identify which interfaces are being selected and if the desired fabric is actually being used.

Why this answer

NCCL (NVIDIA Collective Communications Library) is the primary engine for inter-node communication in distributed training. Misconfiguration of the network interface or the underlying fabric provider (e.g., InfiniBand or RoCE) will lead to significant performance bottlenecks. Investigating the NCCL configuration and the network topology ensures that the GPUs are utilizing the highest bandwidth available, which is vital for preventing training jobs from stalling during gradient synchronization across multiple nodes.

Exam trap

Candidates often try to troubleshoot the model code or the application logic first, ignoring the communication layer (NCCL) and network fabric which are the primary culprits for inter-node training latency.

41
MCQmedium

A platform engineer is preparing a bare-metal Kubernetes cluster to run GPU-accelerated AI workloads using the NVIDIA GPU Operator. The cluster nodes have NVIDIA Ampere GPUs and run Ubuntu 22.04 with containerd as the container runtime. The engineer wants to avoid installing any NVIDIA drivers or CUDA components directly on the host. Which GPU Operator configuration should be used to achieve this?

A.Set driver.enabled=false and rely on pre-installed host drivers.
B.Deploy the GPU Operator with the default configuration, which includes the driver container.
C.Use the operator's 'driver' Helm chart with 'driver.enabled=true' but set 'driver.usePrecompiled=true'.
D.Install the NVIDIA Container Toolkit manually and disable the operator's driver management.
AnswerB

The default GPU Operator deployment includes a driver container that compiles and loads the NVIDIA kernel modules on the host without requiring a pre-installed driver. This satisfies the requirement to avoid manual host driver installation while still enabling GPU access for workloads, as the Operator manages the driver lifecycle automatically.

Why this answer

The NVIDIA GPU Operator is designed to manage the full stack of GPU software, including the driver, by running a driver container on each node. In a default deployment, the driver container builds and loads the kernel modules, so no manual host driver installation is needed. This aligns with the requirement to keep the host clean and let the Operator handle everything.

Exam trap

The trap here is assuming that the GPU Operator always requires pre-installed host drivers, when in fact its default mode deploys a driver container to manage them.

42
MCQmedium

A financial services company is deploying NVIDIA AI Enterprise in an air-gapped data center. They need to install the NVIDIA GPU Operator on their Kubernetes cluster without internet access. Which additional step must they take to ensure a successful installation?

A.Disable the GPU Operator's driver container and manually install the driver on each node.
B.Use a Kubernetes cluster that supports dynamic volume provisioning for GPU drivers.
C.Enable the GPU Operator's built-in proxy to fetch images from the internet.
D.Configure the GPU Operator to use a private container registry with mirrored images.
AnswerD

In an air-gapped environment, the GPU Operator cannot pull images from public registries. The administrator must mirror all required images, including the operator, driver, device plugin, and container toolkit, to a private registry accessible within the network. The GPU Operator must be configured to pull from this registry by setting the appropriate values in the Helm chart. This step is essential for successful installation.

Why this answer

For an air-gapped installation, the critical step is to make all required container images available within the isolated network. This is done by mirroring the images to a private registry and configuring the GPU Operator to use that registry. The operator's Helm chart provides values to specify the registry and image repository.

Without this, the operator cannot pull the necessary components. Other options either do not address the air-gap limitation or propose unnecessary manual steps.

Exam trap

The trap here is thinking that the GPU Operator has a built-in mechanism to fetch images without internet, or that manual driver installation is required, when the actual solution is to use a private registry with mirrored images.

43
MCQhard

An administrator is deploying NVIDIA AI Enterprise on a Kubernetes cluster and must decide how GPU workloads should request accelerators. The environment has a mix of full-GPU training jobs and inference services that share a single A100. Which approach correctly allows a pod to consume a specific MIG-backed slice rather than the whole device?

A.Request nvidia.com/gpu: 1 and add the annotation nvidia.com/mig-profile=2g.10gb to the pod metadata.
B.Set the environment variable NVIDIA_MIG_PROFILE=2g.10gb in the container spec and request nvidia.com/gpu: 1.
C.Deploy a separate RuntimeClass named mig-2g.10gb and reference it from the pod's spec.runtimeClassName.
D.Request the specific extended resource, such as nvidia.com/mig-2g.10gb, in the pod's resource limits.
AnswerD

The NVIDIA device plugin advertises each configured MIG profile as its own extended resource name. A pod that requests nvidia.com/mig-2g.10gb in its limits is scheduled onto a node with a free instance of exactly that profile, giving the inference service a dedicated slice while full-GPU training jobs use other devices.

Why this answer

In Kubernetes, GPU allocation is expressed exclusively through extended resource requests handled by the NVIDIA device plugin. When MIG is enabled, each configured profile is advertised under a distinct resource name, so a pod requesting a specific MIG profile is bound to a matching instance. This is what enables a single A100 to serve both whole-device training and sliced inference workloads.

Exam trap

The trap here is believing that annotations or environment variables can steer GPU allocation, when only extended resource requests in the pod spec determine what the device plugin hands out.

44
MCQmedium

Which component is responsible for exposing the GPU as a schedulable resource in a Kubernetes cluster?

A.NVIDIA Container Toolkit
B.NVIDIA Device Plugin
C.NVIDIA DCGM Exporter
D.The Kubernetes Scheduler.
AnswerB

The device plugin is the crucial component for Kubernetes integration. It polls the host for GPU status and reports capacity to the kubelet, which then tells the scheduler. This bridge is essential for enabling the 'nvidia.com/gpu' resource type, which allows pods to request GPUs as first-class resources.

Why this answer

The Kubernetes device plugin is the interface that allows the kubelet to communicate with the GPU. It advertises the number of available GPUs on each node to the Kubernetes API server. When a pod requests a GPU, the scheduler uses this information to place the workload on the correct node.

Without this plugin, Kubernetes is unaware of the GPU hardware, making it impossible to manage and allocate GPU resources for containerized workloads.

Exam trap

Candidates often select the NVIDIA Container Toolkit or GPU Operator instead of the specific component that directly advertises resources to the kubelet scheduler.

45
Multi-Selecthard

A research lab is deploying NVIDIA AI Enterprise on an air-gapped Kubernetes cluster. The cluster has no internet access, and all software must be installed from a local registry. The administrator plans to use the NVIDIA GPU Operator. Which two actions must be performed to ensure a successful deployment in this environment? (Choose two.)

Select 2 answers
A.Install the NVIDIA GPU Operator using the --offline flag.
B.Disable the GPU Operator's driver management and install the driver manually on each node.
C.Configure the GPU Operator to use the private registry by setting the appropriate Helm chart values.
D.Mirror all required NVIDIA container images to a private registry accessible by the cluster.
E.Ensure the cluster nodes have direct access to the NVIDIA licensing server for vGPU.
AnswersC, D

The GPU Operator must be configured to pull images from the private registry instead of the default public registries. This is done by setting Helm values such as operator.repository, driver.repository, toolkit.repository, and others to point to the private registry. Without this configuration, the Operator would attempt to pull from nvcr.io and fail due to no internet access. Thus, this is a required action for air-gapped deployment.

Why this answer

In an air-gapped environment, the GPU Operator cannot reach public registries to pull container images. Therefore, all necessary images must be mirrored to a private registry, and the GPU Operator must be configured to use that registry via Helm values. These two actions ensure that the Operator can deploy its components without internet access.

Other options are either not required or invalid for this scenario.

Exam trap

The trap here is thinking that an offline flag exists or that driver management must be disabled, when the key steps are mirroring images and configuring the private registry.

46
MCQeasy

An administrator is preparing a Kubernetes cluster for AI workloads and needs to ensure that the NVIDIA GPU Operator can be installed. The cluster nodes have NVIDIA GPUs, and the administrator wants to verify that the nodes are ready. Which command should the administrator run to check if the NVIDIA driver is already loaded on a node?

A.lspci | grep -i nvidia
B.kubectl get nodes -o wide
C.kubectl describe node <node-name>
D.nvidia-smi
AnswerD

nvidia-smi is the NVIDIA System Management Interface command. Running it on a node displays GPU information and driver version, confirming that the driver is loaded and functional. This is the standard way to verify driver installation on a node.

Why this answer

To verify that the NVIDIA driver is loaded on a node, the administrator should run nvidia-smi. This command queries the driver and displays GPU details such as driver version, GPU utilization, and memory usage. It is the definitive check for driver functionality.

Other commands may show GPU hardware or Kubernetes resources but do not confirm driver operation.

Exam trap

The trap here is confusing hardware detection with driver verification; lspci shows the GPU exists, but only nvidia-smi confirms the driver is active.

47
MCQmedium

A system administrator is installing NVIDIA AI Enterprise on a Kubernetes cluster that will use Multi-Instance GPU (MIG) on A100 GPUs. The administrator wants to ensure that MIG instances are properly exposed as schedulable resources. Which action must be taken after enabling MIG mode on the GPUs?

A.Manually create Kubernetes custom resources for each MIG instance using the NVIDIA MIG Manager.
B.Set the environment variable NVIDIA_MIG_CONFIG_DEVICES to 'all' on the kubelet and restart the kubelet service.
C.Deploy a separate device plugin for each MIG instance using a DaemonSet with node affinity to the specific GPU.
D.Install the NVIDIA GPU Operator with the MIG strategy set to 'mixed' and configure the device plugin to advertise MIG resources.
AnswerD

The GPU Operator supports MIG by deploying a device plugin that advertises MIG instances as resources. Setting the MIG strategy to 'mixed' allows both MIG and non-MIG GPUs in the cluster. The device plugin then exposes each MIG instance as a schedulable resource, enabling pods to request specific MIG profiles. This is the correct way to integrate MIG with Kubernetes scheduling.

Why this answer

To expose MIG instances as schedulable resources in Kubernetes, the NVIDIA GPU Operator must be installed with the MIG strategy configured (e.g., 'mixed' or 'single'). The Operator's device plugin then discovers and advertises each MIG instance as a resource. This allows pods to request specific MIG profiles via resource limits, enabling efficient scheduling.

Exam trap

The trap here is assuming that enabling MIG mode on the GPU is sufficient, when Kubernetes also requires the device plugin to advertise MIG resources.

48
MCQeasy

An administrator is deploying the NVIDIA GPU Operator on a Kubernetes cluster using Helm. The cluster nodes have NVIDIA GPUs and the administrator wants to ensure that the GPU Operator can automatically label nodes with GPU properties and install the device plugin. Which prerequisite must be met on the cluster nodes before installing the GPU Operator?

A.The nodes must have a supported container runtime such as Docker or containerd installed and configured.
B.The NVIDIA driver must be pre-installed on the nodes.
C.The nodes must have the NVIDIA Container Toolkit installed.
D.The nodes must have the NVIDIA DCGM exporter pre-installed for monitoring.
AnswerA

The GPU Operator relies on a container runtime to deploy its components and to run GPU-accelerated workloads. Kubernetes requires a container runtime like Docker or containerd to be installed and configured on each node. This is a fundamental prerequisite for any Kubernetes cluster, and the GPU Operator assumes it is present. Without a container runtime, the Operator cannot schedule its daemonsets, making this the correct prerequisite.

Why this answer

Before installing the NVIDIA GPU Operator, the Kubernetes cluster must have a container runtime such as Docker or containerd installed and configured on all nodes. This is a basic Kubernetes requirement because the Operator deploys its components as containers. Other components like the NVIDIA driver, container toolkit, and DCGM exporter are managed by the Operator itself and do not need to be pre-installed.

Exam trap

The trap here is assuming that NVIDIA-specific components like the driver or container toolkit must be pre-installed, when the GPU Operator is designed to manage them automatically.

49
MCQeasy

An administrator is installing the NVIDIA GPU Operator on a Kubernetes cluster. They want to verify that the GPU Operator's components are running correctly after installation. Which command should they use to check the status of the GPU Operator pods?

A.kubectl get pods -n gpu-operator
B.kubectl get nodes --show-labels
C.helm list -n gpu-operator
D.nvidia-smi -q
AnswerA

The NVIDIA GPU Operator installs its components in the 'gpu-operator' namespace by default. Running 'kubectl get pods -n gpu-operator' lists all pods in that namespace, allowing the administrator to verify that the operator and its managed components are running. This is the standard way to check the status.

Why this answer

The GPU Operator deploys its components into the 'gpu-operator' namespace. Checking the pods in that namespace with 'kubectl get pods -n gpu-operator' directly shows whether the operator and its managed pods are running. Other commands like nvidia-smi or helm list provide different information and do not replace the need to inspect pod status.

Exam trap

The trap here is confusing deployment verification with host-level GPU queries or Helm release checks, when the direct way to see component status is to list the pods in the operator's namespace.

50
MCQeasy

A cloud operations engineer is deploying the NVIDIA GPU Operator on a managed Kubernetes service where the worker nodes already have the NVIDIA data center driver installed by the cloud provider. The team wants the Operator to manage only the device plugin, container toolkit, and monitoring components. Which Helm value should the engineer set during installation?

A.--set driver.enabled=false
B.--set operator.driver.install=false
C.--set mig.strategy=none
D.--set toolkit.enabled=false
AnswerA

Setting driver.enabled=false tells the GPU Operator not to deploy the driver container, so it uses the preinstalled host driver. The Operator still deploys the device plugin, container toolkit, DCGM exporter, and other components. This is the standard approach on managed services where the provider owns the driver lifecycle and node image updates.

Why this answer

The GPU Operator chart exposes driver.enabled to control whether the driver container is deployed. On managed Kubernetes services the node image already contains a validated NVIDIA driver, so disabling the Operator's driver avoids conflicts and duplicate work while still letting the Operator manage the device plugin, container toolkit, DCGM exporter, and related components. The other values either do not exist or disable components the team needs.

Exam trap

The trap here is inventing or misremembering a Helm key for skipping the driver, when the actual chart value is driver.enabled and the goal is to keep all other GPU software components Operator-managed.

51
MCQhard

A financial services company is deploying NVIDIA AI Enterprise on a Kubernetes cluster with strict security policies. They need to ensure that GPU workloads are isolated and that the NVIDIA GPU Operator components are deployed with least privilege. Which feature of the NVIDIA GPU Operator allows administrators to define granular permissions for its components?

A.Security Context Constraints
B.PodSecurityPolicy
C.Role-Based Access Control (RBAC)
D.Network Policies
AnswerC

RBAC in Kubernetes allows administrators to define roles and role bindings that specify which actions are permitted on which resources. The GPU Operator uses RBAC to grant its components the minimum necessary permissions. This enables least-privilege access and is essential for strict security environments.

Why this answer

RBAC is the Kubernetes mechanism for defining granular permissions. The NVIDIA GPU Operator deploys components with specific ServiceAccounts and RBAC roles that grant only the permissions needed to perform their functions. This aligns with least-privilege principles and is critical for security-sensitive deployments.

Other options are either deprecated, network-focused, or platform-specific.

Exam trap

The trap here is assuming that PodSecurityPolicy or Network Policies can restrict API permissions; only RBAC governs what actions components can perform on Kubernetes resources.

52
MCQmedium

An administrator is deploying NVIDIA AI Enterprise on a bare-metal cluster. The workload requires full GPU isolation with minimal latency. Which configuration should the administrator select to achieve this goal?

A.Deploy using NVIDIA vGPU on a KVM hypervisor.
B.Configure the system using NVIDIA License System (NLS) in disconnected mode.
C.Implement bare-metal installation with NVIDIA GPUDirect RDMA enabled.
D.Utilize containerized GPU passthrough with a standard Docker runtime.
AnswerC

GPUDirect RDMA allows direct memory access between the GPU and third-party devices such as NICs, effectively bypassing the host CPU. This architecture eliminates unnecessary data copies and context switches, providing the lowest possible latency for high-speed AI data pipelines and distributed training environments on physical infrastructure.

Why this answer

GPU Direct and bare-metal deployments are essential for latency-sensitive workloads. By avoiding hypervisor overhead, the administrator ensures direct path access to the GPU memory and interconnects. This configuration is critical in high-performance computing environments where jitter and interrupt latency can degrade model training performance.

Selecting the right deployment mode is a foundational step in AI Operations to ensure optimal hardware utilization and predictable execution times for deep learning models.

Exam trap

Candidates often incorrectly choose virtualization or container-only solutions, failing to recognize that 'bare-metal' and 'GPUDirect RDMA' are the specific, non-negotiable requirements for minimizing latency in high-performance computing environments.

53
MCQmedium

Refer to the exhibit. An administrator notices poor performance in an AI training job. What is the most likely cause based on the CLI output?

A.The GPU is overheating due to a fan failure.
B.The power limit is set too low for the current workload.
C.The GPU driver is corrupted or out of date.
D.The workload is waiting for CPU memory allocation.
AnswerB

The 'Sw Power Cap' status confirms that the GPU is limited by the current software configuration. To improve performance, the administrator should evaluate the power policy settings to determine if the wattage limit can be safely increased to allow the GPU to reach its maximum boost clock frequency.

Why this answer

The output indicates that the GPU is currently throttling due to 'Sw Power Cap'. This means the software-defined power limit is restricting the GPU performance to stay within a specific wattage budget. This is common in densely packed servers or cloud environments where power infrastructure is shared.

Monitoring power usage is vital because it directly impacts clock speeds, which in turn bottleneck training throughput and lengthen the time required for model convergence.

Exam trap

Candidates often mistake power capping errors for hardware faults or thermal overheating issues, ignoring the explicit 'Sw Power Cap' message in the telemetry output.

54
MCQeasy

During an NVIDIA AI Enterprise deployment, you are asked to configure the 'NVIDIA Container Toolkit'. What is its primary function?

A.To provide high-level APIs for neural network training.
B.To allow the container runtime to interact with the host's GPU.
C.To act as a package manager for AI model weight files.
D.To automatically optimize neural network hyper-parameters.
AnswerB

The Container Toolkit provides the necessary hooks and libraries to map GPU resources into the container namespace. This allows the application running inside the container to make calls to the GPU hardware, which would otherwise be inaccessible due to the isolation boundaries of the container runtime environment.

Why this answer

The Container Toolkit enables containers to access the GPU by exposing the necessary drivers, libraries, and device files from the host into the container runtime. It acts as the critical bridge between the hardware-level drivers and the containerized applications. Understanding this is foundational for AI Ops, as it ensures that containerized AI models can actually utilize the GPU hardware for training or inference tasks without manual configuration.

Exam trap

Candidates often describe the Toolkit as a library installer for the container, rather than its primary role as the runtime interface enabling GPU resource access.

55
MCQmedium

Refer to the exhibit. An AI engineer observes that a model training job is running slower than expected. Based on the output, what is the primary cause of the performance degradation?

A.The GPU is overheating and the thermal management system has engaged to prevent hardware damage.
B.The GPU is currently idle and the driver has shifted the card to an energy-saving state.
C.An administrator has set a software power limit that is lower than the GPU's maximum performance threshold.
D.The GPU is experiencing a PCIe bus error, forcing the system to reduce clock speeds to maintain stability.
AnswerC

The 'SW Power Cap' indicator explicitly confirms that an administrative policy or software command has capped the power usage of the GPU. This forces the GPU to maintain lower clock speeds, directly impacting the throughput of the training job by limiting the available computational resources.

Why this answer

The 'SW Power Cap' throttle reason indicates that the power limit is configured below the card's maximum design capacity, causing the GPU to downclock to P12 state to stay within the power envelope. This is crucial for AI operations because it signals that the hardware is being artificially throttled, necessitating a review of power policies to ensure optimal training performance in high-compute scenarios.

Exam trap

Candidates frequently mistake power throttling for thermal throttling. They look at temperature metrics instead of examining the specific 'Power Cap' status flag, which indicates an administrative software limit is active.

56
MCQeasy

When deploying NVIDIA containers using the NVIDIA Container Toolkit, what is the primary function of the 'nvidia-container-runtime'?

A.It automatically recompiles the application code for the specific GPU architecture found on the host.
B.It manages the lifecycle of the GPU driver installation on the host operating system.
C.It exposes the host's NVIDIA GPUs and driver libraries to the container environment.
D.It monitors the temperature and power consumption of the GPUs during container execution.
AnswerC

This runtime transparently mounts the host GPU devices and user-mode NVIDIA libraries into the container. It modifies the container's OCI specification during the execution phase, ensuring that the containerized process can communicate with the physical GPU drivers installed on the host operating system.

Why this answer

The nvidia-container-runtime is a crucial component that allows Docker containers to interface with host GPUs. By modifying the container runtime specification, it ensures that the necessary device nodes and NVIDIA driver libraries are injected into the container namespace at startup. This enables seamless hardware acceleration for AI applications without requiring users to manually manage drivers or complex device path configurations inside their container images.

Exam trap

Candidates frequently confuse the container runtime's role with orchestration components like the device plugin, failing to realize the runtime directly injects driver libraries and device nodes into the container namespace.

57
MCQmedium

An AI operations team is installing the NVIDIA GPU Operator on a Kubernetes cluster that uses a custom containerd configuration. They need to ensure that the GPU Operator can properly manage the container runtime. Which action should they take before installing the GPU Operator?

A.Label the nodes with the appropriate container runtime version.
B.Set the default runtime in containerd to nvidia-container-runtime.
C.Disable the containerd systemd service and let the GPU Operator start its own runtime.
D.Ensure that the containerd configuration does not already include conflicting NVIDIA runtime settings.
AnswerD

The GPU Operator manages the NVIDIA Container Toolkit and runtime configuration. If containerd already has manually added NVIDIA runtime settings, these can conflict with the operator's configuration, leading to failures. It is best practice to remove any existing NVIDIA runtime configuration before installation so the operator can set it up cleanly.

Why this answer

Before installing the GPU Operator, any pre-existing NVIDIA runtime configuration in containerd should be removed to avoid conflicts. The operator manages the runtime configuration itself, so manual settings can interfere. Other options like changing the default runtime or disabling containerd are incorrect because the operator integrates with the existing runtime rather than replacing or requiring manual overrides.

Exam trap

The trap here is assuming the GPU Operator requires manual runtime configuration, when in fact it manages the runtime and conflicts can arise from pre-existing settings.

58
MCQmedium

When configuring the NVIDIA Device Plugin for Kubernetes, what is the purpose of the 'time-slicing' configuration?

A.To schedule tasks based on the specific time of day for load balancing.
B.To increase the GPU's clock frequency during high-demand periods.
C.To allow multiple pods to share a single physical GPU through context switching.
D.To restrict access to the GPU based on container process priority.
AnswerC

Time-slicing enables a physical GPU to be oversubscribed by allowing multiple pods to execute in turns on the same hardware. This increases the utilization of the GPU in scenarios where the individual pods do not require the full dedicated throughput of the entire device at all times.

Why this answer

Time-slicing is a technique that allows multiple Kubernetes pods to share a single GPU by rapidly switching context between them. In scenarios where full GPU isolation is not required, this increases resource utilization. It is a vital deployment strategy for optimizing cost and efficiency in shared environments, allowing administrators to balance the workload across limited GPU hardware without needing more expensive virtual machine-based partitioning solutions.

Exam trap

Candidates confuse time-slicing with hardware partitioning (MIG). Time-slicing is software-based context switching, whereas MIG provides true hardware-level isolation of GPU resources.

59
MCQmedium

During an air-gapped installation of the NVIDIA GPU Operator, the administrator must make all required images available to the cluster. Which component is responsible for pulling the Operator's operand images from the private registry?

A.The NVIDIA Container Toolkit must be pointed at the private registry through its config.toml file.
B.The Operator's Helm chart values must include an imagePullSecret for the private registry on every namespace.
C.The kubelet on each node must be configured with a mirror registry that rewrites all container image pulls.
D.The GPU Operator's ClusterPolicy must reference the private registry so operand pods use images from that registry.
AnswerD

In an air-gapped environment, the ClusterPolicy is configured with the private registry path and image repository settings for each operand, such as driver, toolkit, device plugin, and DCGM exporter. This ensures the Operator deploys pods that pull from the internal registry rather than the public NVIDIA registry, which is unreachable.

Why this answer

Air-gapped deployments require the ClusterPolicy to specify the private registry and repository paths for each operand image. This directs the Operator to deploy pods that pull driver, toolkit, device plugin, and monitoring images from the internal registry. Kubelet mirrors, toolkit configuration, and image pull secrets address adjacent concerns but do not select the operand image sources.

Exam trap

The trap here is assuming a kubelet mirror or pull secret is sufficient to redirect operand images, when the ClusterPolicy must explicitly reference the private registry.

60
MCQmedium

An AI platform engineer is preparing a fleet of NVIDIA DGX H100 systems for production workloads using the NVIDIA Base Command Manager (BCM). During initial bare-metal provisioning via PXE boot, the provisioning server successfully hands out IP addresses, but nodes consistently fail during the OS image deployment phase, throwing a kernel panic related to missing storage drivers. Which deployment step must be verified or corrected to ensure successful hardware-specific image deployment?

A.Reconfigure the DHCP server scope options to extend lease times and include custom vendor-class identifiers for the provisioning daemon.
B.Update the system BIOS Unified Extensible Firmware Interface boot order to prioritize internal redundant array of independent disks storage over network interfaces.
C.Verify that the target operating system image profile assigned in Base Command Manager includes the necessary storage controller modules and hardware support packages for the specific server model.
D.Modify the dynamic port forwarding settings on the top-of-rack management switches to allow uninterrupted trivial file transfer protocol block size extensions.
AnswerC

A kernel panic citing missing storage drivers means the assigned OS image profile lacks the storage controller modules and hardware support packages required by that server model. Correcting the image profile in Base Command Manager lets the deployed kernel detect the disks.

Why this answer

NVIDIA Base Command Manager relies heavily on matching the hardware profile of complex servers like DGX H100 with the correct custom software image and specialized storage drivers. If the deployment image lacks the required kernel modules for modern high-performance NVMe controllers, PXE booting fails at the kernel load stage. Ensuring the software image profile incorporates the exact hardware-specific drivers resolves this mismatch, allowing enterprise-grade bare-metal automation to complete seamlessly across the GPU cluster.

Exam trap

Candidates often assume standard Linux server golden images work universally across NVIDIA DGX hardware without customizing kernel modules for specialized high-performance storage and networking controllers.

61
MCQeasy

Which component is strictly necessary for managing NVIDIA AI Enterprise licensing across a distributed cluster of nodes?

A.NVIDIA CUDA Toolkit.
B.NVIDIA License System (NLS) Instance.
C.NVIDIA Container Toolkit.
D.NVIDIA Deep Learning GPU Training System (DIGITS).
AnswerB

The NLS instance acts as the centralized point for distributing and validating licenses to nodes in the cluster. It ensures that the enterprise software features are appropriately licensed, providing the necessary reporting and validation required by the AI Enterprise software subscription models for large-scale distributed computing environments.

Why this answer

The NVIDIA License System (NLS) is the centralized authority for managing product entitlements. In a distributed environment, nodes must reach this server to validate their software keys. This component is crucial because it decouples the license management from the individual worker nodes, allowing for flexible scaling and centralized compliance reporting, which are essential in enterprise AI infrastructure where hardware resources are frequently provisioned and decommissioned.

Exam trap

Candidates mistakenly believe that individual offline license keys stored locally on each worker node are sufficient for managing enterprise-wide distributed NVIDIA AI Enterprise deployments.

62
MCQmedium

Refer to the exhibit. The administrator has deployed the GPU Operator, but the node does not show GPU resources. What is the most likely cause?

A.The node is not part of a valid Kubernetes namespace.
B.The NFD service is not installed or configured correctly.
C.The GPU firmware is out of date.
D.The cluster is using an unsupported CPU architecture.
AnswerB

Node Feature Discovery is responsible for scanning the hardware and applying labels such as 'nvidia.com/gpu.present'. Without these labels, the GPU Operator will not know which nodes to target for driver installation, causing the node to remain without detected GPU resources in the Kubernetes API server.

Why this answer

The lack of GPU resources in the node description indicates that the Node Feature Discovery (NFD) or the GPU Operator has not correctly identified or labeled the GPU hardware. This is a common deployment issue where the hardware-to-software handshake fails due to missing labels. Identifying this early is key to ensuring that the Kubernetes scheduler can correctly place AI workloads on nodes equipped with the necessary compute resources.

Exam trap

Candidates often focus on the GPU Operator itself, forgetting that the GPU Operator relies on NFD to detect and label the hardware so the Kubernetes scheduler knows where to place workloads.

63
Multi-Selecthard

A cloud operations team is deploying the NVIDIA GPU Operator in an environment where the Kubernetes control plane cannot reach the public internet, but worker nodes can access an internal HTTP registry that mirrors required images. The team wants to avoid manual image pulls on each node. Which two configurations should they implement to enable a successful air-gapped installation? (Choose two.)

Select 2 answers
A.Use a private image registry and configure image pull secrets in the `gpu-operator` namespace for authentication.
B.Configure the ClusterPolicy with `operator.defaultRuntime: containerd` and set `driver.repository` to the internal registry path.
C.Disable the Node Feature Discovery (NFD) component to reduce the number of images that need to be mirrored.
D.Mirror all GPU Operator component images to the internal registry and update the ClusterPolicy to reference that registry for each component.
E.Set `driver.enabled: false` to avoid pulling the driver image from the public registry.
AnswersA, D

A private registry often requires authentication. Creating an image pull secret in the operator's namespace and referencing it in the ClusterPolicy or service accounts allows nodes to pull images securely. This is a standard requirement for air-gapped installations where the internal registry is not publicly accessible, ensuring components can authenticate and retrieve images.

Why this answer

Air-gapped installations require that all container images used by the GPU Operator are available in a reachable registry. Mirroring every component image and updating the ClusterPolicy to reference the internal registry ensures pods can start without internet access. Additionally, if the registry requires authentication, an image pull secret must be configured in the operator's namespace to allow secure pulls.

Exam trap

The trap here is thinking that simply disabling a component or setting a repository path is enough for air-gapped operation, when all images must be mirrored and authentication configured if needed.

64
MCQhard

An administrator observes that despite the GPU Operator being installed, the pods cannot access the GPU. What is the most likely cause if the NVIDIA container runtime is properly configured?

A.The Kubernetes API server is down.
B.The host NVIDIA driver is not loaded.
C.The pod has too little memory requested.
D.The container image is missing the CUDA library.
AnswerB

Even with the correct runtime, the container requires the underlying host driver to communicate with the hardware. If the driver is not installed or the kernel module is not loaded, the runtime will fail to map the GPU device, leading to a situation where the GPU is inaccessible.

Why this answer

If the runtime is configured, the most common remaining failure is the lack of the correct NVIDIA drivers on the host node. If the kernel modules are not loaded or the driver version is incompatible with the installed GPU, the runtime will be unable to successfully inject the device nodes. This is a common installation oversight where the operator might be deployed but the driver installation task failed or was skipped.

Exam trap

Candidates often assume the issue is a Kubernetes misconfiguration or a pod error. They fail to check the underlying host kernel, where driver loading issues are the most frequent root cause.

65
MCQmedium

What is the primary role of a private container registry in an NVIDIA AI Enterprise deployment?

A.Managing hardware firmware updates
B.Storing and distributing container images
C.Monitoring real-time GPU thermals
D.Providing GPU driver licensing keys
AnswerB

A private registry acts as a local source of truth for container images. In production environments, it is essential for security and stability, allowing administrators to control exactly which software versions are deployed to the cluster while reducing reliance on external registries that may have latency or availability issues.

Why this answer

A private registry provides a secure, reliable, and high-speed source for container images within the enterprise firewall. By caching NVIDIA-provided containers locally, organizations avoid dependency on public repositories, which is critical for security compliance and offline operational readiness. It ensures that the cluster has consistent, immutable versions of software available, preventing drift and ensuring that deployments remain stable and reproducible across the entire production environment.

Exam trap

Candidates often assume a registry is for model versioning or training data storage. In the context of NVIDIA AI Enterprise, its primary function is strictly container image management and distribution.

66
MCQhard

A company is deploying NVIDIA AI Enterprise on a vSphere cluster with multiple ESXi hosts, each containing NVIDIA A100 GPUs. They need to enable vGPU functionality for virtual machines running AI workloads. Which configuration step is required on the ESXi host to allow vGPU assignment to VMs?

A.Install the NVIDIA vGPU Manager on the ESXi host and set the graphics type to shared or direct.
B.Install the NVIDIA GPU Operator on the ESXi host and create a custom resource for vGPU.
C.Enable SR-IOV on the ESXi host and configure virtual functions for each VM.
D.Install the NVIDIA Container Toolkit on the ESXi host and configure the Docker daemon.
AnswerA

To enable vGPU on ESXi, you must install the NVIDIA vGPU Manager VIB on each host and configure the GPU's graphics type. For vGPU, the graphics type is typically set to 'shared' to allow multiple VMs to share the GPU, or 'direct' for passthrough. This is a required step before vGPU profiles can be assigned to VMs.

Why this answer

Enabling vGPU on ESXi requires the NVIDIA vGPU Manager VIB and setting the GPU graphics type appropriately. This allows the hypervisor to present virtual GPUs to VMs. Alternatives like the Container Toolkit or GPU Operator are not applicable to ESXi host-level configuration, and SR-IOV is not the mechanism used for vGPU in this context.

Exam trap

The trap here is assuming that Kubernetes-centric tools like the GPU Operator or Container Toolkit are used for vSphere vGPU setup, when actually the hypervisor-specific vGPU Manager is required.

67
MCQmedium

A platform team is deploying the NVIDIA GPU Operator on a Kubernetes cluster to manage GPU nodes. They want the Operator to automatically install the NVIDIA driver, the container toolkit, and the device plugin on each GPU node. Which component of the GPU Operator is responsible for installing the NVIDIA driver on the host?

A.NVIDIA Container Toolkit
B.NVIDIA Driver Container
C.NVIDIA Device Plugin
D.NVIDIA GPU Operator Validator
AnswerB

The NVIDIA Driver Container is a containerized version of the NVIDIA driver that the GPU Operator deploys as a DaemonSet on GPU nodes. It installs the driver into the host's filesystem, enabling driver management without pre-installing the driver on the host OS. This is exactly what the team needs for automatic driver installation.

Why this answer

The GPU Operator automates the deployment of all NVIDIA software components on Kubernetes. The Driver Container is specifically designed to install the NVIDIA driver on the host. It runs as a DaemonSet and installs the driver into the host's filesystem, eliminating the need for manual driver installation.

The other components handle container runtime integration, resource advertisement, and validation.

Exam trap

The trap here is assuming that the NVIDIA Container Toolkit installs the driver, when it actually only enables containers to access GPUs and requires the driver to be present.

68
MCQeasy

Which utility is primarily used to monitor and manage NVIDIA GPU power, temperature, and usage statistics in real-time on a Linux-based deployment?

A.top
B.nvidia-smi
C.htop
D.dmesg
AnswerB

This is the official utility for managing NVIDIA GPU devices. It provides comprehensive real-time data on GPU usage, memory allocation, power consumption, and thermal status. It also allows administrators to set performance limits and reset devices, which is essential for maintaining a healthy and performant AI cluster.

Why this answer

nvidia-smi (NVIDIA System Management Interface) is the command-line utility used to interact with the NVIDIA driver for monitoring and managing GPU devices. It is the fundamental tool for AI ops teams to verify that GPUs are properly detected and performing within expected thermal and power envelopes during model training or inference tasks on DGX or workstation systems.

Exam trap

Candidates frequently confuse nvidia-smi with higher-level management tools or cloud-native exporters like DCGM, overlooking its role as the foundational host-level CLI utility.

69
MCQmedium

Refer to the exhibit. An administrator attempts to deploy a GPU-based pod, but it remains in the 'Pending' state. What is the most likely cause based on the error log?

A.The pod is requesting more CPU cores than the node can provide.
B.The NVIDIA device plugin is not correctly advertising GPU resources.
C.The node is in a 'NotReady' state due to high disk latency.
D.The pod security policy prohibits the use of GPU-accelerated containers.
AnswerB

The device plugin is the component responsible for telling the Kubernetes scheduler how many GPUs are available on a specific node. If this process fails or is not running, the scheduler will not see any allocatable GPU resources, resulting in the 'Insufficient' error even if physical GPUs are present.

Why this answer

The error 'Insufficient nvidia.com/gpu' indicates that the Kubernetes scheduler cannot find a node with available GPU resources that match the pod's request. This typically happens when the device plugin is not correctly reporting resource availability to the scheduler, or the cluster is over-provisioned. Resolving this requires ensuring the GPU Operator is running and that the device plugin has successfully registered the GPUs with the Kubelet on the worker nodes.

Exam trap

Candidates often blame pod resource limits or application errors. They fail to recognize the 'Insufficient' resource error as a direct signal that the device plugin is not communicating with the Kubelet.

70
MCQeasy

An administrator is preparing to install the NVIDIA GPU Operator on a new Kubernetes cluster. The cluster uses containerd as the container runtime. Which prerequisite must be satisfied on each GPU node before the operator can successfully deploy the driver container?

A.The Linux kernel headers for the running kernel must be available on the node.
B.The node must have the `nvidia.com/gpu.present` label applied manually by the administrator.
C.The NVIDIA Container Toolkit must be installed manually on each node.
D.The node must have the NVIDIA driver preinstalled at the exact version specified by the operator.
AnswerA

The driver container compiles the NVIDIA kernel modules against the running kernel, so the matching kernel headers and build tools must be present. Without them, the driver build fails and the GPU cannot be initialized. This is a fundamental prerequisite for any node where the operator will install the driver via container.

Why this answer

The driver container builds NVIDIA kernel modules on the host, which requires kernel headers and a compiler toolchain matching the running kernel. Without these, the container cannot compile the modules and the driver fails to load. This prerequisite is essential for any node where the operator manages the driver, and it is often overlooked during cluster preparation.

Exam trap

The trap here is assuming the NVIDIA Container Toolkit or driver must be preinstalled, when the operator handles those components and instead requires kernel headers for module compilation.

71
Multi-Selectmedium

An administrator is preparing to deploy an NVIDIA AI Enterprise solution on an OpenShift cluster. Which TWO steps must be completed to ensure the NVIDIA drivers are loaded correctly on the worker nodes?

Select 2 answers
A.Deploy the NVIDIA GPU Operator via the OperatorHub.
B.Manually install NVIDIA drivers on every host OS.
C.Configure the Node Feature Discovery (NFD) operator.
D.Install the NVIDIA Triton Inference Server.
E.Set the Kubernetes scheduler to 'AlwaysPull'.
AnswersA, C

The OperatorHub provides the standard interface for managing lifecycle updates and dependencies within OpenShift. Deploying the GPU Operator from here automatically pulls in required dependencies like NFD, ensuring that the environment is correctly primed to handle the proprietary NVIDIA kernel modules and device plugin communication protocols.

Why this answer

On OpenShift, the NVIDIA GPU Operator leverages the Node Feature Discovery (NFD) and the Machine Config Operator to manage kernel modules. Ensuring these are configured correctly allows the driver container to load the proprietary NVIDIA kernel modules successfully. Proper node labeling is also required so the operator can target the correct hardware, ensuring that the driver is injected into the node's operating system environment during the boot process.

Exam trap

Candidates often overlook Node Feature Discovery (NFD), mistakenly assuming the GPU Operator alone can automatically detect hardware features on OpenShift worker nodes without NFD labeling.

72
MCQeasy

An AI operations engineer is preparing a Kubernetes cluster to run GPU-accelerated inference workloads using the NVIDIA GPU Operator. The cluster nodes already have NVIDIA data center GPUs installed, and the engineer wants to avoid installing the driver manually on each node. Which component of the GPU Operator is responsible for automatically deploying the NVIDIA driver on worker nodes?

A.NVIDIA Device Plugin
B.NVIDIA GPU Operator's driver container
C.NVIDIA DCGM Exporter
D.NVIDIA Container Toolkit
AnswerB

The GPU Operator includes a driver container that runs as a DaemonSet on GPU nodes. It compiles and loads the NVIDIA kernel driver inside a container, eliminating the need for manual host driver installation. This is the intended mechanism for automated driver deployment in Kubernetes environments, ensuring consistent driver versions across nodes without direct host modifications.

Why this answer

The driver container within the GPU Operator automates the deployment and management of NVIDIA drivers on Kubernetes nodes. It runs as a DaemonSet and ensures the correct driver version is loaded without manual intervention. This simplifies operations and maintains consistency across the cluster, which is essential for AI workloads that depend on specific driver capabilities.

Exam trap

The trap here is confusing the NVIDIA Container Toolkit with the driver container; the toolkit exposes GPUs to containers but does not install the kernel driver.

73
MCQmedium

A platform engineer is preparing a Kubernetes cluster to run AI training jobs that require GPU access. The cluster nodes have NVIDIA GPUs, and the engineer wants the GPU Operator to manage the driver lifecycle. Which component must be installed on the host nodes to allow the GPU Operator to load kernel modules and create device nodes?

A.NVIDIA Kubernetes Device Plugin
B.NVIDIA Container Toolkit
C.NVIDIA GPU Operator Driver Container
D.NVIDIA GPU Driver
AnswerC

The GPU Operator deploys a driver container that runs on each node to install the NVIDIA driver, load kernel modules, and create device nodes. This container manages the driver lifecycle, ensuring the correct version is used. It is the component that enables the GPU Operator to handle driver installation and updates automatically, which is essential for this scenario.

Why this answer

The GPU Operator uses a driver container to manage the NVIDIA driver on host nodes. This container loads kernel modules, creates device nodes, and ensures the driver version matches the operator's requirements. The NVIDIA Container Toolkit and Device Plugin are also deployed by the operator but serve different purposes: the toolkit enables container GPU access, and the device plugin advertises GPUs to Kubernetes.

The driver container is specifically responsible for the driver lifecycle, making it the correct component.

Exam trap

The trap here is confusing the driver container with the NVIDIA Container Toolkit or Device Plugin, assuming that any GPU-related component can manage the driver.

74
MCQmedium

An administrator is planning to monitor GPU utilization across a large cluster. Which component should be deployed to collect metrics that are compatible with Prometheus?

A.NVIDIA GPU Operator
B.NVIDIA DCGM Exporter
C.NVIDIA Triton Inference Server
D.NVIDIA CUDA Toolkit
AnswerB

The DCGM Exporter is purpose-built to extract hardware metrics and expose them in a format that Prometheus can scrape. This is the correct choice for integrating GPU telemetry into existing monitoring pipelines, enabling administrators to visualize performance data and effectively manage GPU resources within their Kubernetes infrastructure clusters.

Why this answer

NVIDIA DCGM (Data Center GPU Manager) is the industry-standard tool for collecting GPU health and telemetry data. By deploying the DCGM Exporter, the administrator enables the conversion of hardware telemetry into a Prometheus-friendly format. This integration is vital for observability, allowing teams to set up alerts and dashboards to track GPU usage, power consumption, and memory allocation across the entire fleet of accelerated computing nodes.

Exam trap

Candidates frequently confuse general Kubernetes metrics servers with GPU-specific telemetry tools, choosing standard kube-state-metrics instead of the dedicated NVIDIA DCGM Exporter.

75
MCQmedium

What is the primary benefit of using an NVIDIA-Certified System for AI Enterprise deployments?

A.Guaranteed 100% network uptime
B.Validated compatibility and performance
C.Automatic software license renewal
D.Free access to all AI models
AnswerB

NVIDIA-Certified Systems are tested against a strict validation suite to ensure that hardware, firmware, and software drivers work seamlessly together. This reduces the risk of deployment failure and performance degradation, providing a predictable and supported environment for AI Enterprise software workloads across diverse enterprise data center environments.

Why this answer

NVIDIA-Certified Systems have passed rigorous testing to ensure optimal performance and compatibility with the NVIDIA AI Enterprise software stack. This certification minimizes the risk of hardware-software conflicts, performance bottlenecks, and stability issues during production. For AI Ops professionals, this translates to faster deployment times, lower maintenance costs, and high confidence in the reliability of the underlying infrastructure for mission-critical deep learning and inference workloads.

Exam trap

Candidates often confuse 'NVIDIA-Certified' with 'any GPU-enabled server.' They incorrectly assume any hardware with an NVIDIA GPU provides the same level of validated performance, software stack compatibility, and production-grade support.

Page 1 of 2 · 92 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Installation and Deployment questions.