Courseiva

CCNA Installation and Deployment Questions

17 of 92 questions · Page 2/2 · Installation and Deployment · Answers revealed

76
MCQhard

When deploying NVIDIA AI Enterprise, why is the selection of the correct CUDA version in the container image critical during the installation phase?

A.It determines the speed of the GPU's memory bus.
B.The CUDA version must be compatible with the host driver version.
C.It enables the use of the NVIDIA License System.
D.It is required for the installation of the GPU Operator.
AnswerB

NVIDIA drivers follow a backward-compatibility model where the driver must support the CUDA version used by the application. Using a container with a newer CUDA version than the driver supports will cause the application to fail to initialize, as it cannot properly map the required kernel functions.

Why this answer

The CUDA version dictates which APIs and features are available to the AI application. Because the driver on the host must support the CUDA version used by the container (backward compatibility), mismatching these leads to runtime failures. This is a crucial AI Ops consideration as it directly affects the stability of the entire stack, ensuring that the software environment aligns with the underlying hardware capabilities for maximum performance and reliability.

Exam trap

Candidates often assume that the container image includes its own driver, failing to realize that the host driver must be compatible with the CUDA version installed inside the container.

77
MCQhard

An AI operations engineer is deploying a multi-node Kubernetes cluster with NVIDIA A100 GPUs for distributed training. The engineer wants to ensure that GPUs are correctly discovered and that workloads can request GPU resources. After installing the NVIDIA GPU Operator, the engineer notices that the GPU nodes are not advertising any 'nvidia.com/gpu' resources. Which component should the engineer verify first to resolve this issue?

A.NVIDIA DCGM Exporter
B.NVIDIA GPU Operator's driver container
C.NVIDIA Container Toolkit
D.NVIDIA Device Plugin
AnswerD

The NVIDIA Device Plugin is the component that discovers GPUs and advertises them as schedulable resources (e.g., nvidia.com/gpu) to the Kubernetes API server. If GPUs are not appearing as resources, the device plugin is the first component to check. It could be failing to start, unable to communicate with the kubelet, or missing driver dependencies.

Why this answer

When GPU resources are not advertised, the NVIDIA Device Plugin is the primary component to investigate. It runs as a DaemonSet on GPU nodes and registers GPUs with the kubelet. Issues such as pod failures, missing driver libraries, or misconfigured kubelet settings can prevent resource advertisement.

Checking its logs and status will reveal the root cause.

Exam trap

The trap here is focusing on the driver or toolkit first; however, the device plugin is the component that directly advertises GPU resources to Kubernetes.

78
MCQmedium

Which TWO of the following steps are essential when deploying the NVIDIA GPU Operator on a Kubernetes cluster to ensure that GPU resources are discoverable by the scheduler?

A.Enable the node feature discovery (NFD) service in the operator configuration.
B.Manually install the NVIDIA Container Toolkit on every worker node prior to operator deployment.
C.Configure the GPU device plugin to register NVIDIA-specific resources with the Kubelet.
D.Modify the Kubernetes API server manifest to include the NVIDIA-specific admission controller.
E.Disable the default Kubernetes scheduler to allow the NVIDIA scheduler plugin to take over.
AnswerA, C

NFD is essential for labeling nodes based on hardware features, such as GPU architecture. The GPU Operator uses these labels to identify nodes suitable for GPU-accelerated workloads. Without NFD, the cluster cannot dynamically detect and categorize the GPU hardware capabilities required to schedule pods accurately across the node pool.

Why this answer

The NVIDIA GPU Operator automates the lifecycle of NVIDIA software components, including the driver, toolkit, and device plugin. By configuring the operator to deploy the GPU device plugin and the Node Feature Discovery (NFD) service, the cluster gains the ability to identify GPU hardware and advertise it as an allocatable resource. Without these, the Kubernetes scheduler cannot place pods requiring GPU resources, leading to 'Pending' status for those workloads.

Exam trap

Candidates often assume that installing the GPU driver is sufficient. They miss that the Kubernetes scheduler requires explicit registration via NFD and the device plugin to actually 'see' the GPU as a resource.

79
MCQmedium

Which security configuration is necessary when deploying NVIDIA GPUs in a multi-tenant environment to prevent unauthorized access between containers?

A.Increase the Kubernetes pod memory limit.
B.Enable MIG (Multi-Instance GPU).
C.Use a privileged container.
D.Install the latest version of CUDA.
AnswerB

MIG provides hardware-level partitioning, ensuring that individual GPU instances have isolated compute and memory. This is the recommended security practice for multi-tenancy, preventing cross-tenant resource leakage and providing strong isolation that software-based approaches cannot match, which is critical for compliance and security in shared infrastructure deployments.

Why this answer

In multi-tenant environments, using MIG (Multi-Instance GPU) is the standard method for hardware-level isolation. MIG allows a single physical GPU to be partitioned into multiple isolated instances, each with its own dedicated compute and memory. This ensures that one tenant's workload cannot interfere with or access the resources of another, providing a robust security boundary that software-only isolation cannot reliably guarantee in accelerated computing environments.

Exam trap

Candidates often suggest software-based container isolation (like namespaces or cgroups). While these provide process boundaries, they do not provide the hardware-level memory and compute isolation required for secure GPU multi-tenancy.

80
MCQmedium

What is the primary function of the NVIDIA Persistence Daemon in an AI deployment?

A.It automatically updates the NVIDIA driver when a new version is released on the web.
B.It keeps the GPU driver initialized to reduce latency during application startup.
C.It monitors GPU health and automatically initiates a reboot if an error is detected.
D.It manages the network traffic between the GPU and the storage backend.
AnswerB

By maintaining the driver's state in memory, the Persistence Daemon eliminates the overhead associated with the driver unloading and reloading process. This ensures that GPU resources are always ready for immediate use, which is essential for performance-sensitive AI applications that require rapid task execution.

Why this answer

The Persistence Daemon ensures that the NVIDIA driver remains loaded even when no applications are using the GPU. This prevents the driver from unloading and then reloading when a job starts, which significantly reduces the startup latency of AI models. It is a critical configuration for high-performance environments where frequent job scheduling would otherwise incur unnecessary overhead from repeated driver and device initialization.

Exam trap

Candidates often think the persistence daemon is used for saving machine learning model checkpoints, confusing storage persistence with driver state persistence.

81
MCQmedium

An AI operations team is deploying NVIDIA AI Enterprise on a Kubernetes cluster using the NVIDIA GPU Operator. They need to ensure that GPU metrics such as utilization, memory usage, and temperature are collected and exposed to Prometheus for monitoring. Which component of the GPU Operator is responsible for this?

A.NVIDIA GPU Operator Validator
B.NVIDIA GPU Device Plugin
C.NVIDIA Container Toolkit
D.NVIDIA DCGM Exporter
AnswerD

The NVIDIA DCGM Exporter is a component deployed by the GPU Operator that collects GPU telemetry using NVIDIA Data Center GPU Manager (DCGM) and exposes it as Prometheus metrics. It provides metrics on utilization, memory, temperature, power, and more. This is the correct component for integrating GPU monitoring with Prometheus in a Kubernetes environment managed by the GPU Operator.

Why this answer

The NVIDIA DCGM Exporter is the component of the GPU Operator that gathers GPU metrics via DCGM and exposes them in Prometheus format. It is specifically designed for monitoring GPU health and performance in Kubernetes. Other components like the Container Toolkit, Device Plugin, and Validator serve different purposes and do not provide metrics collection for Prometheus.

Exam trap

The trap here is confusing the device plugin, which handles scheduling, with the DCGM Exporter, which handles monitoring metrics.

82
Multi-Selectmedium

An administrator is deploying the NVIDIA GPU Operator on a Kubernetes cluster in an air-gapped environment. The cluster nodes have no internet access, and all container images must be pulled from a private registry. Which two actions are required to ensure a successful deployment? (Choose two.)

Select 2 answers
A.Ensure that the private registry supports manifest lists for multi-architecture images.
B.Configure the GPU Operator to use the private registry by setting the image repository in the Helm chart values.
C.Mirror all required NVIDIA GPU Operator images to the private registry.
D.Disable the driver container and install drivers manually on each node.
E.Set the environment variable 'AIRGAP=true' in the operator's deployment.
AnswersB, C

The GPU Operator Helm chart allows specifying a custom image registry and repository for all components. By setting the appropriate values, the Operator will pull images from the private registry instead of NVIDIA's public registry. This configuration is essential to direct the Operator to the mirrored images, enabling successful deployment in an air-gapped setup.

Why this answer

Air-gapped deployments require that all container images are available in a private registry, and the GPU Operator must be configured to pull from that registry. Mirroring the images ensures availability, and setting the image repository in the Helm values directs the Operator to the correct location. These two actions are essential for a successful deployment without internet access.

Exam trap

The trap here is assuming that a special air-gap flag exists in the Operator, when in reality air-gapped support is achieved by mirroring images and configuring the registry settings.

83
MCQmedium

When configuring a node for NVIDIA AI Enterprise in a Kubernetes environment, what is the primary function of the NVIDIA Container Toolkit?

A.Provisioning virtual machines
B.Exposing GPUs to containers
C.Optimizing neural network layers
D.Managing Kubernetes network policies
AnswerB

The NVIDIA Container Toolkit provides the necessary runtime libraries and hooks that allow container orchestrators to inject GPU resources into a container. By modifying the container runtime specification, it ensures that device nodes and necessary drivers are mapped into the container's namespace during startup.

Why this answer

The NVIDIA Container Toolkit allows container engines, such as Docker or containerd, to interface with physical GPUs. It enables the exposure of GPUs inside containers, ensuring that applications can access CUDA libraries and hardware acceleration. Without this toolkit, the container runtime cannot bridge the gap between the host's GPU hardware and the containerized process, effectively rendering the GPU unusable for AI workloads.

Exam trap

Candidates often confuse the Container Toolkit with the GPU Operator itself. They assume the Toolkit manages the entire cluster lifecycle rather than its primary, specific purpose of exposing host GPUs to containers.

84
MCQeasy

A platform team is preparing a bare-metal Kubernetes cluster to run AI workloads. They want the GPU Operator to install and manage the NVIDIA driver automatically on each node. Which prerequisite must be satisfied on every worker node before the GPU Operator can succeed?

A.The node must have the NVIDIA Container Toolkit preinstalled and configured.
B.The node must be running the NVIDIA Data Center GPU Manager (DCGM) as a systemd service before joining the cluster.
C.The node must have a supported Linux kernel with matching kernel headers available for the driver container to build against.
D.The node must have the NVIDIA vGPU Manager installed so the physical GPU can be partitioned before driver installation.
AnswerC

The GPU Operator's driver container compiles the NVIDIA kernel module against the running kernel, so the matching kernel headers or development package must be present on the host. Without them the driver build fails during installation. This is a documented prerequisite for nodes where the Operator manages the driver rather than relying on a preinstalled one.

Why this answer

The driver container builds the NVIDIA kernel module at runtime against the host kernel, so matching kernel headers or development packages must be present on each node. The GPU Operator manages the container toolkit, DCGM, and device plugin itself, so those do not need to be preinstalled. vGPU Manager applies to virtualized deployments, not bare-metal driver automation.

Exam trap

The trap here is assuming the GPU Operator installs the driver without needing host-level build dependencies such as matching kernel headers.

85
Multi-Selecthard

A cloud architect is deploying NVIDIA AI Enterprise on a vSphere cluster with multiple ESXi hosts, each containing NVIDIA A100 GPUs. The architect must enable GPU virtualization using NVIDIA vGPU. Which two components are required to support vGPU on the ESXi hosts? (Choose two.)

Select 2 answers
A.NVIDIA GPU Operator for Kubernetes
B.NVIDIA CUDA Toolkit installed on the ESXi host
C.NVIDIA Container Toolkit installed on the ESXi host
D.NVIDIA vGPU guest driver installed in each VM
E.NVIDIA vGPU Manager for VMware ESXi
AnswersD, E

Each VM that uses a vGPU must have the NVIDIA vGPU guest driver installed. This driver communicates with the vGPU Manager on the host to access the GPU. Without the guest driver, the VM cannot utilize the virtual GPU. It is a required component for vGPU functionality.

Why this answer

To enable NVIDIA vGPU on VMware ESXi, the host must have the NVIDIA vGPU Manager VIB installed, which allows the hypervisor to partition the physical GPU. Each VM that uses a vGPU must also have the NVIDIA vGPU guest driver installed. These two components are essential; other tools like the GPU Operator or Container Toolkit are not used on ESXi.

Exam trap

The trap here is confusing Kubernetes-focused tools like the GPU Operator with hypervisor-level components required for vGPU on ESXi.

86
MCQhard

When configuring a multi-tenant environment on an NVIDIA DGX system, how are MIG (Multi-Instance GPU) instances best provisioned?

A.Through the BIOS settings, creating physical partitions on the GPU die.
B.Using the nvidia-smi command to define the desired MIG profiles.
C.By editing the /etc/nvidia/mig-config.json file directly.
D.Via the OS kernel boot parameters in the GRUB configuration.
AnswerB

nvidia-smi is the correct utility for creating and managing MIG profiles. By applying these profiles, the GPU is partitioned into distinct instances. This provides hardware-level isolation, which is critical for security and performance when multiple users share the same physical GPU infrastructure for different training tasks.

Why this answer

MIG allows a single GPU to be partitioned into multiple isolated instances, each with dedicated compute and memory resources. Configuring MIG via the nvidia-smi tool at the OS level ensures that these partitions are visible to the container runtime. This is crucial for AI operations, as it enables safe, secure multi-tenancy where different users can run jobs simultaneously without resource contention or cross-tenant interference.

Exam trap

Candidates frequently assume MIG is configured via Kubernetes manifests or YAML files. While Kubernetes manages the pods, the actual hardware partitioning must be defined at the OS level via nvidia-smi.

87
MCQmedium

A platform team is installing the NVIDIA GPU Operator on a Kubernetes cluster that runs a mix of GPU and non-GPU nodes. They want the operator to manage the driver lifecycle only on nodes that actually have NVIDIA GPUs, without requiring manual taints on non-GPU nodes. Which configuration should they apply to the ClusterPolicy to achieve this?

A.Set `driver.enabled: false` and rely on preinstalled drivers on all nodes.
B.Set `driver.rdma.enabled: true` to limit driver operations to GPU nodes.
C.Set `driver.nodeSelector` in the ClusterPolicy to match a label such as `nvidia.com/gpu.present: "true"`.
D.Set `driver.enabled: true` and `driver.useNvidiaDriverRoot: true` in the ClusterPolicy.
AnswerC

The `driver.nodeSelector` field in the ClusterPolicy restricts where the driver daemonset is scheduled. By selecting only nodes labeled with `nvidia.com/gpu.present: "true"`, the operator manages drivers exclusively on GPU nodes. This avoids manual taints on non-GPU nodes and matches the desired automatic scoping to GPU hardware.

Why this answer

The GPU Operator uses node selectors to determine where components are deployed. Configuring `driver.nodeSelector` with a label that identifies GPU-equipped nodes ensures driver management is scoped to those nodes only. This avoids the need for manual taints on non-GPU nodes and aligns with the operator's declarative model for heterogeneous clusters.

Exam trap

The trap here is confusing driver configuration fields such as `driver.enabled` or `driver.rdma.enabled` with node placement controls, when only `driver.nodeSelector` governs which nodes are targeted.

88
MCQhard

An administrator is tasked with deploying a multi-node training job using the NVIDIA GPU Operator. Which configuration must be present to ensure that pods are scheduled on nodes with identical GPU architectures to prevent performance degradation?

A.Enable the 'auto-scaling' feature in the NVIDIA GPU Operator.
B.Use node affinity labels based on NFD-provided hardware information.
C.Increase the timeout values for the NCCL collective communication operations.
D.Set the 'nvidia.com/gpu' resource limit to zero on all nodes except the master node.
AnswerB

Node Feature Discovery (NFD) labels nodes with their GPU model and architecture. By using these labels in the training pod's affinity configuration, the administrator forces the scheduler to select only nodes that match the desired hardware profile, ensuring consistent performance for distributed training jobs that rely on identical hardware capabilities.

Why this answer

To maintain high-performance, synchronized training, all participating nodes should share the same GPU architecture and interconnect type (e.g., NVLink or InfiniBand). Using Kubernetes node affinity or anti-affinity rules combined with the hardware labels automatically generated by the Node Feature Discovery (NFD) service ensures that the scheduler places the job exclusively on compatible nodes. This prevents the training from falling back to slower, sub-optimal communication paths or heterogeneous hardware modes.

Exam trap

Candidates often rely on default scheduling, forgetting that Kubernetes is unaware of GPU architecture differences. They fail to use NFD labels, leading to heterogeneous nodes that cause significant performance degradation.

89
Multi-Selecthard

An AI infrastructure team is deploying NVIDIA AI Enterprise on a Kubernetes cluster using the NVIDIA GPU Operator. They need to ensure that the GPU Operator can successfully manage GPUs and that workloads can consume GPU resources. Which two components does the GPU Operator deploy to enable GPU scheduling and container GPU access? (Choose two.)

Select 2 answers
A.NVIDIA Device Plugin
B.NVIDIA Container Toolkit
C.NVIDIA Persistence Daemon
D.NVIDIA MIG Manager
E.NVIDIA GPU Driver
AnswersA, B

The NVIDIA Device Plugin is deployed by the GPU Operator to advertise GPU resources to the Kubernetes API server. It allows the scheduler to allocate GPUs to pods that request them. Without it, Kubernetes would not know about the GPUs, and pods would not be scheduled with GPU resources. This component is essential for GPU scheduling.

Why this answer

The NVIDIA GPU Operator deploys the NVIDIA Device Plugin to enable Kubernetes to schedule pods with GPU resources, and the NVIDIA Container Toolkit to allow containers to access GPUs. The device plugin registers GPUs with the Kubernetes API, while the container toolkit configures the container runtime. Together, they provide the necessary integration for GPU-accelerated workloads.

Other components like the driver and MIG manager are important but not the specific enablers for scheduling and container access.

Exam trap

The trap here is assuming that the GPU driver alone is sufficient for GPU scheduling and container access, overlooking the need for the device plugin and container toolkit.

90
MCQmedium

An administrator is deploying the NVIDIA GPU Operator into an existing Kubernetes cluster where the NVIDIA driver is already installed and maintained by the node image. The team wants the Operator to manage only the container runtime, device plugin, and monitoring components. Which configuration should be applied to the GPU Operator's ClusterPolicy?

A.Set devicePlugin.enabled to false so the Operator does not advertise GPU resources.
B.Set driver.enabled to false so the Operator does not deploy the driver container.
C.Set dcgmExporter.enabled to false so the Operator does not conflict with the existing driver.
D.Set toolkit.enabled to false so the Operator does not modify the container runtime configuration.
AnswerB

Setting driver.enabled to false tells the GPU Operator to skip driver management and rely on the preinstalled host driver. The Operator then proceeds to deploy the container toolkit, device plugin, DCGM exporter, and other managed components. This is the documented way to integrate the Operator with nodes whose drivers are maintained outside the Operator's lifecycle.

Why this answer

To use the GPU Operator with drivers already managed by the node image, the ClusterPolicy must disable driver management so the Operator skips the driver container. The container toolkit, device plugin, and DCGM exporter remain enabled so the Operator still configures runtime integration, advertises GPU resources, and exposes telemetry as the team intends.

Exam trap

The trap here is confusing driver management with runtime or device plugin management and disabling a component the scenario actually wants the Operator to control.

91
MCQmedium

An administrator is deploying NVIDIA AI Enterprise on a bare-metal cluster. Which component must be installed first to ensure proper communication between the Kubernetes scheduler and the underlying GPU hardware?

A.NVIDIA Triton Inference Server
B.NVIDIA GPU Operator
C.NVIDIA NeMo Framework
D.NVIDIA Base Command Manager
AnswerB

The GPU Operator automates the installation of the NVIDIA driver, the Kubernetes device plugin, the DCGM monitoring agent, and other necessary components. Establishing this layer first ensures the cluster is GPU-aware and that the scheduler can effectively identify and assign physical resources to incoming application workloads.

Why this answer

The NVIDIA GPU Operator is essential for automating the management of all NVIDIA software components in Kubernetes. By installing it first, the administrator ensures that the device plugin, monitoring tools, and drivers are correctly configured. This foundation is critical for scheduling GPU-accelerated pods, as the Kubernetes scheduler requires the device plugin to advertise available GPU resources to the cluster's API server, enabling seamless workload orchestration across the infrastructure.

Exam trap

Candidates often think monitoring tools or device plugins must be installed individually first, overlooking that the GPU Operator automates and manages all underlying subcomponents.

92
MCQeasy

Which component in the NVIDIA AI Enterprise stack is responsible for providing the necessary user-space libraries and binaries to run GPU-accelerated applications inside containers?

A.NVIDIA GPU Operator
B.NVIDIA Container Toolkit
C.NVIDIA License System
D.NVIDIA Unified Fabric Manager
AnswerB

The NVIDIA Container Toolkit is specifically designed to provide the libraries and binaries required for containerized applications to perform GPU acceleration. It includes the runtime hook that makes the GPU visible to the container environment, ensuring the application can utilize the hardware for compute or graphics tasks.

Why this answer

The NVIDIA Container Toolkit is essential because it allows the container runtime to interact with the host's NVIDIA drivers. It provides the necessary libraries and the container runtime wrapper to expose GPUs inside the container environment. Without this toolkit, applications inside a container cannot access the GPU hardware, even if the host has the correct drivers installed, making it a critical deployment component.

Exam trap

Candidates often confuse the NVIDIA driver itself with the Container Toolkit. They assume drivers alone allow containers to access GPU hardware, ignoring the necessary user-space library mapping provided by the toolkit.

← PreviousPage 2 of 2 · 92 questions total

Ready to test yourself?

Try a timed practice session using only Installation and Deployment questions.