Courseiva
Software Development →mediumMultiple Select

NCA-GENL Software Development Practice Question

A developer is packaging a NeMo-based LLM application into a container for deployment on an NVIDIA GPU node. Which two practices are required to ensure the container can access the GPU and run inference efficiently? (Choose two.)

⚠ Common exam trap

The trap here is thinking the driver belongs inside the image, when in reality the host driver is injected by the NVIDIA Container Toolkit and only the user-space CUDA runtime should be pinned in the image.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Install the NVIDIA Container Toolkit on the host and run the container with the --gpus flag or equivalent runtime configuration.

GPU access in containers requires the NVIDIA Container Toolkit on the host plus a runtime flag, and efficient NeMo inference requires matching CUDA runtime, driver, and framework versions. Bundling drivers, disabling persistence mode, or shrinking shared memory does not enable GPU access and can degrade performance.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Install the NVIDIA Container Toolkit on the host and run the container with the --gpus flag or equivalent runtime configuration.

    Why this is correct

    The NVIDIA Container Toolkit exposes host GPU devices and driver libraries to containers. Without it and the corresponding runtime flag, the container sees no GPU and falls back to CPU or fails. This is a prerequisite for any GPU-accelerated NeMo inference workload on the node.

  • ✓

    Pin the CUDA runtime and relevant libraries in the image to versions compatible with the host driver and NeMo release.

    Why this is correct

    Compatibility between the CUDA runtime in the image, the host driver, and the NeMo framework release determines whether kernels load correctly. Pinning versions avoids silent ABI mismatches that cause runtime errors. This is essential for reproducible, efficient inference in the containerized NeMo application.

  • ✗

    Set the container's shared memory size to the minimum allowed value to conserve RAM.

    Why it's wrong here

    Reducing shared memory is counterproductive for GPU workloads that use it for inter-process communication and data staging; too small a value can cause crashes. Conserving RAM is not a goal that improves GPU access or inference efficiency, and the default is often already too low.

  • ✗

    Disable NVIDIA persistence mode on the host to reduce memory overhead.

    Why it's wrong here

    Persistence mode keeps the driver loaded and reduces initialization latency for repeated GPU work; disabling it generally hurts startup performance and is not a requirement for container GPU access. It also has no bearing on whether the container can see the device.

  • ✗

    Bake the full NVIDIA data center driver into the container image.

    Why it's wrong here

    Container images should not bundle the kernel-mode driver; the driver must match the host kernel and is injected by the toolkit. Baking it in creates version conflicts and bloats the image. The correct pattern is to rely on host driver injection plus a compatible CUDA runtime in the image.

About these practice questions

One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.