NCA-GENL Software Development Practice Question
A developer is preparing a container for an LLM microservice that will run on an NVIDIA GPU node and must be deployable through NVIDIA NIM. They want the image to be portable across supported GPU generations while still using NVIDIA's optimized inference stack. Which two practices should they follow? (Choose two.)
⚠ Common exam trap
The trap here is assuming that installing drivers or precompiling engines inside the image increases portability, when both actually tie the artifact to a specific host or GPU SKU.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Base the service on an NGC-hosted NIM container image and pin a specific version tag.
Portable NIM deployment rests on using the NGC-published, version-pinned NIM image and exposing the GPU through the NVIDIA Container Toolkit. Bundling weights or drivers into the image, or shipping one SKU-specific TensorRT engine, either breaks the supported deployment model or destroys the cross-generation portability the team requires.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Base the service on an NGC-hosted NIM container image and pin a specific version tag.
Why this is correct
NGC-published NIM images already bundle the optimized runtime, CUDA libraries, and model-serving layer that NIM expects. Pinning an explicit version tag instead of latest keeps deployments reproducible and lets the team roll forward deliberately when a new image is validated against their traffic.
- ✓
Pass the GPU through with the NVIDIA Container Toolkit so the container can see the device.
Why this is correct
The NVIDIA Container Toolkit is what exposes the host driver and GPU devices inside the container. Without it the container cannot initialize CUDA, so the NIM process fails at startup regardless of how well the image itself was built or how the weights are supplied.
- ✗
Compile the model into a TensorRT engine for one exact GPU SKU and ship only that engine.
Why it's wrong here
A TensorRT engine is built for a specific compute capability and can fail or fall back on other GPU generations. Shipping a single SKU-specific engine contradicts the portability goal; NIM instead selects or builds an appropriate profile for the GPU it detects at runtime.
- ✗
Install a full CUDA toolkit and matching driver inside the image for maximum portability.
Why it's wrong here
GPU drivers belong on the host, and the container runtime injects the necessary user-space driver components. Installing a driver inside the image conflicts with the host driver, produces version mismatches, and is explicitly unsupported, so it reduces rather than improves portability across nodes.
- ✗
Bake the model weights directly into the image layer during the build step.
Why it's wrong here
Embedding multi-gigabyte weights into the image bloats every registry push and pull and couples model updates to image rebuilds. NIM expects weights and profiles to be provided through configuration and mounted cache volumes, so bundling them defeats the intended separation and slows deployment.
About these practice questions
This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.