NCA-GENL Software Development Practice Question
A developer is containerizing an inference service built with TensorRT-LLM and NVIDIA NIM for a Kubernetes cluster. They want the deployment to start reliably and use the GPU efficiently. Which two practices should they follow? (Choose two.)
⚠ Common exam trap
The trap here is thinking that a restart policy or CrashLoopBackOff can substitute for proper GPU resource advertisement and pre-staged engine files.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Request nvidia.com/gpu resources in the pod spec and ensure the NVIDIA GPU Operator or device plugin is installed so the scheduler can allocate the device.
Reliable GPU-backed startup requires that Kubernetes can schedule the device through the NVIDIA device plugin or GPU Operator, and that the heavy engine artifacts are already present locally rather than downloaded at boot. Restart loops, bypassing the NVIDIA container runtime, and CPU-only emulation all prevent the engine from loading or waste resources instead of improving reliability and GPU efficiency.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Request nvidia.com/gpu resources in the pod spec and ensure the NVIDIA GPU Operator or device plugin is installed so the scheduler can allocate the device.
Why this is correct
Kubernetes does not expose GPUs to pods unless the NVIDIA device plugin or GPU Operator advertises nvidia.com/gpu as a schedulable resource. Without that resource request, the container starts without a GPU and inference fails. Declaring the resource also lets the scheduler place the pod on a node with a free device, which is fundamental to reliable startup.
- ✗
Disable the NVIDIA container runtime so the pod uses the host's default runtime and accesses the GPU through /dev/nvidia devices directly.
Why it's wrong here
Direct /dev access without the NVIDIA container runtime bypasses the libraries and driver mounts that CUDA needs inside the container, so inference fails even if the device node is visible. The container runtime is what injects the correct driver capabilities. Disabling it is a common misconfiguration that produces confusing CUDA initialization errors during pod startup.
- ✗
Set the pod's restartPolicy to Always and rely on CrashLoopBackOff to warm the GPU driver on each restart.
Why it's wrong here
Restart loops do not warm drivers or fix missing GPU allocation; they merely repeat the same failure. CrashLoopBackOff is a symptom of a misconfiguration, not a startup strategy. Relying on automatic restarts hides root causes such as a missing device plugin or an incompatible engine, and it wastes cluster resources while the pod thrashes.
- ✓
Mount the TensorRT-LLM engine files and model weights from a persistent volume or baked image layer rather than downloading them at pod startup.
Why this is correct
Engine files are large and architecture-specific, so fetching them during startup lengthens boot time and risks timeouts or partial downloads that crash the container. Baking them into the image or mounting them from a volume makes startup deterministic and repeatable. This also avoids recompiling engines on every pod restart, which preserves GPU efficiency and reliability.
- ✗
Pin the pod to a CPU-only node pool and use CUDA emulation to run the TensorRT-LLM engine.
Why it's wrong here
TensorRT-LLM engines require real NVIDIA GPU hardware; there is no CUDA emulation mode that executes a compiled engine. Scheduling onto a CPU-only node guarantees that the engine cannot load. This choice also defeats the efficiency goal, since the workload would never accelerate, and it contradicts the requirement to use the GPU efficiently.
About these practice questions
This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.