Courseiva
Software Development →hardMultiple Choice

NCA-GENL Software Development Practice Question

A developer is packaging a generative AI application for NVIDIA AI Enterprise deployment on Kubernetes. The application must run an LLM served by NVIDIA NIM, an embedding model, and a vector database, and must support rolling upgrades without dropping in-flight inference requests. Which design choice best meets these requirements?

⚠ Common exam trap

The trap here is assuming that atomic or all-at-once upgrades simplify operations, when they actually guarantee downtime for in-flight inference.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Run each component as its own Kubernetes Deployment with readiness probes, multiple replicas, and a Service, using rolling update strategy and preStop hooks for graceful drain.

Independent Deployments per component, each with readiness probes, multiple replicas, and a stable Service, allow rolling updates that replace pods gradually while healthy capacity remains. A preStop hook combined with a termination grace period lets the NIM LLM finish in-flight generations before the process exits. This architecture also lets the LLM, embedding, and vector database tiers scale separately according to their distinct resource needs.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Run each component as its own Kubernetes Deployment with readiness probes, multiple replicas, and a Service, using rolling update strategy and preStop hooks for graceful drain.

    Why this is correct

    Separate Deployments let each tier scale and upgrade independently, readiness probes gate traffic until a replica is warm, and rolling updates replace pods gradually so capacity is maintained. A preStop hook plus termination grace period lets in-flight requests finish before shutdown. This directly satisfies the rolling-upgrade and no-dropped-requests requirements across the LLM, embedding, and database components.

  • ✗

    Use a single Deployment with a Recreate strategy so old pods are fully terminated before new ones start, guaranteeing no version mixing.

    Why it's wrong here

    The Recreate strategy tears down all existing pods before starting new ones, producing a period with zero serving capacity and guaranteed request failures. It avoids version mixing but at the cost of availability, which contradicts the requirement that in-flight inference requests not be dropped. Rolling updates with surge capacity are the appropriate strategy here.

  • ✗

    Deploy all three components in a single pod so they share a network namespace and can be upgraded together atomically.

    Why it's wrong here

    Co-locating components in one pod couples their lifecycles: any upgrade restarts the whole pod and drops in-flight requests for every component. It also prevents independent scaling of the LLM, embedding, and database tiers, which have very different resource profiles. Atomic upgrades sound appealing but conflict directly with the requirement for rolling upgrades without dropped requests.

  • ✗

    Place the NIM LLM and embedding model behind one Service and rely on client-side retries to mask any requests lost during upgrades.

    Why it's wrong here

    Client-side retries do not prevent dropped requests; they merely hide some of them and add latency while the service is unavailable. Sharing one Service across two different model endpoints also complicates routing and independent scaling. The requirement is to avoid dropping in-flight requests in the first place, which calls for graceful shutdown and rolling replacement rather than retry logic.

About these practice questions

This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.