Courseiva
Software Development →hardMultiple Choice

NCA-GENL Software Development Practice Question

A team fine-tunes a Llama model with NVIDIA NeMo Framework and must serve it behind an OpenAI-compatible endpoint with no Python glue code. They want the adapter weights kept separate from the base model so several adapters can share one loaded base. Which deployment approach fits these constraints?

⚠ Common exam trap

The trap here is treating any OpenAI-compatible endpoint as sufficient and overlooking that merging the adapter destroys the ability to share one base across several adapters.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Deploy the base model on NVIDIA NIM and load the fine-tuned adapter at runtime through multi-LoRA support in the NIM container.

The constraints are an OpenAI-compatible endpoint, no custom Python service code, and adapter weights stored separately from a shared base model. NIM's multi-LoRA serving satisfies all three: the container exposes standard APIs and loads adapters against one resident base. Merged engines, custom wrappers, and per-adapter Python backends each break at least one stated requirement.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Export the merged model to a TensorRT-LLM engine and serve it through NVIDIA Triton Inference Server with the TensorRT-LLM backend.

    Why it's wrong here

    This produces a high-performance engine and an OpenAI-compatible front end, but merging the adapter into the base bakes one adapter into each engine. Serving several adapters over a shared base would require a separate engine per adapter, multiplying GPU memory. The requirement to keep adapters separate from the base is therefore violated even though the endpoint works.

  • ✗

    Package the fine-tuned model inside a custom FastAPI service that wraps the NeMo checkpoint and exposes chat completions.

    Why it's wrong here

    A hand-written FastAPI wrapper is exactly the Python glue code the team wants to avoid, and it adds a serving layer the team must maintain, scale, and secure. It also does not provide adapter isolation from the base model unless the developer implements that logic. The approach works technically but fails the stated constraint about no custom glue code.

  • ✓

    Deploy the base model on NVIDIA NIM and load the fine-tuned adapter at runtime through multi-LoRA support in the NIM container.

    Why this is correct

    NIM microservices expose OpenAI-compatible APIs without custom glue code, and their multi-LoRA capability lets a single loaded base model host multiple adapters selected per request. This keeps adapter weights separate from the base, satisfies the shared-base requirement, and avoids building per-adapter engines. It directly matches every stated constraint.

  • ✗

    Convert the adapter to ONNX and deploy it with NVIDIA Triton Inference Server using the Python backend for each adapter.

    Why it's wrong here

    The Triton Python backend is still custom Python code that the team must author and maintain, so the no-glue-code constraint is not met. ONNX conversion of LoRA adapters also adds a fragile export step, and this path does not deliver OpenAI-compatible chat semantics out of the box. It is more moving parts than the requirement allows.

About these practice questions

Courseiva writes every NCA-GENL question from scratch — 367 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.