Courseiva
Software Development →hardMultiple Choice

NCA-GENL Software Development Practice Question

A team fine-tunes a Llama-3 8B model with NVIDIA NeMo Framework and must ship an inference artifact that a C++ service can load without a Python runtime. They want maximum throughput on Hopper GPUs and plan to serve many concurrent requests with in-flight batching. Which artifact and runtime pairing best satisfies these constraints?

⚠ Common exam trap

The trap here is treating any GPU-capable runtime as equivalent, when only the TensorRT-LLM C++ path combines a Python-free deployment with in-flight batching on Hopper.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Export a TensorRT-LLM engine and run it through the TensorRT-LLM C++ runtime with in-flight batching enabled.

The requirement combination of no Python runtime, Hopper-tuned kernels, and in-flight batching points to a TensorRT-LLM engine executed by its C++ runtime. Checkpoint formats, ONNX, and the transformers pipeline each fail at least one of the three constraints, most commonly the C++ deployment or the continuous batching capability.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Convert the checkpoint to ONNX and execute it with the ONNX Runtime CUDA execution provider.

    Why it's wrong here

    ONNX Runtime can run transformer graphs on CUDA, but its generative decoding path lacks the fused attention and paged KV cache kernels that TensorRT-LLM uses for in-flight batching. Throughput on Hopper would be far below a compiled engine, so this choice trades away the primary requirement.

  • ✗

    Save the fine-tuned checkpoint in NeMo .nemo format and load it with the NeMo inference toolkit in a Python worker.

    Why it's wrong here

    The .nemo checkpoint is a training and fine-tuning artifact that requires the NeMo Python stack to instantiate the model. It does not provide a C++ load path, and running it through the Python inference toolkit would reintroduce exactly the runtime dependency the team is trying to eliminate.

  • ✓

    Export a TensorRT-LLM engine and run it through the TensorRT-LLM C++ runtime with in-flight batching enabled.

    Why this is correct

    TensorRT-LLM compiles the fine-tuned weights into a GPU-specific engine and ships a C++ runtime that supports in-flight (continuous) batching, so new requests join an active batch without waiting. This removes the Python dependency, targets Hopper kernels at compile time, and is the intended path for high-throughput production serving.

  • ✗

    Merge the LoRA weights into the base model and serve the merged Hugging Face checkpoint with the standard transformers pipeline.

    Why it's wrong here

    Merging the adapter produces a valid model, but the transformers pipeline is Python-only and does not implement continuous batching with a paged KV cache. It would leave the C++ service dependent on a Python process and deliver much lower concurrency than a compiled engine.

About these practice questions

One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.