NCA-GENL Software Development Practice Question
A developer is packaging a fine-tuned Llama 3 model as a TensorRT-LLM engine for an on-premises inference service. The model was trained with a custom tokenizer that adds four new special tokens beyond the base vocabulary. When the engine is built and the service is started, the model outputs garbled text and repeats the same fragment regardless of the prompt. Which action should the developer take to resolve this?
⚠ Common exam trap
The trap here is blaming generation settings such as temperature or repetition penalty for output corruption that actually originates from a tokenizer vocabulary mismatch.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Rebuild the TensorRT-LLM engine using the fine-tuned checkpoint's updated tokenizer vocabulary and matching special-token IDs.
Garbled, looping text from a fine-tuned model almost always traces back to a tokenizer or vocabulary mismatch between training and inference. Because the fine-tuned checkpoint added special tokens, the TensorRT-LLM engine must be compiled against the checkpoint's tokenizer metadata so token IDs align with the embedding table. Rebuilding the engine with the updated vocabulary is the only option that restores correct prompt encoding.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Convert the fine-tuned checkpoint back to a Hugging Face format and serve it through a generic Python HTTP wrapper instead of a compiled engine.
Why it's wrong here
Changing the serving framework does not address the root cause. The tokenizer mismatch exists in the model artifacts themselves, so a Python wrapper using the same stale tokenizer would produce the same garbled behavior. Abandoning the optimized engine also sacrifices the latency and throughput benefits the developer originally sought, making this a costly non-solution.
- ✗
Increase the engine's maximum batch size and rebuild the engine so the additional special tokens can be processed in parallel.
Why it's wrong here
Batch size affects throughput and memory, not token identity. The reported symptom is semantic corruption of every prompt, not a throughput shortfall. Raising the maximum batch size does nothing to reconcile the vocabulary difference between the fine-tuned checkpoint and the tokenizer used when the engine was built, so the garbled and repeating output would persist unchanged.
- ✓
Rebuild the TensorRT-LLM engine using the fine-tuned checkpoint's updated tokenizer vocabulary and matching special-token IDs.
Why this is correct
The garbled, repeating output is the classic signature of an input token ID mapping mismatch. The fine-tuned checkpoint extends the vocabulary by four tokens, so the engine must be built with that same tokenizer configuration and the same special-token ID assignments. Building the engine from the updated tokenizer keeps the prompt-to-embedding mapping consistent with training and restores coherent generation.
- ✗
Set the sampling temperature to zero and add a repetition penalty in the runtime generation config to stabilize the output.
Why it's wrong here
Sampling parameters shape the distribution over logits but cannot repair a broken token-to-embedding mapping. If the prompt is tokenized with the wrong vocabulary, the model receives embeddings that never corresponded to those inputs during fine-tuning. Deterministic decoding would simply produce the same incorrect fragment more consistently, masking rather than fixing the underlying tokenizer mismatch.
Visual reference
About these practice questions
One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.