NCA-GENL Experimentation Practice Question
A team has fine-tuned a small LLM with NeMo and now wants to quantify how much the fine-tuning improved performance on a domain question-answering task relative to the base model. They have a curated set of 500 question-answer pairs that were never used during training. What is the most appropriate next step?
⚠ Common exam trap
The trap here is using training loss or training-set accuracy as evidence of improvement, when only held-out evaluation isolates the generalization benefit of fine-tuning.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Evaluate both the base model and the fine-tuned model on the held-out 500-question set using the same metric and decoding settings.
Quantifying the benefit of fine-tuning requires an unbiased comparison between the adapted model and its base counterpart on data neither has seen in training. Using the same held-out questions, prompts, decoding settings, and metric for both models yields a clean estimate of the improvement. Training loss, training-set accuracy, or a second training run do not measure generalization gain.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Evaluate both the base model and the fine-tuned model on the held-out 500-question set using the same metric and decoding settings.
Why this is correct
A held-out set that was excluded from training provides an unbiased estimate of generalization. Running both the base and fine-tuned models with identical prompts, decoding parameters, and scoring metric isolates the effect of fine-tuning. Comparing the two scores directly answers how much the fine-tuning improved domain question answering, which is the stated goal.
- ✗
Run the fine-tuned model on the original training questions and compare its answers to the stored labels.
Why it's wrong here
Evaluating on training data measures memorization rather than generalization, and the model has already seen those labels. High accuracy there says nothing about how the model handles new domain questions. The team already has a curated held-out set, so using training questions wastes the opportunity for an unbiased comparison.
- ✗
Fine-tune the base model a second time with a different seed and compare the two training curves.
Why it's wrong here
Comparing two fine-tuning runs reveals training stability but does not quantify improvement over the base model on the task. Training curves also do not substitute for task-level evaluation on unseen data. The scenario asks for the gain from fine-tuning, which requires scoring both the base and adapted models on the held-out questions.
- ✗
Report the final training loss of the fine-tuned model as the measure of improvement.
Why it's wrong here
Training loss reflects fit to the training data, not performance on unseen questions, and it has no comparable value for the base model. A lower training loss can coexist with worse generalization. Using it as the improvement metric would not answer the question and could actively mislead the team about the fine-tuning benefit.
About these practice questions
One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.