NCA-GENL Experimentation Practice Question
An ML engineer is comparing two fine-tuned variants of the same base LLM on an internal question-answering benchmark. Variant 1 scores higher on exact-match accuracy, but Variant 2 produces answers that human reviewers judge more helpful and better grounded. The engineer must decide which variant to promote. Which evaluation approach is most appropriate?
⚠ Common exam trap
The trap here is treating the decision as a choice between objective and subjective evaluation, when the two measure different quality dimensions and should be combined.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Combine the automated metric with a structured human or LLM-judge rubric that scores helpfulness and grounding, then decide using both.
When automated and human signals disagree, the right move is usually to measure both dimensions rather than discard one. A rubric-based human or LLM-judge evaluation for helpfulness and grounding complements exact-match accuracy, which is reproducible but blind to paraphrase and partial credit. Deciding with both signals yields a promotion choice that reflects real user value while retaining an objective regression check.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Promote Variant 1 because exact-match accuracy is an objective, reproducible metric that removes human subjectivity.
Why it's wrong here
Exact-match is objective but narrow; for open-ended question answering it penalizes paraphrases and partial correctness that humans accept. Choosing it solely for objectivity ignores the qualitative signal reviewers provided, so the promoted model may score well on the benchmark while disappointing users in real interactions.
- ✗
Promote Variant 2 without further measurement because human reviewers are the ultimate authority on answer quality.
Why it's wrong here
Human judgment is valuable but can be noisy, expensive, and inconsistent across reviewers, and it may miss regressions on tasks the reviewers did not sample. Discarding the automated metric entirely removes a cheap regression check and makes the promotion decision dependent on a small, possibly unrepresentative review set.
- ✓
Combine the automated metric with a structured human or LLM-judge rubric that scores helpfulness and grounding, then decide using both.
Why this is correct
Open-ended QA quality is multi-dimensional, so pairing a reproducible automated metric with a rubric-based judgment captures both correctness and the qualities reviewers valued. Using both signals lets the engineer weigh benchmark accuracy against real helpfulness and grounding, producing a decision that reflects actual user experience rather than a single narrow number.
- ✗
Retrain both variants on the benchmark questions until exact-match accuracy converges, then promote the higher scorer.
Why it's wrong here
Training on the evaluation set contaminates the benchmark through leakage, so the resulting scores no longer estimate generalization. The comparison becomes meaningless, and the promoted model may perform well only on the leaked questions. This approach also does nothing to address the helpfulness and grounding qualities the reviewers observed.
About these practice questions
One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.