NCA-GENL Experimentation Practice Question
An ML engineer is comparing two fine-tuned variants of the same base LLM on an internal question-answering benchmark. Variant A scores higher on exact-match accuracy, while Variant B produces answers that human reviewers rate as more helpful and better formatted. The team must choose one variant to ship to production. Which evaluation approach is most appropriate for making this decision?
⚠ Common exam trap
The trap here is treating exact-match accuracy as inherently superior because it is numeric, when generative quality often depends on human-perceived helpfulness.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Combine automated metrics with a structured human-preference evaluation, and choose based on the product's primary success criterion.
For generative LLM tasks, automated metrics and human preference capture different quality dimensions. The right decision combines both and weights them by the product's primary success criterion. Shipping on accuracy alone ignores user-perceived helpfulness, shipping on preference alone ignores objective regressions, and averaging without justified weights hides the real trade-off.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Combine automated metrics with a structured human-preference evaluation, and choose based on the product's primary success criterion.
Why this is correct
Generative LLM evaluation typically pairs task-specific automated metrics with human or model-based preference judgments, because neither alone captures the full quality picture. Weighting them according to the product's success criterion, such as helpfulness for a user-facing assistant, yields a defensible shipping decision. This approach uses both signals the scenario provides.
- ✗
Average the exact-match score and a normalized human-preference score into a single number and pick the higher value.
Why it's wrong here
Blindly averaging incomparable metrics assumes they are on the same scale and equally important, which is rarely true. The weighting should reflect the product's priorities, not an arbitrary equal split. This option sounds quantitative but hides a subjective decision inside an unjustified formula, making the choice less defensible than a deliberate weighting.
- ✗
Ship Variant A because exact-match accuracy is an objective metric and human ratings are subjective.
Why it's wrong here
Exact-match accuracy is objective but often misaligned with user-perceived quality in generative tasks, where phrasing can vary while meaning is preserved. Discarding human preference entirely ignores the actual product goal of helpfulness and formatting. The scenario explicitly notes reviewers prefer Variant B, so ignoring that signal risks shipping a model users find less useful.
- ✗
Ship Variant B because human preference is the only metric that matters for generative models.
Why it's wrong here
Human preference is important, but treating it as the sole criterion ignores the risk of rater bias, small sample sizes, and regressions on factual correctness that exact-match can reveal. A balanced evaluation combines automated and human signals rather than discarding either. This option overcorrects in the opposite direction from the accuracy-only approach.
About these practice questions
This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.