NCA-GENL Experimentation Practice Question
A team is comparing two fine-tuned variants of the same base LLM for a customer-support summarization task. Variant A was trained on 10,000 examples and variant B on 2,000 examples drawn from the same distribution. To make a fair comparison of their generalization, which evaluation practice should the team adopt?
⚠ Common exam trap
The trap here is treating training loss as a proxy for quality, when it mostly reflects how much data the model memorized rather than how well it generalizes.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Evaluate both variants on the same held-out test set that neither saw during training
Fair comparison requires controlling the evaluation data. When both variants are scored on the same held-out set that neither encountered during training, differences in their scores reflect the training choices rather than variation in test difficulty. Training loss and parameter counts are process metrics that do not measure generalization to unseen customer-support summaries.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Evaluate both variants on the same held-out test set that neither saw during training
Why this is correct
Holding the evaluation data constant isolates the effect of the training-set size from the effect of the evaluation data. If each variant were scored on a different test set, differences in difficulty or domain mix would confound the comparison. A shared, untouched held-out set gives both models the same challenge, so observed differences can be attributed to training choices rather than evaluation noise.
- ✗
Report the training loss of each variant at the final step
Why it's wrong here
Training loss measures fit to data the model has already seen and is heavily influenced by dataset size. The variant trained on fewer examples will typically show higher training loss regardless of which generalizes better, so this metric cannot answer the comparison question and may actively mislead the team into preferring the wrong variant.
- ✗
Compare the number of parameters each variant updated during fine-tuning
Why it's wrong here
Parameter counts describe the training procedure, not the resulting quality on the summarization task. Two variants could update the same number of parameters yet differ substantially in output quality. This metric offers no signal about generalization and cannot substitute for measuring task performance on data the models have not memorized.
- ✗
Let each variant be scored on its own randomly split test set
Why it's wrong here
Separate random splits create different evaluation samples with different difficulty and topic composition. A model could appear superior simply because its test split happened to contain easier or more repetitive summaries. To attribute performance differences to the training regime, both variants must face an identical, fixed evaluation set.
About these practice questions
This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.