Courseiva
Experimentation →hardMultiple Choice

NCA-GENL Experimentation Practice Question

A team runs a NeMo fine-tuning experiment and observes that validation loss decreases for several epochs and then steadily rises while training loss keeps falling. They want to confirm whether the checkpoint from the best validation epoch is genuinely better than the final checkpoint. Which action provides the strongest evidence?

⚠ Common exam trap

The trap here is reusing the validation metric that drove checkpoint selection as if it were an independent judge, which inflates confidence in the selected model.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Evaluate both checkpoints on a separate held-out test set and compare task metrics

When a checkpoint is selected using validation loss, that metric becomes optimistically biased for the chosen model. The strongest evidence comes from an untouched test set that was not involved in any selection decision. Evaluating both checkpoints there with task-relevant metrics yields an unbiased comparison, whereas training loss and repeated validation checks cannot settle the question.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Compare the training loss values of the two checkpoints

    Why it's wrong here

    Training loss measures fit to data the model has already optimized against, so the later checkpoint will almost always show lower training loss. This is exactly the pattern described and tells the team nothing about generalization. Using training loss to choose between checkpoints would systematically favor the overfit model, which is the opposite of what is needed.

  • ✓

    Evaluate both checkpoints on a separate held-out test set and compare task metrics

    Why this is correct

    Validation loss guided model selection, so it can be optimistically biased for the checkpoint chosen by that same metric. A separate held-out test set that played no role in selecting the checkpoint gives an unbiased estimate of generalization. Comparing task metrics on this set directly answers whether the earlier checkpoint is truly better rather than an artifact of selection.

  • ✗

    Continue training the final checkpoint for more epochs and recheck validation loss

    Why it's wrong here

    Extending training does not compare the two existing checkpoints; it creates a new model and changes the question. If the model is overfitting, additional epochs are likely to worsen validation loss further. This action also consumes compute without producing the unbiased comparison the team needs to decide which checkpoint to deploy.

  • ✗

    Pick whichever checkpoint has the lower validation loss

    Why it's wrong here

    The rising validation loss already indicates the earlier epoch is favored by that metric, so this merely restates the selection rule. Because the checkpoint was chosen using validation loss, that same number is biased in its favor and cannot independently confirm superiority. An unbiased estimate requires data not used for selection.

About these practice questions

One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.