NCA-GENL Core Machine Learning and AI Knowledge Practice Question
A team trains a transformer language model on a large corpus but the model achieves very low training loss while performing poorly on held-out text. Which action most directly addresses this outcome?
⚠ Common exam trap
The trap here is reacting to poor validation results by training longer or tuning the learning rate, when the low training loss already proves the model is memorizing rather than underfitting.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Apply dropout and weight decay to regularize the model
Low training loss combined with weak held-out performance indicates overfitting, meaning the model memorizes training text instead of learning generalizable patterns. Regularization techniques such as dropout and weight decay constrain model capacity and penalize reliance on specific training examples, which directly reduces the train-validation gap and improves performance on unseen text.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Lower the learning rate for the final training steps
Why it's wrong here
Learning-rate scheduling affects optimization stability and convergence speed, not the fundamental capacity of the model to memorize. A model already overfitting will still overfit at a lower learning rate, just more slowly. This change does not introduce any constraint that discourages memorization of training sequences.
- ✗
Increase the number of training epochs on the same corpus
Why it's wrong here
Training longer on the same data would push training loss even lower while held-out performance stagnates or worsens, deepening the overfitting already evident. The gap between low training loss and poor validation performance signals the model is memorizing rather than generalizing, so additional epochs on identical data cannot close that gap.
- ✓
Apply dropout and weight decay to regularize the model
Why this is correct
A large gap between near-zero training loss and weak held-out performance is the classic signature of overfitting. Dropout randomly deactivates units during training, preventing co-adaptation, while weight decay constrains parameter magnitude. Together they reduce memorization of training text and improve generalization to unseen sequences, directly targeting the observed symptom.
- ✗
Reduce the size of the training corpus
Why it's wrong here
Shrinking the corpus removes the very data the model needs to learn generalizable language patterns and typically worsens held-out performance. The problem is not too much data but insufficient regularization relative to model capacity. Discarding examples does not fix overfitting in a principled way and may introduce additional bias.
About these practice questions
Courseiva writes every NCA-GENL question from scratch — 367 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.