NCA-GENL Core Machine Learning and AI Knowledge Practice Question
A data scientist is training a transformer model and observes that the training loss is decreasing while the validation loss is increasing. Which technique should be prioritized to address this specific generalization challenge?
⚠ Common exam trap
Candidates often suggest increasing model size or training epochs to fix the loss gap, failing to recognize that these actions typically exacerbate overfitting rather than solving the underlying generalization problem.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Implement L2 regularization (weight decay) on the model weights.
This scenario indicates overfitting, where the model captures noise in the training set rather than the underlying distribution. Regularization techniques like Dropout, weight decay, or early stopping are essential to improve model generalization. By limiting the model's ability to memorize the training data, these methods ensure that the weights remain small and the model learns robust patterns that perform better on unseen, held-out validation datasets.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the depth of the transformer architecture layers.
Why it's wrong here
Adding more layers increases the model capacity, which typically exacerbates overfitting in scenarios where validation loss is already diverging. Complex models require more data to converge properly; increasing depth without sufficient data will likely widen the gap between training and validation performance, worsening the existing generalization issue.
- ✓
Implement L2 regularization (weight decay) on the model weights.
Why this is correct
L2 regularization penalizes large weights by adding a cost term proportional to the square of the magnitude of coefficients. This forces the model to learn simpler patterns, significantly reducing the variance of the model and effectively mitigating overfitting, which is the primary cause of the divergent loss curves described.
- ✗
Reduce the batch size to increase stochasticity during training.
Why it's wrong here
While smaller batch sizes introduce noise that can help escape local minima, they do not inherently solve overfitting. If the model is already overfitting, increasing the noise level through smaller batches might hinder convergence or make training unstable without directly addressing the model's failure to generalize to validation data.
- ✗
Decrease the number of training epochs to save time.
Why it's wrong here
Simply stopping early without a validation-based strategy is ineffective. While decreasing epochs is part of early stopping, it must be triggered by validation metrics rather than an arbitrary time constraint. Without a systematic approach, one might stop before the model has learned the relevant features required for good performance.
About these practice questions
Courseiva writes every NCA-GENL question from scratch — 367 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.