NCA-GENL Core Machine Learning and AI Knowledge Practice Question
A team is pre-training a 7-billion-parameter LLM on a large text corpus. They observe that the training loss decreases steadily but the validation loss begins to increase after a certain number of steps. The training and validation data come from the same distribution, and the model has not yet reached the compute budget. Which action is most appropriate to address this behavior?
⚠ Common exam trap
The trap here is interpreting rising validation loss as a need for more compute or a larger model, when it actually indicates overfitting and calls for regularization or early stopping.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Apply regularization techniques such as dropout or weight decay, and consider early stopping.
A widening gap between decreasing training loss and increasing validation loss signals overfitting. The model is memorizing training-specific patterns rather than generalizing. Applying regularization like dropout or weight decay reduces this tendency, and early stopping captures the best validation checkpoint. These steps are standard and directly target the observed behavior without unnecessary architectural changes.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Reduce the model size to match the dataset size.
Why it's wrong here
Reducing model size is a drastic architectural change that may underfit the data and discard useful capacity. Overfitting can often be mitigated with regularization and early stopping before resorting to a smaller model. For a 7B-parameter LLM, shrinking the model is not the first or most appropriate action when regularization has not been tried.
- ✗
Increase the learning rate to escape the local minimum.
Why it's wrong here
Increasing the learning rate when validation loss is rising would likely worsen overfitting by making the model fit training noise more aggressively. The divergence between training and validation loss indicates the model is memorizing training-specific patterns, not stuck in a local minimum. A higher learning rate could destabilize training further and is not the right response.
- ✗
Collect more training data from the same distribution.
Why it's wrong here
Adding more data from the same distribution can help, but it is not the most immediate or practical action when the team already has a large corpus and is mid-training. Regularization and early stopping directly address the overfitting without requiring new data collection, which may be costly or time-consuming. The question asks for the most appropriate action given the current setup.
- ✓
Apply regularization techniques such as dropout or weight decay, and consider early stopping.
Why this is correct
The described pattern—training loss decreasing while validation loss increases—is classic overfitting. Adding dropout or weight decay constrains the model's capacity to memorize training data, and early stopping halts training at the point of best validation performance. These are standard, effective remedies when validation loss diverges despite ample compute budget remaining.
About these practice questions
One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.