AI0-001 AI Concepts and Techniques Practice Question
A developer is fine-tuning a large language model for a legal document summarization task. They notice that during training, the loss decreases rapidly in the first few epochs but then plateaus with high variance. Which hyperparameter adjustment is MOST likely to help stabilize training?
⚠ Common exam trap
CompTIA often tests the misconception that high variance in loss is always solved by increasing batch size or regularization, when in fact the immediate cause is often an overly aggressive learning rate that prevents convergence.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Decrease the learning rate
A high-variance loss plateau after rapid initial convergence typically indicates that the learning rate is too large, causing the optimizer to overshoot the minima and oscillate. Decreasing the learning rate allows smaller, more stable weight updates, reducing variance and enabling smoother convergence.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Add L1 regularization
Why it's wrong here
L1 regularization penalises large weights to encourage sparsity, which does not target the variance causing the plateau. It is tempting because regularisation is a standard remedy for overfitting, and would be correct if the model were memorising training data and generalising poorly to validation documents.
- ✓
Decrease the learning rate
Why this is correct
A learning rate that is too high causes the optimiser to overshoot minima, producing the plateau with high variance seen after the initial rapid loss drop. Lowering it reduces update step size, letting the model settle into a smoother minimum and stabilising training.
- ✗
Increase the batch size
Why it's wrong here
Increasing batch size reduces gradient noise per update but does not address the plateau itself, and can slow convergence further. It is tempting because larger batches do stabilise noisy gradients, and would be correct if the loss were erratic from the outset rather than plateauing after rapid early descent.
- ✗
Increase the number of epochs
Why it's wrong here
More epochs extend training but cannot resolve high variance; the model simply continues oscillating around the plateau. It is tempting because additional training often improves underfit models, and would be correct if the loss were still declining steadily rather than fluctuating with high variance.
About these practice questions
This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.