AI0-001 AI Concepts and Foundations Practice Question
An AI model achieves high accuracy on training data but performs poorly on new test data. The data scientist suspects the model has memorized noise. Which technique directly adds a penalty term to the loss function to address this?
⚠ Common exam trap
CompTIA often tests the distinction between regularization techniques that modify the loss function (L2) versus those that modify the network architecture or data (dropout, batch normalization, data augmentation), so candidates mistakenly choose dropout because it is a well-known regularization method, even though it does not add a penalty term to the loss function.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
L2 regularization
L2 regularization (also known as weight decay) directly adds a penalty term proportional to the squared magnitude of the model's weights to the loss function. This discourages the model from fitting the noise in the training data by keeping weights small, thereby reducing overfitting and improving generalization to new test data.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Batch normalization
Why it's wrong here
Batch normalisation rescales layer activations to stabilise and accelerate training; it adds no penalty term to the loss function, so it does not directly counter memorised noise. It is tempting because it does improve generalisation somewhat, and would be the right choice for unstable or slow training in deep networks.
- ✗
Data augmentation
Why it's wrong here
Data augmentation expands the training set with transformed copies, which reduces overfitting indirectly but adds no penalty term to the loss function. It is tempting because it does improve generalisation, and would be correct when training data is scarce, particularly for image or text tasks.
- ✗
Dropout
Why it's wrong here
Dropout randomly deactivates neurons during training to prevent co-adaptation; it regularises structurally rather than adding an explicit penalty term to the loss. It is tempting because it does combat overfitting, and would be the right choice for large fully connected layers in deep neural networks.
- ✓
L2 regularization
Why this is correct
L2 regularization adds a penalty proportional to the squared magnitude of weights to the loss function, directly discouraging the large weight values that let a model memorise noise. This constrains model complexity, satisfying the stem's requirement for a technique that penalises the loss function to reduce overfitting.
About these practice questions
This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.