AI0-001 Machine Learning and Deep Learning Practice Question
A dataset contains features on vastly different scales (e.g., age 0-100 vs. income 0-1,000,000). Which preprocessing step is essential before training a neural network?
⚠ Common exam trap
CompTIA AI often tests the misconception that data augmentation or dimensionality reduction can substitute for feature scaling, when in fact scaling is a prerequisite for stable gradient descent in neural networks.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Feature scaling (standardization or normalization)
Neural networks rely on gradient-based optimization, where features with larger scales can dominate the weight updates, causing unstable convergence or slow training. Feature scaling (standardization or normalization) ensures all features contribute equally to the loss function, preventing the model from being biased toward high-magnitude features like income versus age.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Data augmentation
Why it's wrong here
Data augmentation generates extra training examples (rotations, noise) and does not alter feature magnitudes, so income would still swamp age during gradient descent. It is tempting for image or text tasks with limited samples, but the stem's issue is scale disparity, not dataset size.
- ✗
Dimensionality reduction
Why it's wrong here
Dimensionality reduction compresses feature count (for example via PCA), leaving the age-versus-income scale disparity untouched, so large-magnitude features still dominate gradient updates. It is tempting because it aids high-dimensional or noisy data, but the stem requires rescaling, not fewer features.
- ✓
Feature scaling (standardization or normalization)
Why this is correct
Features on vastly different scales cause gradient updates to be dominated by large-magnitude inputs, slowing or destabilising training. Standardisation or normalisation rescales each feature to a comparable range, ensuring the network converges efficiently and weights are not biased toward high-range variables such as income.
- ✗
One-hot encoding
Why it's wrong here
One-hot encoding converts categorical variables into binary vectors; it cannot rescale continuous numeric features like age or income. It is tempting because it is a common preprocessing step, but it applies to nominal categories, whereas the stem needs normalisation or standardisation.
About these practice questions
Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.