AI0-001 AI Models and Data Engineering Practice Question
A data scientist is preparing a dataset for a machine learning model and notices that one feature has a range from 0 to 1,000,000, while another feature ranges from 0 to 1. The model to be used is a k-nearest neighbors (KNN) classifier. Which preprocessing step is MOST important to apply before training?
⚠ Common exam trap
The trap here is assuming that dimensionality reduction or outlier removal is the primary fix, when the core issue is the disparity in feature scales that directly affects distance-based algorithms like KNN.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Feature scaling, such as min-max normalization or standardization.
KNN uses distance metrics like Euclidean distance, which are sensitive to feature scales. A feature ranging from 0 to 1,000,000 will have a much larger impact on distance than one ranging from 0 to 1, effectively drowning out the smaller feature. Applying feature scaling, such as min-max normalization or standardization, ensures that all features contribute proportionally to the distance, leading to a more accurate and balanced model.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Principal component analysis (PCA) to reduce dimensionality.
Why it's wrong here
PCA can help with dimensionality reduction and may implicitly scale data if using correlation matrix, but it is not the most direct or necessary step. The primary problem is scale disparity, which should be addressed explicitly with feature scaling. PCA might be applied after scaling, but it is not the first step to fix the immediate issue.
- ✗
One-hot encoding for all features to convert them into binary vectors.
Why it's wrong here
One-hot encoding is used for categorical variables, not for numerical features with different scales. Applying it to numerical features would create a very high-dimensional sparse representation and lose the ordinal information. It does not address the scale disparity, which is the primary issue for KNN.
- ✗
Removing outliers from the feature with the larger range.
Why it's wrong here
While outliers can affect KNN, the main issue here is the difference in scale between features, not outliers. Removing outliers might help but does not solve the fundamental problem that the larger-range feature dominates distance calculations. Scaling is the appropriate and essential preprocessing step.
- ✓
Feature scaling, such as min-max normalization or standardization.
Why this is correct
KNN relies on distance calculations between data points. If one feature has a much larger range than others, it will dominate the distance metric, making the model effectively ignore the smaller-range features. Scaling ensures all features contribute equally to distance computations, which is essential for KNN to perform well.
About these practice questions
Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.