AI0-001 AI Models and Data Engineering Practice Question
Which TWO data preprocessing techniques reduce the dimensionality of a dataset?
⚠ Common exam trap
CompTIA often tests the distinction between techniques that transform or select features (reducing dimensionality) versus those that prepare data for modeling (like encoding, imputation, or scaling) without changing the number of features.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Principal Component Analysis (PCA)
Principal Component Analysis (PCA) (D) is correct because it projects the original features onto a smaller set of orthogonal principal components that capture most of the variance, thereby reducing the number of dimensions while preserving as much information as possible. Feature selection (E) is also correct because it explicitly removes irrelevant or redundant features and keeps only a subset of the original variables, directly lowering the dataset's dimensionality. By contrast, one-hot encoding (A) increases dimensionality by creating a separate binary column per category, and imputation (B) merely fills in missing values without changing the number of features. Feature scaling (C) standardizes or normalizes feature values (e.g., via min-max or z-score) but leaves the feature count unchanged, so it does not reduce dimensionality.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
One-hot encoding
Why it's wrong here
One-hot encoding expands each categorical feature into multiple binary columns, increasing dimensionality rather than reducing it. It is tempting because it is a standard preprocessing technique, and it is the correct choice when categorical variables must be converted into a numeric representation for algorithms that cannot handle raw categories.
- ✗
Imputation
Why it's wrong here
Imputation fills missing values with estimates such as the mean or median, preserving the existing feature count, so dimensionality is unchanged. It is tempting because it is a common preprocessing step, and it is the right choice when a dataset contains nulls that would otherwise break model training.
- ✗
Feature scaling
Why it's wrong here
Feature scaling rescales numeric values onto a common range, leaving the number of features unchanged, so dimensionality is untouched. It is tempting because it is a standard preprocessing step, and it is the right choice when algorithms such as k-NN or SVM are sensitive to differing feature magnitudes.
- ✓
Principal Component Analysis (PCA)
Why this is correct
PCA projects the original features onto orthogonal principal components ordered by explained variance, so the dataset can be represented with far fewer dimensions while retaining most information. This directly reduces dimensionality rather than merely selecting or transforming individual columns.
- ✓
Feature selection
Why this is correct
Feature selection removes irrelevant or redundant attributes, leaving a smaller subset of the original variables. This reduces dimensionality without creating new features, unlike projection-based methods such as PCA, and directly lowers the dataset's column count.
About these practice questions
Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.