AI0-001 AI Models and Data Engineering Practice Question
Which THREE are common data preprocessing steps in a machine learning pipeline? (Choose 3)
⚠ Common exam trap
CompTIA often tests the distinction between preprocessing steps (data cleaning, transformation) and later pipeline stages (model tuning, evaluation), so candidates mistakenly select hyperparameter tuning or model evaluation as preprocessing steps.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Encoding categorical variables
Encoding categorical variables (B) is a standard preprocessing step because algorithms require numeric input, so techniques like one-hot encoding or label encoding convert strings into numbers. Scaling numeric features (D) is also preprocessing, since methods such as standardization (z-score) or min-max normalization bring features to comparable ranges, which helps distance-based and gradient-based models. Handling missing values (E) is preprocessing too, using imputation (mean, median, mode) or deletion to ensure the dataset is complete before training. Hyperparameter tuning (A) and model evaluation (C) are not preprocessing; they occur later in the pipeline, during model selection/training and after training, respectively.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Hyperparameter tuning
Why it's wrong here
Hyperparameter tuning optimises model configuration after training, selecting values such as learning rate or tree depth; it operates on the model, not the raw dataset. It is tempting because it sits in the same pipeline workflow, but it would be the right answer to a question about model selection or training optimisation, not preprocessing.
- ✓
Encoding categorical variables
Why this is correct
Encoding categorical variables converts non-numeric labels into numeric representations, such as one-hot or ordinal encoding, so algorithms that require numerical input can process them. This satisfies the preprocessing requirement by making categorical features usable during model training, alongside handling missing values and feature scaling.
- ✗
Model evaluation
Why it's wrong here
Model evaluation scores a trained model's performance on held-out data, which happens after preprocessing and training. It is tempting because it is a genuine pipeline stage, but it would be correct in a question about assessing generalisation or choosing between candidate models, not about cleaning or transforming input data.
- ✓
Scaling numeric features
Why this is correct
Scaling transforms numeric features onto a comparable range, such as standardisation or min-max normalisation, preventing features with larger magnitudes from dominating distance-based or gradient-based algorithms. This is a core preprocessing step applied before model fitting.
- ✓
Handling missing values
Why this is correct
Handling missing values removes or imputes incomplete entries so algorithms receive a complete matrix, since many implementations cannot process nulls. This preprocessing step prevents training failures and bias introduced by silently dropped or malformed records.
About these practice questions
This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.