MLA-C01 Data Preparation for Machine Learning Practice Question
Which TWO actions are recommended best practices when preparing training data for a machine learning model in AWS? (Choose two.)
⚠ Common exam trap
Test-takers frequently assume all outliers must be removed (Option A) or that normalization is always required (Option E), but the exam tests nuanced understanding that these steps depend on the algorithm and data characteristics, not blanket rules.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Check for and handle missing values appropriately.
Option C is correct because missing values can bias or break training algorithms, so best practice is to detect them (e.g., with pandas isnull() or Amazon SageMaker Data Wrangler) and handle them via imputation, removal, or model-native handling. Option D is correct because splitting data into training, validation, and test sets lets you fit parameters, tune hyperparameters, and estimate generalization performance on unseen data, avoiding overfitting and data leakage. Option A is not recommended because not all outliers are errors; blindly removing them can discard legitimate signal and distort the distribution. Option B is wrong because training on the entire dataset leaves no held-out data for validation or unbiased testing. Option E is wrong because normalization is not always required and [0,1] scaling is only one of several techniques (e.g., standardization, log transforms) chosen based on the algorithm and feature distribution.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Remove all outliers from the dataset.
Why it's wrong here
Outliers can represent genuine rare events or fraud signals; indiscriminate removal biases the model and discards valid information. It is tempting because outliers distort some statistical measures, but removal is warranted only after investigation confirms they are errors, not as a blanket rule.
- ✗
Train the model on the entire dataset to maximize data usage.
Why it's wrong here
Training on the entire dataset leaves no held-out data to evaluate generalisation, risking undetected overfitting. It is tempting because more data often improves learning, but a validation or test split is needed to measure performance on unseen examples before deployment.
- ✓
Check for and handle missing values appropriately.
Why this is correct
Missing values bias model training and can cause failures in algorithms that reject nulls. Detecting and handling them—via imputation, removal, or indicator flags—preserves data integrity, satisfying the best-practise requirement for robust training data preparation in AWS pipelines such as SageMaker Data Wrangler or Glue.
- ✓
Split the data into training, validation, and test sets.
Why this is correct
Holding out validation and test partitions prevents evaluating the model on data it memorised during training. This separation enables unbiased hyperparameter tuning and final performance estimation, directly satisfying the best-practise requirement for honest generalisation assessment before deployment.
- ✗
Always normalize all features to a [0,1] range.
Why it's wrong here
Normalising every feature to [0,1] discards the distribution shape and outlier information that tree-based and distance-based models may need, and is unnecessary for models such as random forests. It is tempting because scaling genuinely helps gradient descent and distance metrics, but only selected numeric features should be scaled, not all of them unconditionally.
Go deeper
Related to this question
About these practice questions
This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.