Courseiva

Best Practices for Preprocessing Data in Machine Learning

Which TWO of the following are best practices for preparing training data for a machine learning model?

⚠ Common exam trap

The AIF-C01 exam often tests the misconception that removing all outliers is always beneficial, when in fact domain knowledge is required to distinguish between noise and legitimate extreme values that may be critical for model accuracy.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Handle missing values by imputing or removing them.

Option A is correct because missing values can bias or break many ML algorithms, so best practice is to address them explicitly—either by imputation (e.g., mean/median/mode or model-based) or by removing affected rows/columns when appropriate. Option B is correct because splitting data into training, validation, and test sets enables unbiased model fitting, hyperparameter tuning, and final performance estimation, preventing data leakage and overfitting. Option C is not a best practice because outliers may be legitimate signal; blindly removing them can distort the distribution and reduce robustness, so they should be investigated and handled contextually. Option D is not a best practice because training on the entire dataset leaves no held-out data for validation or testing, making it impossible to reliably estimate generalization. Option E is not a best practice because shuffling is typically needed to remove ordering bias and ensure representative mini-batches, especially when data is sorted by class or time.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Handle missing values by imputing or removing them.

    Why this is correct

    Missing values break many algorithms and bias estimates, so imputing sensible substitutes or removing affected rows prevents errors and skewed learning. This is a standard preparation step ensuring the training set is complete and representative before model fitting.

  • ✓

    Split the data into training, validation, and test sets.

    Why this is correct

    Holding out validation and test partitions lets you tune hyperparameters and estimate generalisation on data the model never saw, guarding against overfitting. Training on the full dataset without splits would inflate reported accuracy and hide poor real-world performance.

  • ✗

    Remove all outliers to improve model robustness.

    Why it's wrong here

    Outliers frequently represent genuine fraud or failure cases, so deleting them strips exactly the signal a financial model must detect. Removing them is tempting because extreme values distort means and distances, and it is valid only when an outlier is a confirmed measurement or entry error rather than a real event.

  • ✗

    Use the entire dataset for training to maximize data usage.

    Why it's wrong here

    Holding back a validation or test split is what exposes overfitting; training on every record leaves no unseen data to estimate generalisation error. Using the full dataset is tempting because more samples often help, and it is legitimate only after hyperparameters are frozen, when refitting on all data before deployment.

  • ✗

    Avoid shuffling the data to preserve original order.

    Why it's wrong here

    Shuffling breaks any ordering correlation between consecutive samples, so a model trained on sorted data can learn sequence artefacts rather than the underlying pattern. Preserving order is tempting for time-series or sequential data, where chronological splits matter, but even then shuffling within each split is standard.

About these practice questions

One of 862 original AIF-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.