Courseiva

MLA-C01 Data Preparation for Machine Learning Practice Question

A data engineer needs to prepare a dataset for a fraud detection model. The dataset contains a highly skewed numerical feature with extreme outliers. The engineer decides to apply a logarithmic transformation to this feature before training. Which SageMaker Data Wrangler transform should be used to apply the logarithmic transformation?

⚠ Common exam trap

Watch out — candidates often confuse scaling transforms with distribution-changing transforms; scaling does not alter skewness or outliers.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use the 'Log Transform' transform in Data Wrangler.

The logarithmic transformation is a common technique to reduce right skewness and stabilize variance in numerical data. In SageMaker Data Wrangler, the 'Log Transform' transform applies a natural logarithm to the selected column, effectively compressing the scale of large values and making the distribution more symmetric. This helps models that are sensitive to feature distributions, such as linear models, to perform better. Other transforms like scaling or encoding do not change the distribution shape.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use the 'Standard Scaler' transform in Data Wrangler.

    Why it's wrong here

    The 'Standard Scaler' transform standardizes features by removing the mean and scaling to unit variance. While it helps with features on different scales, it does not change the shape of the distribution or compress outliers. It would not address the skewness and extreme outliers described, making it unsuitable for this scenario.

  • ✗

    Use the 'One-Hot Encoding' transform in Data Wrangler.

    Why it's wrong here

    The 'One-Hot Encoding' transform is used for categorical variables, converting them into binary vectors. It is not applicable to numerical features and does not perform any mathematical transformation to reduce skewness or handle outliers. Therefore, it is incorrect for this numerical feature scenario.

  • ✓

    Use the 'Log Transform' transform in Data Wrangler.

    Why this is correct

    The 'Log Transform' transform in SageMaker Data Wrangler applies a natural logarithm (base e) to the selected numeric column. This is specifically designed to reduce right skewness and mitigate the impact of extreme outliers, making the feature more suitable for models that assume normality or are sensitive to scale. It directly addresses the scenario's need for a logarithmic transformation.

  • ✗

    Use the 'Min-Max Scaler' transform in Data Wrangler.

    Why it's wrong here

    The 'Min-Max Scaler' transform rescales features to a fixed range, typically [0, 1]. This is sensitive to outliers because the range is determined by the minimum and maximum values. Extreme outliers would compress the rest of the data into a small interval, potentially harming model performance. It does not reduce skewness.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.