Courseiva

MLA-C01 Data Preparation for Machine Learning Practice Question

A data engineer is preparing a dataset in Amazon SageMaker Data Wrangler for a binary classification model. The dataset contains missing values in several numeric columns, and the engineer wants a reusable, reproducible transformation that can be applied identically to the training data and to future inference data. The engineer plans to export the transformation and integrate it into a SageMaker Pipeline. Which TWO actions should the engineer take to ensure the imputation is consistent between training and inference? (Choose two.)

⚠ Common exam trap

The trap here is treating imputation as a one-time data cleanup rather than a parameterized transformation whose learned statistics must be persisted and reapplied at inference.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Configure the imputation transform in the Data Wrangler flow so it learns the fill statistic (such as the mean) from the training data during the flow run.

Consistency between training and inference requires that the imputation statistic be learned from training data and that the same transformation be reused at inference. Configuring Data Wrangler to learn the fill statistic from training data captures the parameters, and exporting the flow into a SageMaker Pipeline processing step ensures those exact parameters and logic are applied to both training and inference inputs.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Configure the imputation transform in the Data Wrangler flow so it learns the fill statistic (such as the mean) from the training data during the flow run.

    Why this is correct

    Learning the fill statistic from the training data ensures the imputation parameters are derived from the correct distribution and not from inference data, which would leak information. Data Wrangler captures these learned parameters as part of the flow, making the transformation reproducible. This is the foundation for applying the same imputation consistently at inference time.

  • ✗

    Replace missing numeric values with zero in both the training and inference datasets using a custom Python script.

    Why it's wrong here

    Filling with a constant zero is a valid but arbitrary choice that can distort distributions and does not learn from the training data. A custom script run separately on training and inference risks divergence if the logic is edited between runs. It also bypasses Data Wrangler's captured parameters, undermining reproducibility.

  • ✓

    Export the Data Wrangler flow to a SageMaker Pipeline and include the generated processing step so it runs on both training and inference inputs.

    Why this is correct

    Exporting the flow into a SageMaker Pipeline turns the transforms, including the learned imputation parameters, into a processing step that can be reused. Running that same step on inference data applies the identical fill statistics and transformation logic. This guarantees consistency between training and inference rather than relying on manual re-implementation.

  • ✗

    Compute the mean of each numeric column on the full dataset including inference data, then hard-code those values into the pipeline.

    Why it's wrong here

    Including inference data when computing imputation statistics causes data leakage and makes the pipeline dependent on the inference batch, which is not available at training time. Hard-coding values breaks reproducibility if the training data changes. This approach also cannot scale to a stream of new inference records.

  • ✗

    Delete all rows that contain missing values from the training dataset before training.

    Why it's wrong here

    Dropping rows with missing values discards potentially useful data and can bias the model if missingness is non-random. It also does not produce any reusable imputation logic to apply at inference time, so an inference record with a missing value would still need handling. This action reduces the dataset without ensuring consistency.

About these practice questions

This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.