Courseiva

MLA-C01 Data Preparation for Machine Learning Practice Question

A company is building a real-time inference pipeline for an ML model. The raw data arrives in JSON format via Amazon Kinesis Data Streams. Before invoking the SageMaker endpoint, the data must be preprocessed to match the training data format. Which THREE steps should be included in the preprocessing function? (Select THREE)

⚠ Common exam trap

Candidates often confuse the preprocessing function's scope with broader MLOps tasks like model retraining, or assume a specific serialization format like CSV is required when JSON is natively supported by SageMaker endpoints.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Ensure that missing values are handled consistently with the training phase

The preprocessing function must first parse the JSON payload (E), because the raw records arrive from Kinesis Data Streams in JSON format and the individual feature fields cannot be accessed until the JSON is decoded into a usable structure. It must also apply the same feature engineering transformations used during training (C), such as scaling and encoding, since the SageMaker endpoint expects inputs in the exact distribution and representation the model learned; using different transformations would cause training-serving skew and degrade predictions. Missing values must be handled consistently with the training phase (A), because imputation or drop logic applied at inference must mirror training-time behavior to keep the feature semantics identical. Option B is not required because the model's input format is not stated to be CSV, and forcing a CSV string could conflict with the actual serialization the endpoint expects. Option D is incorrect because retraining the model is a separate MLOps activity, not part of the per-record preprocessing function invoked before calling the endpoint.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Ensure that missing values are handled consistently with the training phase

    Why this is correct

    Handling missing values identically to training preserves the feature distribution the model learned, preventing inference-time skew. Because Kinesis delivers raw JSON that may contain nulls, this step satisfies the requirement that preprocessed data match the training format before the SageMaker endpoint is invoked.

  • ✗

    Convert the data to a CSV string for model input

    Why it's wrong here

    CSV conversion is unnecessary; the preprocessing function must parse JSON, apply the training-time feature transformations, and serialise to the format the SageMaker endpoint expects, typically JSON. It is tempting because CSV is a common training input, correct when the model was trained on CSV.

  • ✓

    Apply the same feature engineering transformations (e.g., scaling, encoding) that were used during training

    Why this is correct

    Reapplying the training-time scaling and encoding ensures inference features occupy the same numerical space the model was fitted on, avoiding training-serving skew. This satisfies the stem's constraint that preprocessed Kinesis records match the training data format before invoking the SageMaker endpoint.

  • ✗

    Re-train the model periodically using new data

    Why it's wrong here

    Re-training is a separate offline process, not part of inference preprocessing.

  • ✓

    Parse the JSON payload

    Why this is correct

    Kinesis Data Streams delivers records as base64-encoded JSON, so the payload must be decoded and parsed into fields before any feature transformation can occur. Parsing satisfies the pipeline's first requirement, converting raw JSON into structured values the preprocessing function can then align with the training format.

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.