Courseiva

MLA-C01 Data Preparation for Machine Learning Practice Question

A company collects sensor data from IoT devices. The data arrives with missing timestamps due to network issues. For anomaly detection, the engineer needs to create features that capture rolling statistics over fixed windows. Which data preprocessing step is essential before feature generation?

⚠ Common exam trap

AWS often tests the distinction between handling missing values (imputation) and handling irregular timestamps (resampling), leading candidates to confuse forward-fill as a solution for time alignment when it only addresses missing data points, not the underlying time index irregularity.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Resample data to a fixed frequency

Resampling the data to a fixed frequency is essential because rolling window statistics require a consistent time index to compute accurate aggregations over fixed windows. Without a uniform timestamp grid, the window boundaries become ambiguous and the resulting features will be misaligned or incomplete, undermining the anomaly detection model.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Remove missing timestamps

    Why it's wrong here

    Deleting rows with missing timestamps discards valid sensor readings and breaks the continuous series that fixed-window rolling statistics require, biasing anomaly detection. Timestamp imputation or resampling preserves the series. Removal suits datasets where unlabelled or corrupt records must simply be excluded before training.

  • ✓

    Resample data to a fixed frequency

    Why this is correct

    Rolling statistics require evenly spaced observations; irregular timestamps from dropped network packets make window boundaries ambiguous. Resampling to a fixed frequency creates uniform intervals, so rolling means and variances are computed over consistent, comparable windows.

  • ✗

    Sort data by device ID

    Why it's wrong here

    Sorting by device ID groups each device's readings but leaves the irregular time spacing untouched, so fixed-window rolling statistics still compute over wrong intervals. Sorting by timestamp is the essential step. Device-ID sorting suits per-device partitioning or grouping operations, not time-window feature generation.

  • ✗

    Impute missing values with forward fill

    Why it's wrong here

    Forward fill propagates the last known value across gaps, fabricating readings and distorting the rolling mean, variance and rate features that anomaly detection depends on. Timestamp reconstruction is what the scenario needs. Forward fill suits categorical or slowly changing state fields, not numeric sensor measurements.

About these practice questions

This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.