Courseiva
hardMultiple Choice

MLA-C01 Practice Question: A data science team is using Amazon SageMaker…

A data science team is using Amazon SageMaker Pipelines to orchestrate a multi-step workflow that includes data preprocessing, training, and model evaluation. They want to reuse the preprocessed data across multiple pipeline executions without re-running the preprocessing step if the source data hasn't changed. What should they configure?

⚠ Common exam trap

A common mix-up: candidates confuse checkpointing (for training resumption) with caching (for step reuse), or assume that Feature Store or Data Wrangler inherently provide caching, when in fact only Processing steps with explicit CacheConfig enable this behavior in SageMaker Pipelines.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use SageMaker Processing steps with caching

SageMaker Processing steps support caching, which allows the pipeline to skip re-execution of the preprocessing step if the input data and pipeline parameters have not changed. This is achieved by configuring a `CacheConfig` with a caching key based on the input data source and step parameters, ensuring that the preprocessed data is reused across multiple pipeline executions without redundant computation.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use SageMaker Training steps with checkpointing

    Why it's wrong here

    Training-step checkpointing saves model state during training so interrupted jobs resume, not preprocessed data reuse across executions. It is tempting because checkpointing genuinely suits long training runs recovering from failures, but the stem concerns skipping an unchanged preprocessing step, which pipeline step caching handles.

  • ✓

    Use SageMaker Processing steps with caching

    Why this is correct

    SageMaker Pipelines caches step outputs keyed on the step's code, inputs, and parameters, so unchanged source data lets downstream executions skip preprocessing and reuse the stored artefacts. This directly satisfies the requirement to avoid re-running preprocessing across executions, while training and evaluation still run normally.

  • ✗

    Use SageMaker Feature Store to store the preprocessed features

    Why it's wrong here

    Feature Store persists engineered features for online and offline serving, but it does not automatically skip a pipeline's preprocessing step when source data is unchanged. It is tempting because Feature Store genuinely suits feature reuse across models and teams, yet the requirement is execution-level step caching, configured on the pipeline step itself.

  • ✗

    Use SageMaker Data Wrangler for the preprocessing

    Why it's wrong here

    Data Wrangler transforms data interactively and exports flows, but it does not cache step outputs across pipeline executions, so preprocessing reruns each time. It is tempting because Data Wrangler genuinely suits exploratory feature preparation, yet the requirement is step-level caching keyed on unchanged inputs, which SageMaker Pipelines caching configuration provides.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.