MLA-C01 SageMaker Autopilot Practice Question
A data scientist is using SageMaker Autopilot for a regression problem. They want to see which data preprocessing steps Autopilot applied. Which TWO sources can they use to find this information?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Candidate definition notebook
The candidate definition notebook (option A) is generated by SageMaker Autopilot for each candidate and contains the full ML pipeline code, including the exact data preprocessing and feature engineering transforms that were applied, so it directly answers the question. The data exploration report (option D) is produced during the Autopilot job's data exploration phase and documents the dataset's characteristics along with the preprocessing and feature engineering steps Autopilot selected, making it another valid source. The model leaderboard (option B) only ranks trained candidates by objective metric and does not describe preprocessing. The Autopilot job description in AWS CloudTrail (option C) records API-level audit events, not the internal preprocessing steps. The explainability report (option E) covers feature attributions and model behavior, not the preprocessing pipeline.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Candidate definition notebook
Why this is correct
The candidate definition notebook documents the exact preprocessing pipeline Autopilot generated for that candidate, including transforms applied to features before training. It satisfies the requirement to inspect applied preprocessing by exposing the reproducible code rather than only summary statistics.
- ✗
Model leaderboard
Why it's wrong here
The model leaderboard ranks candidate models by objective metric; it does not document the preprocessing steps applied to features. It is tempting because it is a primary Autopilot output, but it would be correct when comparing candidate model performance rather than inspecting transformation detail.
- ✗
Autopilot job description in AWS CloudTrail
Why it's wrong here
CloudTrail records API calls for auditing, not the candidate definitions or preprocessing transformations Autopilot applied. It is tempting because the job description sounds like configuration detail, but CloudTrail would be correct for tracking who invoked the Autopilot job and when.
- ✓
Data exploration report
Why this is correct
The data exploration report summarises each candidate's preprocessing, including imputation and encoding choices, generated during the Autopilot job. It satisfies the requirement to inspect applied preprocessing by exposing the transformations alongside data statistics and feature distributions.
- ✗
Explainability report
Why it's wrong here
The explainability report covers feature attributions and model behaviour, not the data preprocessing pipeline. It is tempting because it is an Autopilot-generated artefact, but it would be correct when the goal is interpreting which features drove predictions rather than inspecting transformations.
Go deeper
Related to this question
About these practice questions
This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.