MLS-C01 Exploratory Data Analysis Practice Question
A data scientist is performing exploratory data analysis on a dataset stored in Amazon S3 using Amazon SageMaker Studio. The dataset has missing values in several columns. Which approach is the MOST efficient way to handle missing values within SageMaker Studio?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use SageMaker Data Wrangler to impute missing values with mean, median, or mode.
SageMaker Data Wrangler provides a visual interface to handle missing values efficiently within SageMaker Studio, allowing imputation with mean, median, or mode without writing custom code. Option A is inefficient because it requires moving data out of SageMaker. Option C uses an external service (AWS Glue) which adds complexity and overhead. Option D, while possible, is less efficient than using Data Wrangler's built-in capabilities.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Run a Jupyter notebook on a local machine to clean the data and upload back to S3.
Why it's wrong here
Not efficient as it involves manual steps.
- ✓
Use SageMaker Data Wrangler to impute missing values with mean, median, or mode.
Why this is correct
Data Wrangler provides a visual interface for imputation.
- ✗
Use AWS Glue to run a find-and-replace operation.
Why it's wrong here
Glue is external and less integrated.
- ✗
Write a custom Python script using pandas to drop rows with missing values.
Why it's wrong here
Less efficient than using built-in tools.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This MLS-C01 question is part of Courseiva's 1,672-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLS-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLS-C01 exam.