mediumMultiple Select
MLA-C01 Practice Question: A data engineer needs to assess the quality of a…
A data engineer needs to assess the quality of a dataset containing customer information. The dataset has missing values, outliers, and duplicate records. Which TWO AWS services can be used to perform data quality assessment? (Select TWO.)
⚠ Common exam trap
Watch out — candidates often confuse AWS Glue ETL (option C) with AWS Glue DataBrew (option B), assuming the ETL service includes visual data quality assessment, when in fact DataBrew is the dedicated no-code data preparation and quality tool.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
AWS Glue DataBrew
AWS Glue DataBrew (B) is correct because it is a visual data preparation service that provides over 250 built-in transformations and profiling features, including data quality checks such as detecting missing values, outliers, and duplicate records via its profile view and anomaly detection. Amazon SageMaker Data Wrangler (E) is correct because it offers data quality and insights reports that automatically surface missing values, outliers, duplicate rows, and target leakage, allowing a data engineer to assess dataset quality before ML workflows. Amazon Athena (A) is a serverless query service for analyzing data in S3, not a dedicated data quality assessment tool. AWS Glue ETL (C) is used for building and running extract-transform-load jobs rather than profiling data quality. Amazon QuickSight (D) is a business intelligence visualization service and does not provide built-in data quality assessment capabilities.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Amazon Athena
Why it's wrong here
Athena queries data in S3 but cannot profile columns, detect outliers or deduplicate records; it returns query results, not quality metrics. It is tempting because SQL can count nulls and duplicates, yet that requires hand-written queries rather than an assessment service such as AWS Glue DataBrew or Deequ.
- ✓
AWS Glue DataBrew
Why this is correct
AWS Glue DataBrew profiles datasets and surfaces missing values, outliers and duplicates through built-in statistics and anomaly detection, without writing code. Its data quality rules then validate these conditions at scale, directly satisfying the stem's requirement to assess customer data quality across all three defect types.
- ✗
AWS Glue ETL
Why it's wrong here
AWS Glue ETL transforms and moves data; its job scripts can clean records, but the service itself is not a quality assessment tool with profiling rules. It tempts because Glue DataBrew shares the Glue family, yet DataBrew, not ETL jobs, provides the rule-based quality evaluation required.
- ✗
Amazon QuickSight
Why it's wrong here
QuickSight is a BI visualisation service; it renders dashboards and charts rather than profiling nulls, outliers or duplicates programmatically. It tempts because visual inspection can reveal anomalies, but systematic quality assessment needs Glue DataBrew or Deequ-style rules, not presentation tooling.
- ✓
Amazon SageMaker Data Wrangler
Why this is correct
Amazon SageMaker Data Wrangler profiles datasets and generates quality reports that surface missing values, outliers and duplicate rows, satisfying the stem's assessment requirement. Its built-in analysis transforms let the engineer quantify these defects before cleansing, without writing custom code.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.