DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The engineer needs to ensure that the job handles data quality issues such as duplicate records and missing values before loading. The job must also minimize the amount of data shuffled across the network. Which two actions should the engineer take? (Choose two.)
⚠ Common exam trap
Test-takers frequently confuse data quality transforms with performance tuning; repartitioning and job bookmarks are often mistakenly associated with data cleaning but do not resolve duplicates or missing values.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use the FillMissingValues transform to replace missing values with a default.
DropDuplicates removes duplicate records, and FillMissingValues replaces missing values with defaults, directly addressing the data quality requirements. Both are built-in AWS Glue transforms that can be applied without custom code. The other options do not handle duplicates or missing values, and repartitioning to a single partition would increase network shuffle, contrary to the requirement.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use the FillMissingValues transform to replace missing values with a default.
Why this is correct
FillMissingValues is an AWS Glue transform that fills null or missing values in specified columns with a default value. This handles missing values as required. It is a built-in transform that can be easily added to the job, ensuring data quality before loading to Redshift, and it does not require custom code.
- ✗
Repartition the DynamicFrame to a single partition before writing to Redshift.
Why it's wrong here
Repartitioning to a single partition would force all data through one executor, severely limiting parallelism and increasing the amount of data shuffled across the network. This contradicts the goal of minimizing network shuffle. While it might help with ordering, it is not appropriate for large datasets and does not address data quality issues.
- ✓
Use the DropDuplicates transform to remove duplicate records.
Why this is correct
The DropDuplicates transform in AWS Glue removes duplicate rows based on specified columns or all columns. This directly addresses duplicate records and is a built-in transform that can be applied before loading to Redshift. It helps ensure data quality by eliminating duplicates, and it can be used without custom code, aligning with the requirement to handle duplicates.
- ✗
Use the ResolveChoice transform to handle data type conflicts.
Why it's wrong here
ResolveChoice handles data type conflicts by casting or retaining types, but it does not address duplicate records or missing values. While it can be useful for schema consistency, it is not relevant to the specific data quality issues mentioned. Including it would not help with duplicates or missing values, so it is not a correct action for this scenario.
- ✗
Enable job bookmarks to track processed data and avoid reprocessing.
Why it's wrong here
Job bookmarks track previously processed data to avoid reprocessing, which is useful for incremental loads. However, they do not handle duplicate records within a single run or missing values. They also do not directly affect network shuffle. While valuable for incremental processing, they do not address the data quality issues specified in the scenario.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.