Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The engineer needs to ensure that the job handles data quality issues such as duplicate records and missing values before loading. The job must also minimize the amount of data shuffled across the network. Which two actions should the engineer take? (Choose two.)

⚠ Common exam trap

Test-takers frequently confuse data quality transforms with performance tuning; repartitioning and job bookmarks are often mistakenly associated with data cleaning but do not resolve duplicates or missing values.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use the FillMissingValues transform to replace missing values with a default.

DropDuplicates removes duplicate records, and FillMissingValues replaces missing values with defaults, directly addressing the data quality requirements. Both are built-in AWS Glue transforms that can be applied without custom code. The other options do not handle duplicates or missing values, and repartitioning to a single partition would increase network shuffle, contrary to the requirement.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use the FillMissingValues transform to replace missing values with a default.

    Why this is correct

    FillMissingValues is an AWS Glue transform that fills null or missing values in specified columns with a default value. This handles missing values as required. It is a built-in transform that can be easily added to the job, ensuring data quality before loading to Redshift, and it does not require custom code.

  • ✗

    Repartition the DynamicFrame to a single partition before writing to Redshift.

    Why it's wrong here

    Repartitioning to a single partition would force all data through one executor, severely limiting parallelism and increasing the amount of data shuffled across the network. This contradicts the goal of minimizing network shuffle. While it might help with ordering, it is not appropriate for large datasets and does not address data quality issues.

  • ✓

    Use the DropDuplicates transform to remove duplicate records.

    Why this is correct

    The DropDuplicates transform in AWS Glue removes duplicate rows based on specified columns or all columns. This directly addresses duplicate records and is a built-in transform that can be applied before loading to Redshift. It helps ensure data quality by eliminating duplicates, and it can be used without custom code, aligning with the requirement to handle duplicates.

  • ✗

    Use the ResolveChoice transform to handle data type conflicts.

    Why it's wrong here

    ResolveChoice handles data type conflicts by casting or retaining types, but it does not address duplicate records or missing values. While it can be useful for schema consistency, it is not relevant to the specific data quality issues mentioned. Including it would not help with duplicates or missing values, so it is not a correct action for this scenario.

  • ✗

    Enable job bookmarks to track processed data and avoid reprocessing.

    Why it's wrong here

    Job bookmarks track previously processed data to avoid reprocessing, which is useful for incremental loads. However, they do not handle duplicate records within a single run or missing values. They also do not directly affect network shuffle. While valuable for incremental processing, they do not address the data quality issues specified in the scenario.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.