Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer is using AWS Glue DataBrew to prepare a dataset stored in Amazon S3. The dataset contains missing values, inconsistent date formats, and duplicate rows. The engineer needs to clean the data and produce a transformed output for downstream analytics. (Choose two.)

⚠ Common exam trap

The trap here is selecting operational or diagnostic actions like profiling or scheduling instead of the transformation steps that actually change the data.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Create a recipe step using the Fill missing values transformation to replace nulls with a specified value or strategy.

Cleaning the dataset requires transformation steps that actually modify data. Removing duplicates eliminates redundant rows, and filling missing values addresses nulls. Together these recipe steps resolve two of the stated data quality issues. Date format normalization would need an additional step. Source configuration, profiling, and scheduling support the workflow but do not themselves clean the data.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Schedule a DataBrew recipe job to run on a cron expression so the cleaning steps execute periodically.

    Why it's wrong here

    Scheduling a recipe job automates when the cleaning runs, but scheduling itself does not perform any cleaning. Without the appropriate transformation steps in the recipe, the job would output unchanged data. Scheduling is an operational concern, not a data cleansing action, so it does not meet the stated requirements.

  • ✓

    Create a recipe step using the Fill missing values transformation to replace nulls with a specified value or strategy.

    Why this is correct

    The Fill missing values transformation lets you replace null or missing entries with a constant, a calculated value, or a strategy such as most frequent. It directly resolves the missing values problem in the dataset. Because it is a recipe step, it remains part of the reusable transformation pipeline applied to the S3 data.

  • ✗

    Use the DataBrew profile job to generate statistics about missing values and duplicates in the dataset.

    Why it's wrong here

    A profile job analyzes data and produces statistics such as null counts and duplicate detection, but it does not modify the data. Profiling is diagnostic and helps you understand data quality, yet it leaves the dataset unchanged. The requirement is to clean and produce transformed output, so profiling alone is insufficient.

  • ✗

    Configure a DataBrew dataset to use the S3 bucket as a source and set the file type to CSV with header detection enabled.

    Why it's wrong here

    Configuring the dataset source and file type is necessary to read the data, but it does not clean missing values, normalize dates, or remove duplicates. This is a setup step rather than a transformation. The requirement is to clean the data, so source configuration alone does not satisfy the scenario's data quality goals.

  • ✓

    Apply a recipe step that uses the Remove duplicates transformation to eliminate duplicate rows.

    Why this is correct

    The Remove duplicates transformation in DataBrew identifies and removes rows that are identical across selected columns. Applying it as a recipe step directly addresses the duplicate rows requirement and can be saved and reused. This is a native DataBrew operation designed for exactly this data cleaning task within the dataset preparation workflow.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.