Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer uses AWS Glue DataBrew to clean a dataset stored in Amazon S3. The recipe includes a 'Filter' step that removes rows where the 'status' column equals 'INVALID'. After running the recipe, the engineer notices that the output still contains rows with status 'INVALID'. The recipe was published and the job ran successfully. What is the most likely cause?

⚠ Common exam trap

The trap here is assuming that DataBrew filter conditions are case-insensitive or that a successful job run guarantees the filter matched the intended rows.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The filter condition used a case-sensitive match and the actual values are 'invalid' in lowercase, so the condition did not match any rows.

DataBrew filter conditions are case-sensitive, so a condition matching 'INVALID' will not remove rows containing 'invalid'. The engineer should either standardize the case of the status column or adjust the filter condition to match the actual values. This explains why the job succeeded but the unwanted rows remained.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The filter step was configured as a 'Remove' transformation instead of a 'Filter' transformation, so it only flagged rows without deleting them.

    Why it's wrong here

    In DataBrew, a 'Remove' transformation deletes rows or columns based on a condition, while 'Filter' keeps rows that match. If the engineer intended to remove invalid rows, using Remove would work. The symptom described (invalid rows remain) is not caused by choosing Remove; Remove would actually eliminate them. This option misstates the behavior of the Remove transform.

  • ✗

    The recipe was applied to a sample of the data during job execution, so only a subset was filtered.

    Why it's wrong here

    DataBrew jobs can be configured to run on a sample or full dataset, but when a job runs on the full dataset it applies transformations to all rows. The scenario does not indicate sampling was used, and even with sampling, the filter would still remove matching rows in the sample. Sampling would not cause invalid rows to remain in the full output if the filter matched them.

  • ✗

    The DataBrew job was run in profile mode instead of recipe mode, so transformations were not applied.

    Why it's wrong here

    Profile mode in DataBrew is used to generate statistics and data profiles; it does not apply recipe transformations at all. However, if the job had been run in profile mode, the output would typically not be written as a transformed dataset, and the engineer would likely notice missing output. The scenario states the job ran successfully and produced output with invalid rows, which points to a condition mismatch rather than job mode.

  • ✓

    The filter condition used a case-sensitive match and the actual values are 'invalid' in lowercase, so the condition did not match any rows.

    Why this is correct

    DataBrew filter conditions are case-sensitive by default. If the source data contains 'invalid' rather than 'INVALID', a condition checking for 'INVALID' will not match, so those rows are retained. The engineer should either normalize case or adjust the condition. This is a common cause of filters appearing to have no effect when the job succeeds but output is unchanged.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.