Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer is using AWS Glue Studio to build a job that reads from an Amazon S3 data source, applies a filter transformation, and writes to Amazon S3 in Parquet. The engineer notices that the job is reading all files in the prefix, including files that do not match the expected schema, causing job failures. Which action should the engineer take to ensure only valid files are processed?

⚠ Common exam trap

The trap here is reaching for job bookmarks or crawlers to solve a schema mismatch; those features handle incremental processing and cataloging, not selective file exclusion at read time.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Configure the S3 data source with a path that includes only valid partitions or use a glob pattern to exclude invalid files

The root cause is that the job reads all objects under the prefix, including files that do not match the schema. Restricting the S3 source path with a glob pattern or limiting to valid partitions ensures only conforming files are read. This is a configuration change that directly prevents the failures without adding compute or custom classification logic.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Enable job bookmarks and set the transformation context to filter invalid records

    Why it's wrong here

    Job bookmarks track previously processed data to avoid reprocessing, but they do not filter files by schema validity. Setting a transformation context does not automatically exclude files that fail schema checks. This option addresses incremental processing, not the problem of invalid files causing job failures.

  • ✗

    Increase the number of DPUs and enable auto-scaling to handle the invalid files

    Why it's wrong here

    Adding DPUs or enabling auto-scaling increases compute capacity but does not address schema mismatches. Invalid files would still cause parsing errors or job failures because the data does not conform to the expected schema. This option treats a resource problem rather than the actual data quality issue.

  • ✗

    Use a Glue crawler with a custom classifier to catalog only valid files and update the job to read from the catalog table

    Why it's wrong here

    A crawler with a custom classifier can identify formats, but it does not prevent the Glue job from reading files that do not match the schema. The job would still attempt to read all objects under the prefix unless the source is narrowed. This adds complexity without solving the root cause of invalid files causing failures.

  • ✓

    Configure the S3 data source with a path that includes only valid partitions or use a glob pattern to exclude invalid files

    Why this is correct

    Narrowing the S3 source path with a glob pattern or restricting to specific partitions ensures the Glue job only reads files that match the expected schema. This directly prevents invalid files from being processed and avoids job failures. It is the simplest and most targeted fix for the described problem.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.