DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer is using AWS Glue Studio to build a job that reads from an Amazon S3 data source, applies a filter transformation, and writes to Amazon S3 in Parquet. The engineer notices that the job is reading all files in the prefix, including files that do not match the expected schema, causing job failures. Which action should the engineer take to ensure only valid files are processed?
⚠ Common exam trap
The trap here is reaching for job bookmarks or crawlers to solve a schema mismatch; those features handle incremental processing and cataloging, not selective file exclusion at read time.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Configure the S3 data source with a path that includes only valid partitions or use a glob pattern to exclude invalid files
The root cause is that the job reads all objects under the prefix, including files that do not match the schema. Restricting the S3 source path with a glob pattern or limiting to valid partitions ensures only conforming files are read. This is a configuration change that directly prevents the failures without adding compute or custom classification logic.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable job bookmarks and set the transformation context to filter invalid records
Why it's wrong here
Job bookmarks track previously processed data to avoid reprocessing, but they do not filter files by schema validity. Setting a transformation context does not automatically exclude files that fail schema checks. This option addresses incremental processing, not the problem of invalid files causing job failures.
- ✗
Increase the number of DPUs and enable auto-scaling to handle the invalid files
Why it's wrong here
Adding DPUs or enabling auto-scaling increases compute capacity but does not address schema mismatches. Invalid files would still cause parsing errors or job failures because the data does not conform to the expected schema. This option treats a resource problem rather than the actual data quality issue.
- ✗
Use a Glue crawler with a custom classifier to catalog only valid files and update the job to read from the catalog table
Why it's wrong here
A crawler with a custom classifier can identify formats, but it does not prevent the Glue job from reading files that do not match the schema. The job would still attempt to read all objects under the prefix unless the source is narrowed. This adds complexity without solving the root cause of invalid files causing failures.
- ✓
Configure the S3 data source with a path that includes only valid partitions or use a glob pattern to exclude invalid files
Why this is correct
Narrowing the S3 source path with a glob pattern or restricting to specific partitions ensures the Glue job only reads files that match the expected schema. This directly prevents invalid files from being processed and avoids job failures. It is the simplest and most targeted fix for the described problem.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.