DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer is configuring an AWS Glue ETL job that reads from an Amazon S3 bucket containing nested JSON and writes to a partitioned Parquet table. The engineer wants to reduce job cost and improve read performance. (Choose two.)
⚠ Common exam trap
The trap here is assuming that adding workers or running a crawler reduces cost, when bookmarks and output partitioning are the levers that actually cut DPU time and scan volume.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable job bookmarks to avoid reprocessing previously processed files.
Job bookmarks cut cost by skipping already-processed S3 objects, which is important for incremental pipelines. Partitioning the output Parquet table by low-cardinality filter columns enables partition pruning, reducing I/O and improving read performance for downstream queries. Crawlers, worker scaling, and pre-conversion do not directly deliver both cost reduction and read performance for this job configuration.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Enable job bookmarks to avoid reprocessing previously processed files.
Why this is correct
Job bookmarks track which S3 objects have already been processed and skip them on subsequent runs. For incremental data landing in S3, this avoids re-reading and re-transforming old files, directly reducing DPU hours and cost. It is a supported Glue feature and a common cost-control measure for recurring jobs.
- ✗
Convert the source JSON to Parquet before the Glue job runs.
Why it's wrong here
Converting source JSON to Parquet ahead of the job adds an extra pipeline stage and its own compute cost. While Parquet is more efficient to read, the question is about configuring the Glue job, and pre-converting does not reduce the job's cost directly. It also complicates the architecture by introducing another transformation step.
- ✗
Increase the number of workers to the maximum allowed for the job.
Why it's wrong here
Adding workers increases parallelism but also increases the hourly DPU cost. It can shorten runtime, but the total cost may rise if the job is not compute-bound. The question asks for cost reduction and read performance, and simply scaling out does not guarantee either without addressing data layout or reprocessing.
- ✗
Use the Glue crawler to infer the schema and store it in the Data Catalog.
Why it's wrong here
A crawler populates the Data Catalog with table metadata, which helps query engines and job authoring. It does not reduce the cost of an ETL run or improve read performance of the job itself. The crawler runs separately and charges its own DPU time, so it is not a cost-reduction measure for the ETL job.
- ✓
Partition the output Parquet table by low-cardinality columns used in filters.
Why this is correct
Partitioning the output by columns frequently used in downstream filters enables partition pruning. Readers scan only relevant partitions, cutting I/O and improving query performance. This also reduces the volume of data processed in subsequent jobs, which lowers cost. The key is choosing low-cardinality columns to avoid creating too many small partitions.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.