DEA-C01 Data Store Management Practice Question
A data engineer is using AWS Glue to catalog data stored in Amazon S3. The engineer needs to run an AWS Glue ETL job that reads from a large dataset in Parquet format, performs transformations, and writes the output to Amazon Redshift. The job must handle data skew and optimize performance. Which AWS Glue feature should the engineer use to address data skew during the join operation?
⚠ Common exam trap
The trap here is thinking that adding more resources (DPUs) will automatically resolve data skew, when in fact skew requires specific algorithmic handling.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Configure the job to use the `spark.sql.adaptive.enabled` and `spark.sql.adaptive.skewJoin.enabled` properties.
Data skew in joins can be mitigated by enabling Spark's adaptive query execution (AQE) with skew join optimization. This feature dynamically detects and splits skewed partitions, balancing the workload across executors. Other options do not directly address skew: job bookmarks track processed data, Relationalize flattens nested data, and adding DPUs increases capacity without redistributing data.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Configure the job to use the `spark.sql.adaptive.enabled` and `spark.sql.adaptive.skewJoin.enabled` properties.
Why this is correct
Enabling adaptive query execution (AQE) with skew join optimization allows Spark to dynamically handle data skew during joins. It detects skewed partitions and splits them into smaller sub-partitions, balancing the workload. This directly addresses the skew issue and improves job performance. These properties are supported in AWS Glue jobs running Spark 3.0 and later.
- ✗
Use the AWS Glue DynamicFrame and apply the `Relationalize` transform to flatten nested data.
Why it's wrong here
The Relationalize transform is used to flatten nested structures like JSON into relational tables. It does not address data skew in joins. While flattening can be useful for certain transformations, it does not redistribute data to balance partitions. Therefore, it would not mitigate the skew problem during the join operation.
- ✗
Enable job bookmarks to track processed data and avoid reprocessing.
Why it's wrong here
Job bookmarks help track data that has already been processed in previous job runs, preventing reprocessing of old data. They do not address data skew during joins. Data skew causes some partitions to be much larger than others, leading to uneven workload distribution. Job bookmarks are unrelated to skew mitigation, so they would not solve the performance issue.
- ✗
Increase the number of DPUs allocated to the job to provide more resources for handling skewed data.
Why it's wrong here
Adding more DPUs increases the overall compute capacity but does not redistribute skewed data across partitions. A single large partition can still overwhelm one executor, causing performance bottlenecks. While more DPUs can help with overall throughput, they do not solve the root cause of data skew. Skew-specific optimizations like AQE are required.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.