DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer is using AWS Glue Studio to create a job that joins data from two Amazon S3 sources: a large fact table and a small dimension table. The job performs a join and then writes the result to Amazon S3 in Parquet format. The engineer notices that the job is running slowly and consuming many DPUs. Which optimization technique should the engineer apply to improve performance?
⚠ Common exam trap
The trap here is assuming that adding more DPUs is always the best way to improve performance, without considering algorithmic optimizations like broadcast joins.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a broadcast join to replicate the small dimension table to all worker nodes.
A broadcast join is the appropriate optimization when joining a large fact table with a small dimension table. It avoids shuffling the large table by replicating the small table to all nodes, reducing network traffic and improving speed. The other options either do not address the join inefficiency, may increase cost, or degrade performance by using a less efficient format.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use a broadcast join to replicate the small dimension table to all worker nodes.
Why this is correct
A broadcast join is an optimization technique where the smaller table is replicated to all worker nodes, allowing the join to be performed locally without shuffling the large fact table. This reduces network overhead and improves performance for joins where one table is significantly smaller. AWS Glue supports broadcast joins when one side of the join is small enough to fit in memory.
- ✗
Partition the fact table by join key before the join.
Why it's wrong here
Partitioning the fact table by join key can help with data organization, but it does not directly optimize the join operation in AWS Glue. The join still requires shuffling data across nodes unless a broadcast join is used. Partitioning is more beneficial for query performance in downstream analytics, not for the ETL join itself. It may add an extra step and not improve the join speed.
- ✗
Convert the dimension table to CSV format to reduce file size.
Why it's wrong here
Converting the dimension table to CSV may reduce file size compared to Parquet, but CSV is less efficient for reading and processing in Glue. Parquet is columnar and compressed, which generally improves performance. Changing to CSV would likely degrade performance due to increased I/O and parsing overhead. Format choice should favor Parquet for analytics workloads.
- ✗
Increase the number of DPUs to the maximum to speed up the join.
Why it's wrong here
Increasing DPUs can provide more resources, but it also increases cost. If the join is inefficient due to data shuffling, adding more DPUs may not solve the underlying issue and could lead to unnecessary expenses. The maximum DPU limit is 100, but without optimizing the join strategy, the job may still run slowly. It is better to address the join algorithm first.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.