DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer is using AWS Glue Studio to build a visual ETL job that joins a large Amazon S3 dataset with a small reference dataset of country codes. The join is currently implemented as a standard join, and the job runs slowly and shuffles large amounts of data. The engineer wants to optimize performance without changing the output. Which change should the engineer make?
⚠ Common exam trap
The trap here is thinking a longer timeout fixes a slow join, when the real cost is shuffle volume that only a broadcast join can remove.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Replace the standard join with a broadcast join so the small reference dataset is replicated to each executor.
A broadcast join replicates the small reference dataset to each executor, eliminating the shuffle of the large dataset that a standard join incurs. Because the output rows are unchanged, this is a pure performance optimization. Timeout increases, file consolidation, and catalog views do not alter the join execution strategy and therefore do not address the shuffle bottleneck.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the job's timeout setting so the existing standard join has more time to complete.
Why it's wrong here
Extending the timeout does not reduce shuffle volume or improve throughput; it only allows the slow job to run longer before being terminated. The underlying inefficiency remains, and costs may rise. This change treats the symptom rather than the cause of the performance problem.
- ✗
Move the small reference dataset into the Glue Data Catalog as a view and reference it in the join.
Why it's wrong here
Catalog views are metadata constructs for querying; they do not change the physical join strategy or reduce shuffle. The join would still execute as a standard shuffle join. Cataloging the reference data does not by itself deliver the broadcast optimization needed to speed up the operation.
- ✓
Replace the standard join with a broadcast join so the small reference dataset is replicated to each executor.
Why this is correct
When one side of a join is small, broadcasting it to every executor avoids shuffling the large dataset across the cluster. Glue Studio exposes a join type option that maps to Spark's broadcast join behavior. This reduces network and shuffle cost while producing the same joined output, directly addressing the slowness described.
- ✗
Convert the large S3 dataset to a single file before the join to eliminate partitioning overhead.
Why it's wrong here
Consolidating a large dataset into one file destroys parallelism and forces a single executor to process all data, which typically makes the job slower, not faster. It also increases memory pressure. Shuffle cost is better addressed by broadcasting the small side rather than collapsing the large side.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.