DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer runs an AWS Glue ETL job that joins a 4 TB Parquet dataset in Amazon S3 with a small 40 MB reference lookup table stored as a single CSV file in S3. The join is taking hours and the job frequently fails with executor out-of-memory errors. The engineer wants to reduce shuffle and memory pressure with the least development effort. Which approach should the engineer take?
⚠ Common exam trap
The trap here is assuming that adding more DPU workers will fix a shuffle-heavy join, when the real fix is to change the join strategy so the large dataset is never shuffled.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Convert the small lookup CSV into a broadcast join by loading it with a broadcast hint, avoiding a shuffle of the large dataset.
Because one side of the join is only 40 MB, broadcasting that lookup table to every executor removes the need to shuffle the multi-terabyte dataset, cutting network I/O and eliminating the memory spikes that cause executor OOM failures. This is a minimal code change, satisfying the least-effort constraint, whereas scaling workers or repartitioning the large side still incurs a full shuffle.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the number of Glue DPU workers so more executors can hold the shuffled join partitions.
Why it's wrong here
Adding workers increases total cluster capacity but also the number of partitions that must exchange data, and the shuffle volume for a 4 TB join remains enormous. Memory pressure per executor is not fundamentally resolved by scaling out, and cost rises sharply. The underlying design flaw of shuffling a huge dataset for a tiny lookup remains unaddressed.
- ✓
Convert the small lookup CSV into a broadcast join by loading it with a broadcast hint, avoiding a shuffle of the large dataset.
Why this is correct
Broadcasting the 40 MB lookup table ships it to every executor so the large 4 TB dataset never has to be shuffled for the join. This eliminates the expensive shuffle and the memory pressure caused by co-locating both sides, directly addressing the OOM failures. It requires only a small change to the join logic, satisfying the least-effort requirement.
- ✗
Repartition the large Parquet dataset by the join key using a Glue transform before the join.
Why it's wrong here
Repartitioning on the join key can reduce per-partition skew, but it still forces a full shuffle of the 4 TB dataset, which is exactly the operation causing the runtime and memory problems. It also adds an extra stage and development effort. Broadcasting the tiny lookup table avoids the shuffle altogether with far less work.
- ✗
Enable the AWS Glue job bookmark and rerun the job so only new files are processed.
Why it's wrong here
Job bookmarks track previously processed data so subsequent runs skip already-seen S3 objects, which reduces total input volume over time. However, this job must join the full 4 TB dataset each run, and bookmarks do nothing to reduce shuffle or the memory footprint of the join itself, so the executor OOM and long runtime would persist.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.