Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer runs an AWS Glue ETL job that joins a 4 TB Parquet dataset in Amazon S3 with a small 40 MB reference lookup table stored as a single CSV file in S3. The join is taking hours and the job frequently fails with executor out-of-memory errors. The engineer wants to reduce shuffle and memory pressure with the least development effort. Which approach should the engineer take?

⚠ Common exam trap

The trap here is assuming that adding more DPU workers will fix a shuffle-heavy join, when the real fix is to change the join strategy so the large dataset is never shuffled.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Convert the small lookup CSV into a broadcast join by loading it with a broadcast hint, avoiding a shuffle of the large dataset.

Because one side of the join is only 40 MB, broadcasting that lookup table to every executor removes the need to shuffle the multi-terabyte dataset, cutting network I/O and eliminating the memory spikes that cause executor OOM failures. This is a minimal code change, satisfying the least-effort constraint, whereas scaling workers or repartitioning the large side still incurs a full shuffle.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the number of Glue DPU workers so more executors can hold the shuffled join partitions.

    Why it's wrong here

    Adding workers increases total cluster capacity but also the number of partitions that must exchange data, and the shuffle volume for a 4 TB join remains enormous. Memory pressure per executor is not fundamentally resolved by scaling out, and cost rises sharply. The underlying design flaw of shuffling a huge dataset for a tiny lookup remains unaddressed.

  • ✓

    Convert the small lookup CSV into a broadcast join by loading it with a broadcast hint, avoiding a shuffle of the large dataset.

    Why this is correct

    Broadcasting the 40 MB lookup table ships it to every executor so the large 4 TB dataset never has to be shuffled for the join. This eliminates the expensive shuffle and the memory pressure caused by co-locating both sides, directly addressing the OOM failures. It requires only a small change to the join logic, satisfying the least-effort requirement.

  • ✗

    Repartition the large Parquet dataset by the join key using a Glue transform before the join.

    Why it's wrong here

    Repartitioning on the join key can reduce per-partition skew, but it still forces a full shuffle of the 4 TB dataset, which is exactly the operation causing the runtime and memory problems. It also adds an extra stage and development effort. Broadcasting the tiny lookup table avoids the shuffle altogether with far less work.

  • ✗

    Enable the AWS Glue job bookmark and rerun the job so only new files are processed.

    Why it's wrong here

    Job bookmarks track previously processed data so subsequent runs skip already-seen S3 objects, which reduces total input volume over time. However, this job must join the full 4 TB dataset each run, and bookmarks do nothing to reduce shuffle or the memory footprint of the join itself, so the executor OOM and long runtime would persist.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.