Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer needs to run an AWS Glue for Apache Spark ETL job that joins a 40 GB Amazon S3 Parquet dataset with a small 8 MB reference lookup table stored as CSV in Amazon S3. The reference table is read on every join and the job's executors are spending a large amount of shuffle time on the join. The reference table changes only once per month. Which approach MOST efficiently reduces shuffle overhead in the Glue job?

⚠ Common exam trap

The trap here is assuming that adding compute resources or changing file formats eliminates shuffle cost, when only broadcasting the small side removes the redistribution of the large dataset.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Broadcast the reference table by reading it with the Spark DataFrame API and applying broadcast() before the join.

The small, slowly changing reference table is an ideal broadcast candidate. Replicating it to every executor lets Spark perform a broadcast hash join where the large Parquet side stays in place and is never redistributed. Increasing DPUs, converting formats, or pre-partitioning all leave the shuffle of the large dataset intact, so they fail to address the actual bottleneck.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Partition both datasets by the join key and write them to S3 before the join so Spark can perform a partition-wise merge join.

    Why it's wrong here

    Pre-partitioning both datasets by the join key would require an additional full pass over the 40 GB dataset and rewriting it, which is more expensive than the join itself. Spark does not automatically use pre-partitioned S3 layout as a shuffle-free merge join without explicit bucketBy and matching partition counts. This adds cost and complexity without reliably removing the shuffle.

  • ✗

    Increase the number of Glue DPUs so that more executors are available to parallelize the shuffle stage of the join.

    Why it's wrong here

    Adding DPUs increases parallelism but does not remove the underlying shuffle of the 40 GB dataset. The join still redistributes both sides across the network, so the shuffle cost scales with data volume and may even grow. More executors also raise cost. This addresses symptoms rather than the root cause, which is the unnecessary movement of the large dataset during the join.

  • ✓

    Broadcast the reference table by reading it with the Spark DataFrame API and applying broadcast() before the join.

    Why this is correct

    Broadcasting the small reference DataFrame replicates it to each executor, so the large Parquet dataset never needs to be shuffled. This eliminates the expensive shuffle of the 40 GB dataset that the join currently causes. Because the table is only 8 MB and changes monthly, broadcasting is safe and fits comfortably in executor memory, making it the most efficient fix for shuffle overhead in the Glue job.

  • ✗

    Convert the reference CSV to Parquet and enable Glue job bookmarks so the lookup is only read once per month.

    Why it's wrong here

    Parquet reduces read I/O for the tiny 8 MB reference file, but the join still shuffles the 40 GB dataset, so the dominant shuffle cost remains. Job bookmarks track processed input for incremental processing and do not affect join execution or shuffle behavior. This improves storage format and state tracking but does not solve the shuffle overhead described in the scenario.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.