Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer is using AWS Glue Studio to build a visual ETL job that joins a large Amazon S3 dataset with a small reference dataset of country codes. The join is currently implemented as a standard join, and the job runs slowly and shuffles large amounts of data. The engineer wants to optimize performance without changing the output. Which change should the engineer make?

⚠ Common exam trap

The trap here is thinking a longer timeout fixes a slow join, when the real cost is shuffle volume that only a broadcast join can remove.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Replace the standard join with a broadcast join so the small reference dataset is replicated to each executor.

A broadcast join replicates the small reference dataset to each executor, eliminating the shuffle of the large dataset that a standard join incurs. Because the output rows are unchanged, this is a pure performance optimization. Timeout increases, file consolidation, and catalog views do not alter the join execution strategy and therefore do not address the shuffle bottleneck.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the job's timeout setting so the existing standard join has more time to complete.

    Why it's wrong here

    Extending the timeout does not reduce shuffle volume or improve throughput; it only allows the slow job to run longer before being terminated. The underlying inefficiency remains, and costs may rise. This change treats the symptom rather than the cause of the performance problem.

  • ✗

    Move the small reference dataset into the Glue Data Catalog as a view and reference it in the join.

    Why it's wrong here

    Catalog views are metadata constructs for querying; they do not change the physical join strategy or reduce shuffle. The join would still execute as a standard shuffle join. Cataloging the reference data does not by itself deliver the broadcast optimization needed to speed up the operation.

  • ✓

    Replace the standard join with a broadcast join so the small reference dataset is replicated to each executor.

    Why this is correct

    When one side of a join is small, broadcasting it to every executor avoids shuffling the large dataset across the cluster. Glue Studio exposes a join type option that maps to Spark's broadcast join behavior. This reduces network and shuffle cost while producing the same joined output, directly addressing the slowness described.

  • ✗

    Convert the large S3 dataset to a single file before the join to eliminate partitioning overhead.

    Why it's wrong here

    Consolidating a large dataset into one file destroys parallelism and forces a single executor to process all data, which typically makes the job slower, not faster. It also increases memory pressure. Shuffle cost is better addressed by broadcasting the small side rather than collapsing the large side.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.