Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer is using AWS Glue Studio to create a job that joins data from two Amazon S3 sources: a large fact table and a small dimension table. The job performs a join and then writes the result to Amazon S3 in Parquet format. The engineer notices that the job is running slowly and consuming many DPUs. Which optimization technique should the engineer apply to improve performance?

⚠ Common exam trap

The trap here is assuming that adding more DPUs is always the best way to improve performance, without considering algorithmic optimizations like broadcast joins.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a broadcast join to replicate the small dimension table to all worker nodes.

A broadcast join is the appropriate optimization when joining a large fact table with a small dimension table. It avoids shuffling the large table by replicating the small table to all nodes, reducing network traffic and improving speed. The other options either do not address the join inefficiency, may increase cost, or degrade performance by using a less efficient format.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use a broadcast join to replicate the small dimension table to all worker nodes.

    Why this is correct

    A broadcast join is an optimization technique where the smaller table is replicated to all worker nodes, allowing the join to be performed locally without shuffling the large fact table. This reduces network overhead and improves performance for joins where one table is significantly smaller. AWS Glue supports broadcast joins when one side of the join is small enough to fit in memory.

  • ✗

    Partition the fact table by join key before the join.

    Why it's wrong here

    Partitioning the fact table by join key can help with data organization, but it does not directly optimize the join operation in AWS Glue. The join still requires shuffling data across nodes unless a broadcast join is used. Partitioning is more beneficial for query performance in downstream analytics, not for the ETL join itself. It may add an extra step and not improve the join speed.

  • ✗

    Convert the dimension table to CSV format to reduce file size.

    Why it's wrong here

    Converting the dimension table to CSV may reduce file size compared to Parquet, but CSV is less efficient for reading and processing in Glue. Parquet is columnar and compressed, which generally improves performance. Changing to CSV would likely degrade performance due to increased I/O and parsing overhead. Format choice should favor Parquet for analytics workloads.

  • ✗

    Increase the number of DPUs to the maximum to speed up the join.

    Why it's wrong here

    Increasing DPUs can provide more resources, but it also increases cost. If the join is inefficient due to data shuffling, adding more DPUs may not solve the underlying issue and could lead to unnecessary expenses. The maximum DPU limit is 100, but without optimizing the join strategy, the job may still run slowly. It is better to address the join algorithm first.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.