Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer is using AWS Glue to process a large dataset where a small number of partitions contain disproportionately more rows than others, causing some executors to run much longer than others and the job to take hours. The engineer wants to redistribute the data across partitions before a join operation to improve performance. Which technique should the engineer apply?

⚠ Common exam trap

The trap here is believing that adding more compute capacity, such as more DPUs, automatically resolves data skew, when the bottleneck is the uneven distribution of rows rather than total capacity.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Call repartition on the DataFrame using a join key column before the join, which triggers a full shuffle and distributes rows evenly across partitions.

Repartitioning on the join key before the join causes a shuffle that spreads rows evenly across partitions, which is the direct fix for skewed partitions. Increasing DPUs does not help because the largest partition remains a single bottleneck, coalesce can worsen skew by merging without redistribution, and job bookmarks only reduce data across runs. The key is to redistribute data within the run before the expensive join.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the number of AWS Glue DPUs allocated to the job so that more executors are available to process the skewed partitions.

    Why it's wrong here

    Adding DPUs increases the total number of executors, but if the data is skewed, the largest partition still lands on a single executor and remains the bottleneck. The job's overall runtime is bounded by the slowest task, so more executors do not fix an uneven distribution of rows. Additional DPUs may even increase cost without resolving the underlying skew, making this an ineffective remedy for the described problem.

  • ✗

    Use coalesce to reduce the number of partitions, which merges partitions without a full shuffle and balances the data.

    Why it's wrong here

    Coalesce reduces the number of partitions by merging existing ones without a full shuffle, but it does not redistribute rows to balance skewed keys. If a small number of partitions contain most of the data, coalescing can actually concentrate the skew further. Because coalesce avoids a shuffle, it cannot evenly spread rows for a join key, so it does not solve the performance problem described.

  • ✗

    Enable the AWS Glue job bookmarks feature so that only new data is processed on subsequent runs, reducing the data volume.

    Why it's wrong here

    Job bookmarks track previously processed data to avoid reprocessing it on subsequent runs, which reduces input volume over time but does nothing for skew within a single run. The current job run still faces the same uneven partition distribution during the join. Bookmarks are a state-management feature, not a partitioning or shuffling technique, so they do not address the executor imbalance described.

  • ✓

    Call repartition on the DataFrame using a join key column before the join, which triggers a full shuffle and distributes rows evenly across partitions.

    Why this is correct

    Repartitioning on the join key before the join forces a shuffle that redistributes rows so that each partition contains a roughly equal share of data for that key. This directly addresses the skew that causes some executors to process far more rows than others. While a full shuffle has a cost, it is the standard remedy for skewed joins and can dramatically reduce overall job time when skew is severe.

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.