Courseiva

Boosting AWS Glue ETL Job Performance When Loading to Redshift

A company is using AWS Glue ETL to transform and load data from Amazon S3 to Amazon Redshift. The data engineer notices that the job is taking longer than expected. Which TWO actions can improve the job performance?

Quick Answer

The correct answer is to partition the source data in S3 and increase the number of DPUs for the AWS Glue job. Partitioning the source data in S3 reduces the amount of data scanned by Glue, allowing it to process only relevant subsets in parallel, while increasing DPUs allocates more distributed processing units to the ETL job, directly improving parallelism and throughput when loading to Redshift. On the AWS Certified Data Engineer Associate DEA-C01 exam, this question tests your understanding of how Glue’s serverless architecture scales—common traps include confusing S3 Transfer Acceleration (which only speeds uploads) or Redshift Spectrum (which is for querying, not Glue ETL). Remember that Glue job performance is about Glue-side resources and data organization, not Redshift instance size. A useful memory tip: “Partition and DPU—two levers for Glue throughput.”

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Partition the source data in S3.

Options B and C are correct because partitioning the source data in S3 reduces the amount of data scanned by Glue, improving I/O efficiency, and increasing the number of DPUs adds more parallelism for transformations. Option A is incorrect because Redshift Spectrum is for querying data in S3 directly from Redshift, not for Glue ETL jobs. Option D is incorrect because S3 Transfer Acceleration speeds up uploads to S3 but does not affect Glue job performance during ETL processing. Option E is incorrect because larger Redshift node types do not impact Glue job execution; they only affect Redshift query performance.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use Amazon Redshift Spectrum to query data directly.

    Why it's wrong here

    Redshift Spectrum queries external data held in S3 directly, bypassing the load step, so it does not accelerate a Glue job that transforms and writes into Redshift tables. It is tempting because it genuinely avoids loading data, and would be correct if the requirement were querying S3 data without ingesting it into Redshift.

  • ✓

    Partition the source data in S3.

    Why this is correct

    Partitioning the S3 source data lets AWS Glue read only the relevant partitions rather than scanning the entire dataset, cutting I/O and shuffle volume during the transform stage. This directly addresses the stem's slow job by reducing the bytes read before loading into Amazon Redshift.

  • ✓

    Increase the number of DPUs for the Glue job.

    Why this is correct

    Adding DPUs increases the number of Apache Spark executors and parallel tasks available to the Glue job, raising throughput for shuffle-heavy transforms. It directly addresses the stem's constraint of a job running longer than expected.

  • ✗

    Enable S3 Transfer Acceleration.

    Why it's wrong here

    S3 Transfer Acceleration speeds uploads to S3 over long geographic distances, but Glue reads from S3 within the same Region, so it adds no throughput to the ETL job. It is tempting because it genuinely accelerates S3 data movement, and would be correct if the bottleneck were cross-Region uploads into the source bucket rather than Glue processing.

  • ✗

    Use a larger Redshift node type.

    Why it's wrong here

    A larger Redshift node type scales the destination cluster's compute and storage, but the Glue job's runtime is governed by its own DPU allocation and shuffle configuration, so it cannot shorten the ETL phase. It is tempting because Redshift node sizing genuinely improves query performance, and would be correct if slow Redshift queries, not slow Glue jobs, were the problem.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

1 more way this is tested on DEA-C01

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. A company is using AWS Glue ETL to process data from Amazon RDS for MySQL to Amazon S3. The job runs daily and takes 2 hours to complete. The engineer wants to improve performance without increasing cost significantly. Which TWO actions should the engineer take? (Choose TWO.)

medium
  • A.Switch to a smaller worker type (e.g., G.1X instead of G.2X).
  • B.Use Spark DataFrames instead of DynamicFrames.
  • C.Enable 'Auto Scaling' in the Glue job configuration.
  • ✓ D.Add a partition column to the source table based on a date column.
  • ✓ E.Increase the number of Glue DPUs.

Why D: Option D is correct because adding a partition column based on a date column allows AWS Glue to read only the relevant partitions (partition pruning) instead of scanning the entire source table, which reduces I/O and shuffle volume and speeds up the daily job without materially raising cost. Option E is correct because increasing the number of Glue DPUs adds parallel executors, so the Spark job can process partitions concurrently and finish the 2-hour run faster; the cost scales with DPU-hours, but since the job finishes sooner, the total cost increase is modest and often offset by the shorter runtime. Option A is not appropriate because switching to a smaller worker type (G.1X instead of G.2X) reduces memory and compute per worker, which typically slows the job and can cause spills or OOM errors rather than improving performance. Option B is not the best choice because DynamicFrames are already optimized for Glue ETL and converting to Spark DataFrames does not by itself improve performance for this RDS-to-S3 pipeline. Option C is not correct because Glue Auto Scaling adjusts workers based on workload but does not guarantee the performance improvement requested here and can still incur additional DPU cost, making it less directly effective than partitioning and adding DPUs.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.