Courseiva
Data Operations and Support →mediumMultiple Choice

DEA-C01 Data Operations and Support Practice Question

A data engineer is monitoring an AWS Glue job that reads from Amazon S3 and writes to Amazon Redshift. The job runs daily and recently started taking significantly longer to complete. The engineer checks the job metrics and notices that the number of DPUs used is consistently at the maximum allocated, and the job's Spark UI shows many tasks spilling to disk. Which action should the engineer take to improve performance?

⚠ Common exam trap

The trap here is assuming that adding more DPUs will always solve performance issues, when the real problem is often data distribution and memory management within Spark.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Repartition the data or adjust the number of partitions to reduce skew and improve memory usage.

Disk spilling in Spark indicates that executors are running out of memory during processing, often due to data skew or insufficient partitions. Repartitioning the data or adjusting partition counts distributes the workload more evenly across executors, reducing memory pressure and eliminating spilling. This is a targeted fix for the observed symptom and improves job performance without unnecessary resource increases.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Change the output format from Parquet to CSV to reduce write overhead.

    Why it's wrong here

    Changing to CSV would likely increase write overhead and file size, and it does not address the memory spilling during processing. Parquet is columnar and more efficient for analytics. The spilling occurs during the Spark job's execution, not during the final write. Format change is not a solution for memory issues.

  • ✗

    Increase the number of DPUs allocated to the job.

    Why it's wrong here

    Increasing DPUs adds more compute resources, but the issue is disk spilling due to insufficient memory per executor. Simply adding DPUs without adjusting executor memory or partitioning may not resolve the spilling. It could even worsen it if the data is skewed. The engineer should first address partitioning and memory configuration.

  • ✓

    Repartition the data or adjust the number of partitions to reduce skew and improve memory usage.

    Why this is correct

    Disk spilling occurs when Spark executors run out of memory and write intermediate data to disk. This is often caused by data skew or too few partitions, leading to large tasks. Repartitioning the data or increasing the number of partitions distributes the workload more evenly, reducing memory pressure and spilling. This directly addresses the root cause shown in the Spark UI.

  • ✗

    Enable job bookmarks to avoid reprocessing old data.

    Why it's wrong here

    Job bookmarks help with incremental processing by tracking previously processed data, but they do not affect the performance of processing new data. The job is already running daily and likely processing only new data. The spilling issue is related to how data is partitioned and processed, not to reprocessing old files.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.