Courseiva
Data Operations and Support →mediumMultiple Select

DEA-C01 Data Operations and Support Practice Question

A data engineer is troubleshooting an AWS Glue job that fails with 'java.lang.OutOfMemoryError: Java heap space'. The job processes a large dataset. Which TWO configuration changes should the engineer consider to resolve this issue? (Choose TWO.)

⚠ Common exam trap

The trap is that candidates might think reducing partitions or changing output format would help, but these can worsen memory issues; the key is to increase resources and parallelism.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Increase the Spark shuffle partitions configuration (spark.sql.shuffle.partitions).

Option B is correct because increasing spark.sql.shuffle.partitions creates more, smaller partitions during shuffles, which reduces the amount of data each executor task must hold in memory and helps prevent Java heap space exhaustion on large datasets. Option D is correct because allocating more DPUs adds more executors and memory to the Glue job, giving Spark more heap capacity to process the large dataset without running out of memory. Option A is incorrect because switching from Parquet to CSV increases data size and I/O, worsening memory pressure rather than relieving it. Option C is incorrect because reducing source partitions concentrates more data per partition, increasing per-task memory usage and the risk of OOM. Option E is incorrect because disabling job bookmarks only affects incremental processing state and does not address heap memory consumption.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Change the output format from Parquet to CSV.

    Why it's wrong here

    CSV is row-based and uncompressed, so it inflates I/O and memory compared with columnar Parquet, aggravating the heap exhaustion. Parquet suits analytical scans; CSV would be chosen only when downstream tools demand plain delimited text.

  • ✓

    Increase the Spark shuffle partitions configuration (spark.sql.shuffle.partitions).

    Why this is correct

    Raising spark.sql.shuffle.partitions splits shuffle data into more, smaller partitions, so each task's working set fits within executor heap. This directly addresses the Java heap space exhaustion caused by oversized partitions during wide transformations on the large dataset.

  • ✗

    Reduce the number of partitions in the source data.

    Why it's wrong here

    Reducing partitions concentrates more records into each task, increasing per-executor heap pressure and worsening the OutOfMemoryError. Repartitioning is genuinely useful when small files or excessive partitions cause per-task overhead and scheduling inefficiency, but here the driver and executors need more memory or fewer records per task, not fewer partitions.

  • ✓

    Increase the number of DPUs allocated to the Glue job.

    Why this is correct

    Adding DPUs increases the total executor count and memory available across the cluster, distributing the large dataset's processing load so no single executor exhausts its Java heap. This satisfies the memory constraint the OutOfMemoryError exposes.

  • ✗

    Disable job bookmarks to avoid incremental processing.

    Why it's wrong here

    Disabling job bookmarks increases the volume reprocessed on every run, worsening heap pressure rather than relieving it. Bookmarks exist to process only new data incrementally; they would be the correct choice when the goal is reducing repeated work, not resolving memory errors.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

About these practice questions

This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

1 more way this is tested on DEA-C01

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. A data engineer is troubleshooting an AWS Glue ETL job that fails with a memory error when processing a large dataset. Which approach can help reduce memory usage?

easy
  • A.Set the job to use only one worker
  • B.Reduce the number of partitions in the data source
  • ✓ C.Increase the number of workers for the job
  • D.Increase the worker type to G.2X

Why C: Increasing the number of workers for the job distributes the data processing load across more executors, which reduces the memory burden on each individual worker. In AWS Glue, each worker runs its own Spark executor with a fixed amount of memory; by adding more workers, the data is partitioned across more executors, so each handles a smaller subset of the data. This directly alleviates out-of-memory errors caused by a single worker attempting to process too much data.

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.