A data engineer notices that an AWS Glue job processing data from an Amazon S3 bucket frequently fails with 'OutOfMemoryError'. The job reads CSV files, applies transformations, and writes Parquet to another S3 bucket. The job has 10 workers of type G.1X. Which change is MOST likely to resolve the issue?
Trap 1: Increase the number of workers to 20
Adding workers increases parallel task slots, but each executor still holds its partition in memory, so a skewed or oversized partition still exhausts a single executor's heap. Scaling worker count suits throughput-bound jobs with many small partitions, not per-executor memory pressure.
Trap 2: Change the worker type from G.1X to G.8X
G.8X provides more memory per worker, yet the OutOfMemoryError typically arises from a few oversized partitions or a skewed key concentrated on one executor, which a larger worker only postpones. Upgrading worker type suits uniformly memory-heavy transformations, not partition skew.
Trap 3: Enable the Spark UI to monitor memory and tune the job
Enabling the Spark UI gives visibility into memory usage and stage behaviour but changes no configuration, so the failing job still exhausts executor memory. Monitoring is appropriate for diagnosing an unknown bottleneck before tuning, not as the remediation itself when the cause is already an OutOfMemoryError.
- A
Change the worker type from G.1X to G.2X
G.2X workers provide double the memory and vCPU of G.1X, giving each executor more heap for the CSV read, transformation and Parquet write stages. This extra memory directly addresses the OutOfMemoryError without changing job logic.
- B
Increase the number of workers to 20
Why it fails: Adding workers increases parallel task slots, but each executor still holds its partition in memory, so a skewed or oversized partition still exhausts a single executor's heap. Scaling worker count suits throughput-bound jobs with many small partitions, not per-executor memory pressure.
- C
Change the worker type from G.1X to G.8X
Why it fails: G.8X provides more memory per worker, yet the OutOfMemoryError typically arises from a few oversized partitions or a skewed key concentrated on one executor, which a larger worker only postpones. Upgrading worker type suits uniformly memory-heavy transformations, not partition skew.
- D
Enable the Spark UI to monitor memory and tune the job
Why it fails: Enabling the Spark UI gives visibility into memory usage and stage behaviour but changes no configuration, so the failing job still exhausts executor memory. Monitoring is appropriate for diagnosing an unknown bottleneck before tuning, not as the remediation itself when the cause is already an OutOfMemoryError.