Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data pipeline uses AWS Glue to process large CSV files. The team notices that some jobs fail with out-of-memory errors. Which TWO configuration changes can help mitigate this issue?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Increase the number of DPUs for the Glue job.

Options B and C are correct: increasing the number of DPUs provides more memory, and enabling autoscaling allows the job to automatically scale resources as needed. Option A (reducing DPUs) would worsen the problem by limiting resources. Option D (converting to Parquet) can improve performance but is not a direct configuration change for the Glue job itself. Option E (job bookmarks) is for incremental processing and does not affect memory.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Reduce the number of DPUs to limit concurrency.

    Why it's wrong here

    Fewer DPUs shrink total executor memory and parallelism, worsening out-of-memory failures rather than relieving them; DPUs are the processing capacity unit, so reducing them lowers available memory. Scaling DPUs upward is the correct lever when jobs exhaust memory on large datasets.

  • ✓

    Increase the number of DPUs for the Glue job.

    Why this is correct

    Adding DPUs allocates more executors and memory per worker, directly relieving the heap pressure causing out-of-memory failures when processing large CSV files. Horizontal scaling suits Glue's distributed Spark engine, satisfying the stem's memory constraint without altering the transformation logic itself.

  • ✓

    Enable Glue job autoscaling.

    Why this is correct

    Glue autoscaling dynamically adds or removes workers based on the job's stage-level workload, so memory-intensive shuffle and aggregation stages receive extra executors rather than failing. This directly addresses the out-of-memory constraint by scaling capacity to demand, instead of relying on a fixed worker count sized for average load.

  • ✗

    Convert input files from CSV to Parquet.

    Why it's wrong here

    Parquet is columnar and compressed, reducing bytes read and memory footprint, but the stem asks for configuration changes to a Glue job, and file format conversion is a data preparation step, not a job configuration. It suits long-term storage and query cost optimisation.

  • ✗

    Enable job bookmarks.

    Why it's wrong here

    Job bookmarks track previously processed S3 objects to avoid reprocessing, which addresses incremental runs, not memory exhaustion during a single large CSV read. They are correct when a job must skip already-handled data across repeated scheduled executions.

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.