Courseiva

DEA-C01 Data Operations and Support Practice Question

A data engineer is optimizing an AWS Glue ETL job that reads from Amazon S3 and writes to Amazon Redshift. The job currently uses a single large file and takes hours to complete. The engineer wants to improve performance by using partitioning and parallelism. Which TWO actions should the engineer take? (Choose two.)

⚠ Common exam trap

The trap here is focusing on incremental processing features like job bookmarks or schema handling with DynamicFrames, which do not directly speed up a single large dataset.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Convert the input data to a partitioned format, such as Parquet, and store it in S3 with a partition key.

To improve performance, the engineer should partition the input data and use a columnar format like Parquet, which enables Glue to read only necessary data and process partitions in parallel. Additionally, increasing the number of DPUs provides more compute resources to handle the parallel tasks. Together, these actions address both data layout and compute capacity, leading to faster job execution.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Configure the job to write output to a single file to reduce the number of S3 PUT requests.

    Why it's wrong here

    Writing to a single file reduces the number of output files but can create a bottleneck and limit parallelism during the write phase. It may also cause memory issues if the data is large. This approach contradicts the goal of improving performance through parallelism. It is better to write multiple partitioned files to leverage concurrent writes.

  • ✓

    Convert the input data to a partitioned format, such as Parquet, and store it in S3 with a partition key.

    Why this is correct

    Partitioning the data and using a columnar format like Parquet allows Glue to read only relevant partitions and leverage predicate pushdown. This reduces I/O and enables parallel processing across partitions. It directly addresses the slow performance caused by a single large file by splitting data into manageable chunks that can be processed concurrently.

  • ✗

    Enable job bookmarks to track processed data and avoid reprocessing.

    Why it's wrong here

    Job bookmarks help with incremental processing by tracking previously processed data, but they do not directly improve performance for a single large file. They reduce redundant work in subsequent runs but do not address parallelism or partitioning for the current dataset. The scenario focuses on speeding up the current job, not incremental processing.

  • ✓

    Increase the number of DPUs allocated to the Glue job to allow more parallel tasks.

    Why this is correct

    Increasing DPUs provides more compute resources, which allows Glue to run more concurrent tasks and process data in parallel. This is effective when the job is resource-bound. However, it should be combined with data partitioning to fully utilize the additional capacity. It directly improves performance by scaling the job's execution environment.

  • ✗

    Use the Glue DynamicFrame instead of a DataFrame to enable automatic schema inference.

    Why it's wrong here

    DynamicFrame provides schema flexibility and handles inconsistencies, but it does not inherently improve performance for large datasets. In fact, DynamicFrames can have overhead compared to DataFrames. The scenario requires performance optimization through partitioning and parallelism, not schema handling. This choice does not address the root cause of slow processing.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.