Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The engineer needs to improve the performance of the Glue job, which currently takes several hours to complete. The job reads large Parquet files, performs joins, and writes to Redshift. Which two actions should the engineer take to improve performance? (Choose two.)

⚠ Common exam trap

The trap here is assuming that job bookmarks improve performance for a single large run, when they only help avoid reprocessing data across runs.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Partition the source Parquet data in Amazon S3 by commonly filtered columns and use partition pruning in the Glue job.

Increasing DPUs adds compute resources to parallelize the job, and partitioning the source data enables partition pruning to reduce the amount of data read. Both actions directly address performance bottlenecks in a large Glue job. Other options either do not affect single-run performance or would degrade it.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Partition the source Parquet data in Amazon S3 by commonly filtered columns and use partition pruning in the Glue job.

    Why this is correct

    Partitioning the source data by columns used in filters allows Glue to read only relevant partitions, reducing I/O and the amount of data shuffled during joins. This can dramatically improve performance for large datasets. The engineer should ensure the Glue job uses predicate pushdown and that the partition columns are used in the query. This is a best practice for optimizing Glue ETL on S3.

  • ✗

    Enable job bookmarks to avoid reprocessing previously processed data.

    Why it's wrong here

    Job bookmarks help avoid reprocessing old data on subsequent runs, which can reduce runtime for incremental loads. However, for a single job run that processes a large dataset, bookmarks do not speed up the transformation itself. They are useful for incremental processing but do not address the performance of joins or writes within a run. Thus, they are not a primary performance improvement for the described scenario.

  • ✓

    Increase the number of AWS Glue DPUs allocated to the job to add more Spark executors.

    Why this is correct

    Increasing DPUs provides more compute resources, which can parallelize reads, joins, and writes. For large Parquet datasets and joins, additional executors reduce the time spent on shuffles and processing. This is a direct way to improve performance when the job is resource-bound, though it increases cost. The engineer should monitor CloudWatch metrics to ensure the job is not bottlenecked elsewhere.

  • ✗

    Convert the Parquet files to CSV to reduce storage size and speed up reading.

    Why it's wrong here

    CSV is less efficient than Parquet for analytics; Parquet is columnar, compressed, and supports predicate pushdown. Converting to CSV would increase storage size and slow down reads, not improve performance. This would be a step backward and is not a valid optimization for a Glue job reading large datasets.

  • ✗

    Use the 'Relationalize' transform to flatten the Parquet data before joining.

    Why it's wrong here

    Relationalize is for flattening nested data, not for improving join performance on already flat Parquet files. Applying it would add unnecessary processing and likely increase runtime. The scenario describes large Parquet files that are already suitable for joins; Relationalize is not relevant and would not help performance.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.