DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The engineer needs to improve the performance of the Glue job, which currently takes several hours to complete. The job reads large Parquet files, performs joins, and writes to Redshift. Which two actions should the engineer take to improve performance? (Choose two.)
⚠ Common exam trap
The trap here is assuming that job bookmarks improve performance for a single large run, when they only help avoid reprocessing data across runs.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Partition the source Parquet data in Amazon S3 by commonly filtered columns and use partition pruning in the Glue job.
Increasing DPUs adds compute resources to parallelize the job, and partitioning the source data enables partition pruning to reduce the amount of data read. Both actions directly address performance bottlenecks in a large Glue job. Other options either do not affect single-run performance or would degrade it.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Partition the source Parquet data in Amazon S3 by commonly filtered columns and use partition pruning in the Glue job.
Why this is correct
Partitioning the source data by columns used in filters allows Glue to read only relevant partitions, reducing I/O and the amount of data shuffled during joins. This can dramatically improve performance for large datasets. The engineer should ensure the Glue job uses predicate pushdown and that the partition columns are used in the query. This is a best practice for optimizing Glue ETL on S3.
- ✗
Enable job bookmarks to avoid reprocessing previously processed data.
Why it's wrong here
Job bookmarks help avoid reprocessing old data on subsequent runs, which can reduce runtime for incremental loads. However, for a single job run that processes a large dataset, bookmarks do not speed up the transformation itself. They are useful for incremental processing but do not address the performance of joins or writes within a run. Thus, they are not a primary performance improvement for the described scenario.
- ✓
Increase the number of AWS Glue DPUs allocated to the job to add more Spark executors.
Why this is correct
Increasing DPUs provides more compute resources, which can parallelize reads, joins, and writes. For large Parquet datasets and joins, additional executors reduce the time spent on shuffles and processing. This is a direct way to improve performance when the job is resource-bound, though it increases cost. The engineer should monitor CloudWatch metrics to ensure the job is not bottlenecked elsewhere.
- ✗
Convert the Parquet files to CSV to reduce storage size and speed up reading.
Why it's wrong here
CSV is less efficient than Parquet for analytics; Parquet is columnar, compressed, and supports predicate pushdown. Converting to CSV would increase storage size and slow down reads, not improve performance. This would be a step backward and is not a valid optimization for a Glue job reading large datasets.
- ✗
Use the 'Relationalize' transform to flatten the Parquet data before joining.
Why it's wrong here
Relationalize is for flattening nested data, not for improving join performance on already flat Parquet files. Applying it would add unnecessary processing and likely increase runtime. The scenario describes large Parquet files that are already suitable for joins; Relationalize is not relevant and would not help performance.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.