DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The job must handle upserts (inserts and updates) into an existing Redshift table based on a primary key. The engineer needs to ensure that the job efficiently processes only changed records and minimizes data movement. Which two AWS Glue features or techniques should be used to achieve this? (Choose two.)
⚠ Common exam trap
The trap here is assuming AWS Glue has native upsert capabilities for Redshift, when in fact upserts require manual staging and merge logic.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable job bookmarks to track previously processed data and only process new or changed files in Amazon S3.
To efficiently upsert data into Redshift, job bookmarks reduce processing to only new or changed S3 files, and staging the data in S3 followed by a COPY into a staging table and a MERGE operation applies changes without full reloads. These techniques together minimize data movement and ensure only changed records are processed.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Enable job bookmarks to track previously processed data and only process new or changed files in Amazon S3.
Why this is correct
AWS Glue job bookmarks maintain state information about data that has already been processed. When reading from Amazon S3, bookmarks allow the job to process only new files or files that have changed since the last run. This reduces the amount of data read and transformed, improving efficiency and minimizing data movement, which is essential for incremental upserts.
- ✗
Use the 'write_dynamic_frame.from_jdbc_conf' method with the 'preactions' and 'postactions' parameters to run SQL commands before and after writing.
Why it's wrong here
This option is incorrect because while 'preactions' and 'postactions' can run SQL, they are not designed for upserts. They are typically used for tasks like dropping temporary tables or granting permissions. They do not handle row-level upserts and would not efficiently process only changed records, potentially leading to full table reloads or complex manual logic.
- ✓
Stage the transformed data in Amazon S3 and use a Redshift COPY command to load into a staging table, then perform a MERGE operation.
Why this is correct
A common pattern for upserts in Redshift is to stage the data in Amazon S3, use the COPY command to load it into a staging table, and then execute a MERGE (or DELETE and INSERT) operation to apply changes to the target table. This approach minimizes data movement by only loading changed records and efficiently updates the target table based on the primary key.
- ✗
Use the 'glueContext.write_dynamic_frame.from_options' with 'upsert' mode and specify the primary key.
Why it's wrong here
AWS Glue's 'write_dynamic_frame.from_options' does not have a built-in 'upsert' mode. While Glue can write to Redshift, upserts require additional steps such as staging the data and then running a merge operation. This method alone would not handle upserts and could result in duplicates or errors.
- ✗
Configure the Glue job to use the 'redshift-upsert' connection type and set the 'mergeKey' parameter.
Why it's wrong here
AWS Glue does not have a 'redshift-upsert' connection type or a 'mergeKey' parameter. Upserts must be implemented manually using staging tables and SQL operations. Relying on non-existent features would result in job failure and not achieve the required upsert functionality.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.