DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer needs to run an AWS Glue extract, transform, and load (ETL) job that joins an Amazon S3-based Parquet dataset with a slowly changing dimension table in Amazon Redshift. The Redshift cluster is in a private subnet and cannot be reached over the public internet. The engineer wants the Glue job to read from Redshift without exposing credentials in the job script. Which combination of actions should the engineer take to meet these requirements?
⚠ Common exam trap
The trap here is assuming that adding a NAT gateway to a VPC connection is enough to reach a private Redshift cluster, when the job actually needs network interfaces in the same VPC.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Create an AWS Glue connection of type JDBC to the Redshift cluster, attach the job to a VPC connection in the same VPC, and store the Redshift credentials in AWS Secrets Manager for the connection to use.
A JDBC connection plus a VPC connection gives the Glue job a private network path to the Redshift cluster, which is required because the cluster is in a private subnet. Storing credentials in Secrets Manager and referencing them through the connection avoids hardcoding secrets in the job script. Together these satisfy both the private connectivity and credential management requirements.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Copy the Redshift data to Amazon S3 using an UNLOAD command, then have the Glue job read the S3 copy, and delete the S3 copy after the job completes.
Why it's wrong here
UNLOAD to S3 is a valid pattern for bulk export, but it introduces an extra copy and does not provide the direct read the engineer wants. It also requires a place for the UNLOAD to run and manage lifecycle of the temporary data, adding complexity without addressing the credential-in-script concern directly.
- ✗
Attach the Glue job to a VPC connection with a NAT gateway, store the Redshift credentials in AWS Secrets Manager, and reference the secret in the Glue job's connection options.
Why it's wrong here
A NAT gateway allows outbound internet access but does not provide a private path to a Redshift cluster in a private subnet; the Glue job could still fail to reach the cluster. Secrets Manager alone does not solve the network reachability requirement, so this combination does not satisfy the private-connectivity constraint.
- ✗
Use the Redshift Data API from the Glue job, attaching an IAM role to the job that has redshift-data:ExecuteStatement permissions, and run all transformations in Redshift.
Why it's wrong here
The Redshift Data API is useful for running SQL statements without a persistent JDBC connection, but it does not let the Glue job directly read Redshift tables as a source for Spark transformations. Pushing all transformations into Redshift also changes the architecture and may not be feasible for the Parquet join.
- ✓
Create an AWS Glue connection of type JDBC to the Redshift cluster, attach the job to a VPC connection in the same VPC, and store the Redshift credentials in AWS Secrets Manager for the connection to use.
Why this is correct
A JDBC Glue connection combined with a VPC connection places the Glue job's elastic network interfaces in the same VPC as Redshift, enabling private connectivity. Storing credentials in Secrets Manager lets the connection retrieve them at runtime without embedding them in the script, satisfying both the network and credential requirements.
Visual reference
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.