DEA-C01 Data Ingestion and Transformation Practice Question
A marketing analytics team needs to ingest customer transaction data from an on-premises PostgreSQL database into Amazon S3 for analysis. The data volume is about 10 GB daily, and the team wants to perform full refresh daily (truncate and load) into S3 as Parquet files. The company has a Direct Connect connection to AWS. The team needs a simple, managed solution that minimizes operational overhead. What should the team use?
⚠ Common exam trap
DEA-C01 often tests the choice between DMS and Glue for batch ingestion; DMS is for continuous replication, while Glue is for batch ETL with transformations.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use an AWS Glue ETL job with a JDBC connection to the PostgreSQL database, extract data, and write to S3 in Parquet format.
AWS Glue ETL jobs provide a serverless, managed environment to extract data from JDBC sources like PostgreSQL, transform it, and write to S3 in Parquet format. It minimizes operational overhead as it handles provisioning, scaling, and job execution. For a daily full refresh, a Glue job can be scheduled to truncate and load data.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Set up AWS Database Migration Service (DMS) to continuously replicate data to S3 in Parquet format.
Why it's wrong here
DMS continuous replication applies ongoing change data capture rather than the daily truncate-and-load full refresh the team requires, and its S3 target writes require extra configuration for Parquet. It is tempting because DMS is the managed choice when ongoing replication, not daily full reloads, is the goal.
- ✗
Use Amazon EMR with a Spark job that reads from PostgreSQL and writes to S3.
Why it's wrong here
Configuring and managing EMR clusters introduces significant operational overhead for a 10 GB daily workload, failing the requirement for a managed solution. While Spark on EMR is ideal for complex large-scale data transformations or massive datasets requiring distributed processing power, it requires manual cluster tuning and lifecycle management. In this scenario, AWS Glue provides the necessary serverless orchestration to automate the ETL process without managing underlying infrastructure.
- ✓
Use an AWS Glue ETL job with a JDBC connection to the PostgreSQL database, extract data, and write to S3 in Parquet format.
Why this is correct
AWS Glue is serverless and managed, using a JDBC connection to read PostgreSQL and writing Parquet to S3, which minimises operational overhead for the daily 10 GB full refresh. It satisfies the stem's simplicity and managed-solution constraints without managing servers.
- ✗
Use AWS Data Pipeline with a SQLActivity to extract data and copy to S3.
Why it's wrong here
AWS Data Pipeline has been placed in maintenance mode and cannot natively write Parquet or perform the truncate-and-load refresh described. It is tempting because it once orchestrated scheduled extract-and-copy jobs between on-premises sources and S3, which was its intended purpose before AWS Glue and DMS replaced it.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.