DEA-C01 Data Ingestion and Transformation Practice Question
A data engineer is designing a pipeline to ingest data from an Amazon RDS for PostgreSQL database into Amazon S3 using AWS Database Migration Service (AWS DMS). The source database has a high volume of transactions and the engineer needs to capture ongoing changes with minimal impact on the source. The target S3 bucket must store the data in Parquet format for querying with Amazon Athena. Which two actions should the engineer take to meet these requirements? (Choose two.)
⚠ Common exam trap
The trap here is assuming that a periodic full load is sufficient for ongoing changes, when it actually causes high source load and misses intermediate changes; and overlooking that DMS can natively output Parquet to S3.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use AWS DMS to replicate directly to Amazon S3 in Parquet format by selecting the Parquet output option in the target endpoint settings.
To capture ongoing changes from PostgreSQL with minimal impact, AWS DMS must use CDC, which requires logical replication on the source and a replication instance sized appropriately. To store data in Parquet for Athena, the DMS target endpoint for S3 must be configured to output Parquet. These two actions together satisfy the requirements: near-real-time replication with low source overhead and a columnar format optimized for query performance.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable multi-AZ on the replication instance and use a single large replication instance to handle all tables.
Why it's wrong here
Multi-AZ improves availability but does not directly address change capture or minimize source impact. Using a single large instance may be sufficient, but the key requirements are CDC and Parquet output. Multi-AZ is not required for the stated goals and does not replace the need for logical replication and proper target configuration.
- ✓
Use AWS DMS to replicate directly to Amazon S3 in Parquet format by selecting the Parquet output option in the target endpoint settings.
Why this is correct
AWS DMS supports Amazon S3 as a target and can write data in Parquet format when configured in the endpoint settings. This eliminates the need for a separate conversion step and allows Athena to query the data directly. The engineer must specify the appropriate serialization format and ensure the S3 bucket has the necessary permissions.
- ✗
Set up a full load only and schedule it to run every hour to capture changes.
Why it's wrong here
A full load every hour would repeatedly read the entire table, causing significant load on the source database and increasing latency. It also does not capture intermediate changes, leading to data loss or inconsistency. CDC is the appropriate method for ongoing changes with minimal impact.
- ✗
Configure AWS DMS to write to Amazon S3 in CSV format and then use AWS Glue to convert the data to Parquet.
Why it's wrong here
While this approach can work, it adds an extra ETL step and increases complexity and cost. Since DMS can write Parquet directly to S3, this intermediate conversion is unnecessary. The requirement is to store data in Parquet for Athena, which DMS can do natively when configured correctly.
- ✓
Configure AWS DMS to use change data capture (CDC) with a replication instance that has sufficient resources and enable logical replication on the source PostgreSQL database.
Why this is correct
CDC allows DMS to capture ongoing changes without repeatedly querying the entire table, minimizing impact on the source. For PostgreSQL, logical replication must be enabled (e.g., setting wal_level to logical) and the replication instance must have enough CPU and memory to handle the transaction volume. This setup ensures near-real-time replication with low overhead.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
One of 1,321 original DEA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.