Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A marketing analytics team needs to ingest customer transaction data from an on-premises PostgreSQL database into Amazon S3 for analysis. The data volume is about 10 GB daily, and the team wants to perform full refresh daily (truncate and load) into S3 as Parquet files. The company has a Direct Connect connection to AWS. The team needs a simple, managed solution that minimizes operational overhead. What should the team use?

⚠ Common exam trap

DEA-C01 often tests the choice between DMS and Glue for batch ingestion; DMS is for continuous replication, while Glue is for batch ETL with transformations.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use an AWS Glue ETL job with a JDBC connection to the PostgreSQL database, extract data, and write to S3 in Parquet format.

AWS Glue ETL jobs provide a serverless, managed environment to extract data from JDBC sources like PostgreSQL, transform it, and write to S3 in Parquet format. It minimizes operational overhead as it handles provisioning, scaling, and job execution. For a daily full refresh, a Glue job can be scheduled to truncate and load data.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Set up AWS Database Migration Service (DMS) to continuously replicate data to S3 in Parquet format.

    Why it's wrong here

    DMS continuous replication applies ongoing change data capture rather than the daily truncate-and-load full refresh the team requires, and its S3 target writes require extra configuration for Parquet. It is tempting because DMS is the managed choice when ongoing replication, not daily full reloads, is the goal.

  • ✗

    Use Amazon EMR with a Spark job that reads from PostgreSQL and writes to S3.

    Why it's wrong here

    Configuring and managing EMR clusters introduces significant operational overhead for a 10 GB daily workload, failing the requirement for a managed solution. While Spark on EMR is ideal for complex large-scale data transformations or massive datasets requiring distributed processing power, it requires manual cluster tuning and lifecycle management. In this scenario, AWS Glue provides the necessary serverless orchestration to automate the ETL process without managing underlying infrastructure.

  • ✓

    Use an AWS Glue ETL job with a JDBC connection to the PostgreSQL database, extract data, and write to S3 in Parquet format.

    Why this is correct

    AWS Glue is serverless and managed, using a JDBC connection to read PostgreSQL and writing Parquet to S3, which minimises operational overhead for the daily 10 GB full refresh. It satisfies the stem's simplicity and managed-solution constraints without managing servers.

  • ✗

    Use AWS Data Pipeline with a SQLActivity to extract data and copy to S3.

    Why it's wrong here

    AWS Data Pipeline has been placed in maintenance mode and cannot natively write Parquet or perform the truncate-and-load refresh described. It is tempting because it once orchestrated scheduled extract-and-copy jobs between on-premises sources and S3, which was its intended purpose before AWS Glue and DMS replaced it.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.