Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineer needs to ingest JSON data from an on-premises relational database into Amazon S3 every hour. Which AWS service should be used to set up a scheduled, incremental data transfer?

⚠ Common exam trap

Watch out — candidates often confuse AWS Glue's ETL capabilities with DMS's managed database migration, assuming Glue's JDBC connections can handle incremental transfers, but Glue lacks built-in change data capture and requires custom logic for scheduled incremental loads.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

AWS Database Migration Service (DMS) with S3 as target.

AWS DMS is purpose-built for migrating databases to AWS targets, including Amazon S3. It supports ongoing replication (change data capture) and scheduled full-load tasks, making it ideal for hourly incremental transfers from an on-premises relational database to S3 without custom scripting.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Amazon S3 Transfer Acceleration with a cron job.

    Why it's wrong here

    Transfer Acceleration only speeds up transfers to S3 over long distances; it does not connect to an on-premises relational database or perform incremental extraction. It tempts because it accelerates S3 uploads. AWS Database Migration Service with a scheduled task performs the required incremental, hourly database-to-S3 ingestion.

  • ✓

    AWS Database Migration Service (DMS) with S3 as target.

    Why this is correct

    AWS DMS performs ongoing, incremental replication from on-premises relational databases to Amazon S3, tracking changes via CDC rather than full reloads. Scheduled hourly tasks satisfy the stem's requirement for recurring incremental transfer, which SCT or DataSync alone cannot provide from a database.

  • ✗

    AWS Glue with a JDBC connection and a scheduled crawler.

    Why it's wrong here

    A scheduled AWS Glue crawler discovers schema and populates the Data Catalog, but it does not perform incremental data ingestion from a JDBC source; it lacks the change-data-capture mechanism required to transfer only new or updated rows each hour. This option is tempting because Glue with a JDBC connection can extract full table snapshots, which would be correct for a one-time or daily bulk load where schema evolution tracking is the primary need.

  • ✗

    Amazon Kinesis Data Firehose with a database source.

    Why it's wrong here

    Kinesis Data Firehose has no relational database source; it ingests from producers such as Kinesis Data Streams, MSK or direct PUT. Firehose suits streaming delivery to S3, Redshift or OpenSearch, not scheduled incremental extraction from an on-premises database.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This DEA-C01 question is part of Courseiva's 1,321-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

1 more way this is tested on DEA-C01

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. A data engineer needs to ingest data from multiple on-premises relational databases into Amazon S3 for analytics. The data must be transformed and loaded daily. Which THREE AWS services should the engineer use together to build this pipeline? (Choose THREE.)

easy
  • ✓ A.AWS Glue
  • ✓ B.AWS Glue Data Catalog
  • C.Amazon Athena
  • ✓ D.AWS Database Migration Service (DMS)
  • E.Amazon Kinesis Data Streams

Why A: AWS Glue is correct because it provides a serverless ETL (Extract, Transform, Load) service that can read data from Amazon S3, apply transformations (e.g., using PySpark or Scala), and write the transformed data back to S3. In this pipeline, AWS Glue jobs can be scheduled to run daily to perform the required transformations on the ingested data.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.