Courseiva

DEA-C01 Data Ingestion and Transformation Practice Question

A data engineering team is building a data lake on Amazon S3. They need to ingest data from multiple sources: (1) streaming IoT data, (2) daily CSV exports from an on-premises system via SFTP, and (3) change data capture (CDC) from an Amazon Aurora database. Which THREE services should the team use to ingest these data sources?

⚠ Common exam trap

Watch out — candidates often confuse AWS Glue ETL (a batch ETL tool) with AWS DMS (a database migration and CDC service), and assuming Amazon EMR is an ingestion service rather than a processing framework for large-scale data transformations.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Amazon Kinesis Data Streams for IoT data ingestion.

Amazon Kinesis Data Streams (A) is the right choice for streaming IoT data because it is a fully managed, scalable service designed for real-time ingestion of high-volume streaming records into S3 via Kinesis Data Firehose or consumers. AWS Database Migration Service (B) supports ongoing change data capture (CDC) from Amazon Aurora (and other relational databases), replicating inserts, updates, and deletes continuously, which matches the CDC requirement. AWS Transfer Family (C) provides a fully managed SFTP endpoint backed by Amazon S3, so the on-premises system can push daily CSV exports over SFTP directly into the data lake. AWS Glue ETL (D) is a batch/transform service, not a CDC ingestion mechanism, so it does not satisfy the Aurora CDC requirement. Amazon EMR (E) is a big data processing cluster, not a managed ingestion service for daily SFTP CSV files, so it is not the appropriate ingestion tool here.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Amazon Kinesis Data Streams for IoT data ingestion.

    Why this is correct

    Kinesis Data Streams ingests continuous, high-volume IoT telemetry with durable ordering and replay, matching the streaming source. It feeds downstream consumers that land the data in Amazon S3, satisfying the real-time ingestion requirement of the data lake.

  • ✓

    AWS Database Migration Service (DMS) for CDC from Aurora.

    Why this is correct

    AWS Database Migration Service reads Aurora's replication log to capture ongoing inserts, updates and deletes as CDC, then streams those changes to Amazon S3. This satisfies the continuous change-data-capture requirement without impacting the source database's performance.

  • ✓

    AWS Transfer Family for SFTP-based file ingestion.

    Why this is correct

    AWS Transfer Family provides a managed SFTP endpoint backed by Amazon S3, so the on-premises system uploads CSV exports using its existing SFTP client. No servers to patch, satisfying the daily file-ingestion requirement with minimal operational overhead.

  • ✗

    AWS Glue ETL for CDC from Aurora.

    Why it's wrong here

    AWS Glue ETL lacks native change data capture (CDC) integration for Aurora; it requires custom logic or additional services like AWS DMS to capture ongoing row-level changes. This option is tempting because Glue ETL is commonly used for batch transformations and could handle initial Aurora snapshots, but it cannot continuously stream CDC events without external tooling.

  • ✗

    Amazon EMR for daily CSV ingestion.

    Why it's wrong here

    Amazon EMR processes and analyses data on S3 or HDFS; it does not poll an SFTP endpoint or schedule file transfers. It would be correct for transforming or querying the ingested CSV data at scale, but ingestion from SFTP needs AWS Transfer Family or a scheduled DataSync task instead.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.