DEA-C01 Data Ingestion and Transformation Practice Question
A data engineering team is building a data lake on Amazon S3. They need to ingest data from multiple sources: (1) streaming IoT data, (2) daily CSV exports from an on-premises system via SFTP, and (3) change data capture (CDC) from an Amazon Aurora database. Which THREE services should the team use to ingest these data sources?
⚠ Common exam trap
Watch out — candidates often confuse AWS Glue ETL (a batch ETL tool) with AWS DMS (a database migration and CDC service), and assuming Amazon EMR is an ingestion service rather than a processing framework for large-scale data transformations.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Amazon Kinesis Data Streams for IoT data ingestion.
Amazon Kinesis Data Streams (A) is the right choice for streaming IoT data because it is a fully managed, scalable service designed for real-time ingestion of high-volume streaming records into S3 via Kinesis Data Firehose or consumers. AWS Database Migration Service (B) supports ongoing change data capture (CDC) from Amazon Aurora (and other relational databases), replicating inserts, updates, and deletes continuously, which matches the CDC requirement. AWS Transfer Family (C) provides a fully managed SFTP endpoint backed by Amazon S3, so the on-premises system can push daily CSV exports over SFTP directly into the data lake. AWS Glue ETL (D) is a batch/transform service, not a CDC ingestion mechanism, so it does not satisfy the Aurora CDC requirement. Amazon EMR (E) is a big data processing cluster, not a managed ingestion service for daily SFTP CSV files, so it is not the appropriate ingestion tool here.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Amazon Kinesis Data Streams for IoT data ingestion.
Why this is correct
Kinesis Data Streams ingests continuous, high-volume IoT telemetry with durable ordering and replay, matching the streaming source. It feeds downstream consumers that land the data in Amazon S3, satisfying the real-time ingestion requirement of the data lake.
- ✓
AWS Database Migration Service (DMS) for CDC from Aurora.
Why this is correct
AWS Database Migration Service reads Aurora's replication log to capture ongoing inserts, updates and deletes as CDC, then streams those changes to Amazon S3. This satisfies the continuous change-data-capture requirement without impacting the source database's performance.
- ✓
AWS Transfer Family for SFTP-based file ingestion.
Why this is correct
AWS Transfer Family provides a managed SFTP endpoint backed by Amazon S3, so the on-premises system uploads CSV exports using its existing SFTP client. No servers to patch, satisfying the daily file-ingestion requirement with minimal operational overhead.
- ✗
AWS Glue ETL for CDC from Aurora.
Why it's wrong here
AWS Glue ETL lacks native change data capture (CDC) integration for Aurora; it requires custom logic or additional services like AWS DMS to capture ongoing row-level changes. This option is tempting because Glue ETL is commonly used for batch transformations and could handle initial Aurora snapshots, but it cannot continuously stream CDC events without external tooling.
- ✗
Amazon EMR for daily CSV ingestion.
Why it's wrong here
Amazon EMR processes and analyses data on S3 or HDFS; it does not poll an SFTP endpoint or schedule file transfers. It would be correct for transforming or querying the ingested CSV data at scale, but ingestion from SFTP needs AWS Transfer Family or a scheduled DataSync task instead.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.