MLS-C01 Data Engineering Practice Question
A data engineering team is designing a data pipeline to process streaming data from social media feeds. The data must be deduplicated, enriched with customer information from a relational database, and stored in Amazon S3 in Parquet format. Which AWS services should the team use to build this pipeline? (Select TWO.)
⚠ Common exam trap
AWS often tests the distinction between data ingestion services (Kinesis Data Firehose) and data processing/ETL services (AWS Glue), leading candidates to mistakenly select Firehose for deduplication and enrichment tasks that it cannot natively perform.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
AWS Glue
AWS Glue is correct because it provides a serverless ETL service that can transform streaming data stored in Amazon S3 into Parquet format. It can also connect to a relational database via JDBC to enrich the data with customer information, and its built-in deduplication capabilities (e.g., using DropDuplicates in PySpark) handle the deduplication requirement.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
AWS Glue
Why this is correct
Glue ETL can transform and enrich data from streams and databases.
- ✗
Amazon Kinesis Data Firehose
Why it's wrong here
Firehose cannot perform enrichment with a relational database.
- ✗
Amazon Athena
Why it's wrong here
Athena is for querying, not ETL.
- ✗
Amazon SageMaker
Why it's wrong here
SageMaker is for ML models, not data pipeline.
- ✓
Amazon Kinesis Data Streams
Why this is correct
Ingests streaming social media data.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLS-C01 question from scratch — 1,672 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLS-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLS-C01 exam.