mediumMultiple Select
MLA-C01 Practice Question: Building a real-time fraud detection system using…
A company is building a real-time fraud detection system using Amazon Kinesis Data Streams. The data must be joined with a reference table (e.g., customer profile) that is stored in Amazon DynamoDB and updated frequently. The enriched data will be used for ML predictions. Which THREE AWS services should the company use to build this streaming pipeline? (Select THREE.)
⚠ Common exam trap
MLA-C01 often tests whether candidates pick Firehose for 'real-time' processing when the requirement involves joins or enrichment — Firehose is delivery, not compute, and cannot join streams to DynamoDB.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Amazon Kinesis Data Analytics for Apache Flink
Amazon Kinesis Data Streams (C) is the ingestion backbone for the real-time fraud events, providing durable, low-latency, ordered streaming with shards and configurable retention so downstream consumers can process records continuously. Amazon Kinesis Data Analytics for Apache Flink (A) is the correct processing engine because it supports stateful stream processing and can perform asynchronous lookups against DynamoDB (via the Flink DynamoDB connector or async I/O) to enrich each event with the frequently updated customer profile before feeding ML predictions. Amazon DynamoDB (E) is the reference table store, offering single-digit-millisecond key-value reads at scale so the Flink job can join each streaming record with current customer profile data. Amazon Kinesis Data Firehose (B) is not appropriate here because it is a managed delivery service for loading data into destinations like S3, Redshift, or Splunk and does not perform stateful joins or ML enrichment. AWS Glue ETL (D) is a batch/serverless Spark-based ETL service and is not designed for sub-second, continuously running stream joins against a rapidly changing DynamoDB table.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Amazon Kinesis Data Analytics for Apache Flink
Why this is correct
Kinesis Data Analytics for Apache Flink performs the stateful stream join between incoming transactions and the DynamoDB reference table, then emits enriched records for ML inference. It satisfies the real-time enrichment requirement by running continuous SQL or Flink queries over streaming data.
- ✗
Amazon Kinesis Data Firehose
Why it's wrong here
Kinesis Data Firehose delivers stream records to destinations such as S3, Redshift or Splunk; it cannot perform the stateful join against a frequently updated DynamoDB reference table that enrichment requires. It is tempting because Firehose is the standard choice for loading streaming data into analytics stores without managing consumers.
- ✓
Amazon Kinesis Data Streams
Why this is correct
Kinesis Data Streams ingests the high-volume real-time transaction events that feed the fraud detection pipeline, providing the durable, ordered transport layer. It satisfies the streaming ingestion constraint, delivering records to downstream enrichment and ML prediction consumers with low latency.
- ✗
AWS Glue ETL
Why it's wrong here
AWS Glue ETL runs batch or micro-batch Spark jobs on schedules or triggers, so it cannot join each incoming Kinesis record against a DynamoDB table at real-time latency. It is tempting because Glue is the usual choice for joining datasets and transforming data in serverless ETL pipelines.
- ✓
Amazon DynamoDB
Why this is correct
DynamoDB supplies the frequently updated customer profile reference table that Kinesis Data Streams records must be joined against. Its low-latency key-value lookups let the enrichment stage retrieve current profile attributes per event, satisfying the requirement for a rapidly changing reference store.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.