Real-Time Streaming with Apache Flink on AWS
A company is designing a new application that will process streaming data from thousands of IoT devices. The data must be ingested in real time and then processed using Apache Flink. Which services should be used? (Choose TWO.)
Quick Answer
The correct answer is Amazon Kinesis Data Streams for ingestion and Amazon Kinesis Data Analytics for Apache Flink for processing. Kinesis Data Streams is the ideal ingestion layer because it provides durable, scalable shards that can handle the high throughput from thousands of IoT devices while retaining data for up to 365 days, enabling Apache Flink to consume records with exactly-once semantics and low latency. On the SAP-C02 exam, this pairing tests your understanding of real-time streaming ingestion and Apache Flink processing within AWS’s managed ecosystem, often appearing in scenarios that require decoupling data capture from compute. A common trap is selecting Amazon MSK or Kinesis Data Firehose—MSK adds operational overhead for Kafka management, while Firehose lacks the replay and fine-grained consumption needed for Flink’s stateful processing. Remember the mnemonic “Streams for streams, Analytics for Flink” to quickly recall that Kinesis Data Streams handles ingestion, and Kinesis Data Analytics for Apache Flink handles the processing logic.
⚠ Common exam trap
Many candidates confuse Kinesis Data Firehose with Kinesis Data Streams, not realizing that Firehose is a delivery service that does not support Apache Flink's requirement for per-record replay and checkpointing, while Data Streams provides the necessary persistent, ordered stream.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Amazon Kinesis Data Streams
Amazon Kinesis Data Streams (A) is correct because it provides a highly scalable, real-time ingestion service that can capture data from thousands of IoT devices with low latency, making it ideal for streaming ingestion. Amazon Kinesis Data Analytics for Apache Flink (C) is correct because it is the managed service that runs Apache Flink applications to process and analyze streaming data in real time, directly integrating with Kinesis Data Streams as a source. AWS Lambda (B) is not designed for continuous stream processing with Apache Flink and is better suited for event-driven, short-lived functions. Amazon Kinesis Data Firehose (D) is primarily for loading streaming data into destinations like S3, Redshift, or Elasticsearch, not for running Flink processing. Amazon SQS (E) is a message queue for decoupling components, not a real-time streaming ingestion or Flink processing service.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Amazon Kinesis Data Streams
Why this is correct
Amazon Kinesis Data Streams ingests thousands of IoT device events in real time with low latency and high throughput, satisfying the stem's real-time ingestion constraint. It integrates natively with Apache Flink through the Kinesis connector, letting Flink consume the stream directly for processing without intermediate storage or batch staging.
- ✗
AWS Lambda
Why it's wrong here
Lambda caps execution at 15 minutes and lacks the stateful, checkpointed operators Apache Flink needs for windowed stream processing, so it cannot host the Flink application itself. It is tempting because Lambda handles event-driven ingestion well, and would be correct for lightweight per-record transformation rather than running Flink.
- ✓
Amazon Kinesis Data Analytics for Apache Flink
Why this is correct
Amazon Kinesis Data Analytics for Apache Flink runs managed Flink applications, satisfying the real-time processing requirement. It integrates natively with Kinesis Data Streams ingestion, so thousands of IoT devices stream in and Flink transforms the data continuously without provisioning clusters.
- ✗
Amazon Kinesis Data Firehose
Why it's wrong here
Kinesis Data Firehose delivers batched records to destinations such as S3, Redshift or Splunk; it cannot run Apache Flink applications, so the required stream processing is impossible. It is tempting because it ingests streaming data, and it would be correct for loading raw device data into a data lake without transformation.
- ✗
Amazon Simple Queue Service (SQS)
Why it's wrong here
SQS is a pull-based queue for decoupling components, not a real-time streaming ingestion service; it lacks the ordered shards and consumer model Apache Flink requires. It would suit asynchronous task distribution between microservices, where durability and backpressure matter more than continuous low-latency stream delivery.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every SAP-C02 question from scratch — 984 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
Same concept, more angles
1 more way this is tested on SAP-C02
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. A company is designing a new application that will process streaming data from IoT devices. The data must be ingested in real-time and stored in Amazon S3 for long-term analytics. Which AWS service should be used to ingest the streaming data?
easy- A.Amazon Simple Notification Service (SNS)
- B.Amazon Simple Queue Service (SQS)
- C.AWS Database Migration Service (DMS)
- ✓ D.Amazon Kinesis Data Streams
Why D: Amazon Kinesis Data Streams is designed for real-time data ingestion and can stream data directly to Amazon S3. Option A is wrong because SNS is a pub/sub messaging service, not intended for real-time data ingestion. Option B is wrong because SQS is a message queue service, not optimized for streaming ingestion. Option C is wrong because AWS DMS is used for database migration, not for ingesting streaming data.
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This SAP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SAP-C02 exam.