A company collects clickstream data from millions of users in real time and needs to process and analyse this data as it arrives — not after storing it — to detect patterns within seconds. Which AWS service is designed for real-time data streaming and processing?
Amazon Kinesis Data Streams is purpose-built for real-time streaming ingestion of high-volume data, such as clickstreams, IoT telemetry, and application logs. It uses shards to scale throughput and supports multiple consumers reading the same stream concurrently via the Kinesis Client Library (KCL), enabling sub-second analytics, pattern detection, and live dashboards. Data is durably stored for up to 365 days, allowing replay and processing by multiple applications independently. Its design directly matches the requirement for real-time streaming data ingestion and consumption.
Why this answer
Amazon Kinesis Data Streams is purpose-built for real-time data streaming and processing, enabling you to ingest and analyze clickstream data as it arrives with sub-second latency. It supports custom processing using AWS Lambda, Kinesis Data Analytics, or Kinesis Data Firehose, making it ideal for detecting patterns in real-time without waiting for data to be stored.
Exam trap
The trap here is that candidates often confuse Amazon SQS or DynamoDB Streams as streaming services, but SQS is a queue for decoupling and DynamoDB Streams is a change data capture mechanism, neither of which is designed for real-time, high-throughput data streaming and processing like Kinesis Data Streams.
How to eliminate wrong answers
Option A is wrong because Amazon SQS is a message queue service designed for decoupling application components and asynchronous message delivery, not for real-time streaming analytics or processing of high-throughput clickstream data. Option B is wrong because Amazon S3 is an object storage service that stores data after it is collected, not designed for real-time processing or streaming ingestion. Option D is wrong because Amazon DynamoDB Streams captures changes to DynamoDB tables in near-real-time but is limited to change data capture from a single table, not designed for ingesting and processing high-volume, real-time clickstream data from millions of users.