Courseiva
Data EngineeringmediumMultiple ChoiceObjective-mapped

Building a Real-Time Fraud Detection Pipeline with Kinesis and SageMaker

A data science team is building a real-time fraud detection system. Transactions are streamed via Amazon Kinesis Data Streams, and a Lambda function performs feature engineering and invokes an Amazon SageMaker endpoint for predictions. The team notices that the Lambda function is timing out and causing data loss. Which solution should the team implement to process the stream reliably and at low latency?

Quick Answer

The answer is to use Amazon Kinesis Data Analytics for Apache Flink to consume the stream, perform feature engineering, and invoke the SageMaker endpoint with exactly-once processing. This is correct because Apache Flink provides a stateful, long-running stream processing engine that eliminates the timeout and data loss issues inherent in short-lived Lambda functions, enabling reliable, low-latency real-time ML inference on streaming data. On the AWS Certified Machine Learning Specialty MLS-C01 exam, this scenario tests your understanding of choosing the right streaming compute service for stateful, exactly-once processing versus stateless, ephemeral functions like Lambda, which are a common trap for simple tasks but fail under sustained throughput. A key memory tip: think “Flink for stateful streaming, Lambda for lightweight triggers”—if the task requires maintaining state or invoking ML models on every record without timeouts, Flink is the durable choice.

⚠ Common exam trap

Many candidates assume increasing Lambda resources (timeout/memory) or moving to a batch-based approach (Firehose/S3) can solve real-time streaming issues, but the exam tests the understanding that stateful, long-running stream processing engines like Flink are required for reliable, low-latency, exactly-once processing in production.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

Use Amazon Kinesis Data Analytics for Apache Flink to consume the stream, perform feature engineering, and invoke the SageMaker endpoint with exactly-once processing.

Amazon Kinesis Data Analytics for Apache Flink provides a stateful, low-latency stream processing engine that can consume from Kinesis Data Streams, perform feature engineering in real-time, and invoke SageMaker endpoints with exactly-once processing semantics. This eliminates Lambda timeouts and data loss by using a long-running, scalable application instead of a short-lived function.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • Use Amazon Kinesis Data Analytics for Apache Flink to consume the stream, perform feature engineering, and invoke the SageMaker endpoint with exactly-once processing.

    Why this is correct

    Kinesis Data Analytics provides stateful stream processing with checkpointing, ensuring no data loss and low-latency integration with SageMaker.

  • Use the Kinesis Client Library (KCL) to process the stream in an Amazon EC2 instance, and store the predictions in Amazon DynamoDB.

    Why it's wrong here

    KCL-based processing on EC2 adds operational overhead and does not directly address Lambda timeout issues; DynamoDB storage is not necessary for the pipeline.

  • Increase the Lambda function timeout to 15 minutes and allocate more memory to reduce processing time.

    Why it's wrong here

    Lambda has a maximum timeout of 15 minutes, but this increases cost and still risks data loss if failures occur; it does not provide checkpointing.

  • Configure Amazon Kinesis Firehose to deliver the stream to an Amazon S3 bucket, then trigger a Lambda function to process the data in batches.

    Why it's wrong here

    Kinesis Firehose introduces minutes of delay, which is unsuitable for real-time fraud detection.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

About these practice questions

One of 1,672 original MLS-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

1 more way this is tested on MLS-C01

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. A data engineering team is building a real-time fraud detection system. Transactions are ingested via Amazon Kinesis Data Streams, and a machine learning model (deployed on Amazon SageMaker) scores each transaction. The team needs to store the raw transactions and the model's predictions in Amazon S3 for later analysis. Which architecture should the team use?

medium
  • A.Use AWS Lambda to read from Kinesis, invoke SageMaker, and write directly to S3.
  • B.Use Amazon Kinesis Data Firehose with a transformation Lambda to call SageMaker.
  • C.Use Amazon Kinesis Data Analytics for Apache Flink to enrich records with SageMaker predictions, then output to Firehose for S3.
  • D.Use AWS Lambda to invoke the SageMaker endpoint for each record, then write to S3 via Firehose.

Why C: It uses Amazon Kinesis Data Analytics for Apache Flink to perform real-time enrichment by invoking the SageMaker endpoint for each transaction, then streams the enriched records to Kinesis Data Firehose for reliable, batched delivery to S3. This architecture handles the asynchronous nature of model inference without blocking the ingestion stream, and Firehose provides automatic retry and compression for S3 storage.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLS-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLS-C01 exam.