Courseiva
Continuous Improvement for Existing SolutionshardMultiple ChoiceObjective-mapped

Step Functions Execution Timeout: How to Fix Lambda Timeouts in State Machines

A media company runs a video processing pipeline on AWS. The pipeline uses AWS Step Functions to orchestrate multiple AWS Lambda functions. The first Lambda function downloads a video file from an S3 bucket, the second transcodes it using AWS Elemental MediaConvert, and the third uploads the transcoded files to a different S3 bucket. Recently, the pipeline has been failing intermittently with 'State machine execution timed out' errors. The Step Functions execution history shows that the first Lambda function takes up to 25 minutes to complete for large video files. The Step Functions state machine has a default execution timeout of 5 minutes. The company wants to fix the timeout issue without redesigning the entire pipeline. Which solution should the solutions architect recommend?

Quick Answer

The correct answer is to increase the timeoutSeconds value in the Step Functions state machine definition to 1800 seconds or more. This directly resolves the "State machine execution timed out" error because the default execution timeout of 5 minutes is being exceeded by the first Lambda function, which can take up to 25 minutes for large video files. The key technical concept here is that Step Functions has its own execution timeout that is independent of individual Lambda function timeouts; even if you increase the Lambda timeout to its 15-minute maximum, the state machine will still terminate the entire workflow once its own timeout is reached. On the AWS Certified Solutions Architect Professional SAP-C02 exam, this question tests your understanding of the separation between service-level timeouts and orchestration-level timeouts—a common trap is confusing Lambda’s 15-minute limit with the state machine’s configurable timeout. Remember the mnemonic: "State machine timeout is the ceiling; Lambda timeout is the floor."

⚠ Common exam trap

Candidates often forget that Lambda has a hard timeout of 15 minutes. Simply increasing the state machine timeout does not fix the underlying Lambda timeout.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

Replace the Lambda function with an Amazon SQS queue and have the Step Functions wait for a callback.

The Lambda function has a maximum timeout of 15 minutes, so increasing the state machine timeout alone (Option A) does not solve the Lambda function's inability to run for 25 minutes. Replacing the Lambda with an SQS queue allows the long-running task to be processed asynchronously, with Step Functions waiting for a callback, avoiding the Lambda timeout limit and the state machine timeout error. Options B and C are invalid because Lambda cannot be configured beyond 15 minutes: setting it to 30 or 15 minutes still fails for a 25-minute task. Option D changes only the problematic component, meeting the requirement to fix the issue without redesigning the entire pipeline.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • Increase the 'timeoutSeconds' value in the Step Functions state machine definition to 1800 (30 minutes) or more.

    Why it's wrong here

    Increasing the state machine timeout to 30 minutes does not solve the issue because the Lambda function itself has a maximum timeout of 15 minutes. The function will time out before the state machine does, and execution will still fail.

  • Increase the Lambda function timeout to 30 minutes in the Lambda configuration.

    Why it's wrong here

    The Lambda function timeout is separate from the state machine timeout; increasing it does not prevent the state machine from timing out after 5 minutes.

  • Increase the Lambda function timeout to 15 minutes and increase the state machine execution timeout to 30 minutes.

    Why it's wrong here

    Lambda has a maximum timeout of 15 minutes, so setting it to 15 minutes does not cover the 25-minute runtime; the state machine timeout must be increased as in option A.

  • Replace the Lambda function with an Amazon SQS queue and have the Step Functions wait for a callback.

    Why this is correct

    Replacing the Lambda function with an SQS queue and using a callback pattern allows the long-running download task to be processed asynchronously by a worker that can run for more than 15 minutes, without redesigning the entire pipeline. Step Functions waits for the callback, avoiding both the Lambda timeout and the state machine execution timeout.

Visual reference

Client Server SYN (seq=100) SYN-ACK (seq=200, ack=101) ACK (ack=201) Connection established — data transfer begins

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every SAP-C02 question from scratch — 1,660 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

1 more way this is tested on SAP-C02

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. A media company runs a video processing pipeline on AWS. Videos are uploaded to an S3 bucket (input-bucket), which triggers an AWS Lambda function that starts an AWS Glue job. The Glue job processes the video metadata and stores results in a DynamoDB table. Then, a second Lambda function triggers an Amazon ECS Fargate task to transcode the video into multiple formats. The transcoded videos are stored in another S3 bucket (output-bucket). Recently, the company started receiving complaints about delays in video availability. The operations team notices that CloudWatch Logs show no errors, but the ECS tasks often take longer than expected. They also see that the DynamoDB table has a high number of throttled write events. The video upload rate has increased by 50% in the last month. The team needs to improve the pipeline's performance and reduce delays. What should they do?

hard
  • A.Enable DynamoDB auto scaling on the table with a target utilization of 70%.
  • B.Increase the Lambda function timeout for both functions to 15 minutes.
  • C.Introduce an Amazon SQS queue between the second Lambda and ECS to buffer requests.
  • D.Set reserved concurrency on the first Lambda function to 10 to control throttling.

Why A: The primary bottleneck is DynamoDB throttling due to increased write load. Enabling DynamoDB auto scaling (Option A) dynamically adjusts read/write capacity to match demand, reducing throttling and delays. Option B (increasing Lambda timeout) does not address DynamoDB throttling. Option C (SQS queue) improves decoupling but does not directly solve the DynamoDB issue. Option D (reserved concurrency) limits Lambda concurrency, which could reduce load on DynamoDB but also slows down processing and is not the best solution.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This SAP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SAP-C02 exam.