SAA-C03 Design High-Performing Architectures Practice Question
An API team runs an AWS Lambda function behind an Application Load Balancer (ALB). During predictable hourly traffic spikes, p95 response latency increases due to occasional cold starts. The team wants stable latency during those spikes without permanently overprovisioning resources for all functions. Which configuration is the most appropriate way to reduce cold starts for this Lambda function?
⚠ Common exam trap
Test-takers frequently confuse reserved concurrency (which limits concurrency but does not prevent cold starts) with provisioned concurrency (which pre-warms environments), leading candidates to select Option C as a cost-saving measure that fails to address latency.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Publish a version of the function and configure provisioned concurrency on an alias, using autoscaling for the alias.
Provisioned concurrency initializes a specified number of execution environments in advance, keeping them warm and ready to handle requests without cold start latency. By configuring provisioned concurrency on an alias with autoscaling, the team can dynamically adjust the number of pre-warmed environments to match predictable traffic spikes, avoiding permanent overprovisioning while ensuring stable p95 latency.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Publish a version of the function and configure provisioned concurrency on an alias, using autoscaling for the alias.
Why this is correct
Provisioned concurrency pre-initializes execution environments for a specific published function version. By attaching provisioned concurrency to an alias, you can control warm capacity and (with the right settings) autoscale the provisioned capacity for predictable spike patterns, reducing cold-start-driven latency increases.
- ✗
Increase the function memory size and rely on faster initialization to reduce cold starts.
Why it's wrong here
Allocating more memory to the Lambda function increases the proportional CPU allocation and can shorten the time needed for certain CPU-bound initialization tasks, such as loading libraries or connecting to databases. However, memory does not pre-provision any execution environments, so during a sharp traffic spike each new concurrent invocation still requires a full cold start—including container spin-up, runtime bootstrapping, and code download—which remains unpredictable. The option also lacks any mechanism to scale warm capacity with demand, so it does not address the specific latency spike pattern the team is trying to eliminate.
When this WOULD be correct
A question where the goal is to reduce the duration of cold starts (e.g., 'Which configuration minimizes the impact of cold starts on latency?') and the candidate is constrained from using provisioned concurrency due to cost or other restrictions.
- ✗
Set reserved concurrency equal to the expected peak requests per second for the function.
Why it's wrong here
Setting reserved concurrency to match expected peak requests per second does not warm any execution environments; it merely defines an upper bound on the number of simultaneous invocations for the function and reserves that capacity against other functions in the account. In fact, if the reserved concurrency value is still lower than the actual burst of concurrent requests, lambda will throttle new invocations, causing request failures rather than preventing cold starts. Reserved concurrency is a governance and throttling control, not a way to maintain pre-initialized runtime instances, so it leaves the cold-start latency problem entirely unresolved.
When this WOULD be correct
In a scenario where a Lambda function must not exceed a specific concurrency limit to avoid throttling downstream resources (e.g., a database with connection limits), and the goal is to cap concurrent invocations rather than reduce latency.
- ✗
Use an event source mapping with a higher batch size so Lambda triggers earlier and keeps the runtime warm.
Why it's wrong here
Event source mappings apply to event sources such as SQS, Kinesis, or DynamoDB streams. For an ALB-triggered Lambda (synchronous request/response), batch size is not a mechanism to guarantee warm environments or to mitigate cold starts.
When this WOULD be correct
For a Lambda function processing messages from an SQS queue or DynamoDB Streams, increasing the batch size can reduce the number of invocations and keep the runtime warm by processing more records per invocation, thereby mitigating cold starts during traffic spikes.
Option-by-option analysis
Why each answer is right or wrong
Understanding why wrong answers are wrong — and when they would be correct — is what separates a 750 score from a 900. The SAA-C03 exam frequently reuses these exact scenarios with slightly different constraints.
✓Publish a version of the function and configure provisioned concurrency on an alias, using autoscaling for the alias.Correct answer▾
Why this is correct
Provisioned concurrency pre-initializes execution environments for a specific published function version. By attaching provisioned concurrency to an alias, you can control warm capacity and (with the right settings) autoscale the provisioned capacity for predictable spike patterns, reducing cold-start-driven latency increases.
✗Increase the function memory size and rely on faster initialization to reduce cold starts.Wrong answer — click to see why▾
Why this is wrong here
Increasing memory size improves CPU speed and can reduce initialization time, but it does not eliminate cold starts; it only shortens them. The question asks to reduce cold starts, not just mitigate their duration.
★ When this WOULD be the correct answer
A question where the goal is to reduce the duration of cold starts (e.g., 'Which configuration minimizes the impact of cold starts on latency?') and the candidate is constrained from using provisioned concurrency due to cost or other restrictions.
Why candidates choose this
Candidates know that more memory often means faster Lambda execution, and they may incorrectly assume that faster initialization eliminates cold starts entirely, rather than just shortening them.
✗Set reserved concurrency equal to the expected peak requests per second for the function.Wrong answer — click to see why▾
Why this is wrong here
Reserved concurrency limits the maximum number of concurrent executions but does not pre-warm instances; cold starts still occur during traffic spikes when new execution environments are needed.
★ When this WOULD be the correct answer
In a scenario where a Lambda function must not exceed a specific concurrency limit to avoid throttling downstream resources (e.g., a database with connection limits), and the goal is to cap concurrent invocations rather than reduce latency.
Why candidates choose this
Candidates may confuse reserved concurrency with provisioned concurrency, thinking that reserving capacity eliminates cold starts, but reserved concurrency only guarantees available capacity, not pre-initialized environments.
✗Use an event source mapping with a higher batch size so Lambda triggers earlier and keeps the runtime warm.Wrong answer — click to see why▾
Why this is wrong here
Event source mappings (e.g., SQS, DynamoDB Streams) are not used with ALB triggers; ALB invokes Lambda synchronously via a function URL or alias ARN. Increasing batch size does not apply to ALB-triggered functions and does not prevent cold starts during predictable spikes.
★ When this WOULD be the correct answer
For a Lambda function processing messages from an SQS queue or DynamoDB Streams, increasing the batch size can reduce the number of invocations and keep the runtime warm by processing more records per invocation, thereby mitigating cold starts during traffic spikes.
Why candidates choose this
Candidates may confuse event source mappings with ALB triggers and think that batching can pre-warm functions, not realizing that ALB invokes Lambda synchronously without batch processing.
Analysis generated from the official SAA-C03blueprint and verified against question context. The “when correct” sections are what AI assistants cite when candidates ask “what’s the difference between these options?”
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
This SAA-C03 question is part of Courseiva's 935-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This SAA-C03 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SAA-C03 exam.