SAA-C03 Design Cost-Optimized Architectures Practice Question
A latency-sensitive API is implemented with AWS Lambda. The team enabled provisioned concurrency to avoid cold starts, setting provisioned concurrency to 50 because marketing campaigns occasionally cause spikes. However, during most weekdays the API receives little traffic (near zero), and the team is seeing high monthly Lambda costs from idle provisioned capacity. What is the best cost-optimized strategy that still meets the requirement of fast initial responses during traffic spikes?
⚠ Common exam trap
Many candidates assume provisioned concurrency must be set to a static high value to handle spikes, ignoring AWS's native Auto Scaling capabilities that can dynamically adjust capacity based on schedule or metrics, thus missing the cost-optimization aspect of the question.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Application Auto Scaling scheduled actions to increase provisioned concurrency on the Lambda alias before campaign windows and reduce it to a minimal baseline afterward.
It uses Application Auto Scaling scheduled actions to dynamically adjust provisioned concurrency, scaling up to 50 before marketing campaigns and reducing to a minimal baseline (e.g., 1-5) during low-traffic weekdays. This eliminates idle capacity costs while ensuring fast initial responses during spikes, as provisioned concurrency keeps Lambda environments warm and ready to handle requests without cold starts.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase provisioned concurrency to 100 so that cold starts never occur, regardless of traffic patterns.
Why it's wrong here
Setting provisioned concurrency to a fixed 100 means you are billed for 100 initialized environments around the clock, even when no requests arrive during low-traffic periods. This creates significant idle cost without solving the actual problem, because any spike beyond 100 concurrent invocations will still trigger cold starts for the extra executions. Moreover, a static value ignores the predictable timing of campaign windows, so you pay for capacity even when it is not needed instead of aligning cost with demand.
- ✓
Use Application Auto Scaling scheduled actions to increase provisioned concurrency on the Lambda alias before campaign windows and reduce it to a minimal baseline afterward.
Why this is correct
Application Auto Scaling scheduled actions let you define time-based policies on a Lambda alias, so provisioned concurrency is raised to a high value just before the campaign starts and lowered to a minimal baseline once it ends. During the campaign, requests are served by pre-initialized execution environments, eliminating the cold-start latency that would otherwise be noticeable. Because provisioned concurrency is billed while allocated even when idle, the schedule ensures you only pay for warm capacity during the actual spike window, not during all other hours.
- ✗
Turn provisioned concurrency off permanently and rely on retries at the client side to mask cold starts.
Why it's wrong here
Disabling provisioned concurrency reintroduces cold starts for every new execution environment, and client-side retries do not remove the initialization delay—they only send the request again after a timeout, which degrades user experience and may cause duplicate processing. During a spike, the Lambda service has to initialize many new environments at once, and you can hit regional burst concurrency limits, leading to throttling errors that retries will only amplify by adding more load. This approach trades latency guarantees for cost savings, which is the opposite of what a latency-sensitive API requires.
- ✗
Replace Lambda with a single always-on EC2 instance sized for peak demand to eliminate cold starts.
Why it's wrong here
Replacing Lambda with a single always-on EC2 instance sized for peak demand means you pay for compute capacity continuously, even when traffic is near zero, which is usually more expensive than paying only for Lambda invocations. That instance is also a single point of failure: if it fails or needs patching, the API goes down completely, and it cannot handle traffic beyond its fixed size without you manually intervening. In contrast, scheduling provisioned concurrency on Lambda gives you the same low-latency behavior during spikes while preserving the serverless benefits of automatic scalability and no idle cost.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
One of 935 original SAA-C03 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This SAA-C03 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SAA-C03 exam.