Courseiva

SAA-C03 Design High-Performing Architectures Practice Question

Exhibit

Lambda monitoring and deployment notes:
- Function: checkout-api-prod
- Current alias: live
- Invocations per day: low except weekdays 09:00-09:15 UTC
- REPORT log sample at 09:00 UTC:
  Init Duration: 842.31 ms
  Duration: 128.42 ms
  Billed Duration: 1000 ms
- REPORT log sample at 09:05 UTC:
  Init Duration: 0.00 ms
  Duration: 121.77 ms

Traffic pattern:
- Spikes are predictable and last about 15 minutes
- No need to keep high concurrency all day

Based on the exhibit, a serverless API on AWS Lambda experiences a predictable cold-start penalty every weekday at 09:00 UTC when a marketing campaign begins. The team wants the first requests to stay fast while minimizing extra cost during quiet periods. What is the best approach?

⚠ Common exam trap

A common mix-up: candidates confuse increasing memory or timeout with solving cold starts, or they mistakenly think an ALB can pre-warm Lambda, when in fact only provisioned concurrency guarantees warm containers for the first requests in a predictable traffic spike.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable provisioned concurrency on the published version and schedule it to scale up shortly before the spike.

Provisioned concurrency pre-warms a specified number of Lambda execution environments so that incoming requests do not incur a cold start. By scheduling the provisioned concurrency to scale up just before the 09:00 UTC spike and scale down afterward, the team eliminates the cold-start penalty during the campaign while minimizing cost during quiet periods. This directly addresses the predictable, time-bound traffic pattern without requiring code changes or over-provisioning.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Enable provisioned concurrency on the published version and schedule it to scale up shortly before the spike.

    Why this is correct

    Provisioned concurrency pre-initializes Lambda execution environments for a specific version or alias, eliminating the cold-start delay for the first request. By configuring a scheduled scaling action (via Application Auto Scaling) to raise provisioned concurrency just before the spike, the environments are already warm and ready to handle the burst. This approach directly targets the initialization penalty shown in the exhibit, and because provisioned concurrency is billed only while it is enabled, scheduling it around the known spike avoids paying for idle resources the rest of the day.

  • ✗

    Increase the Lambda timeout so cold starts have more time to complete.

    Why it's wrong here

    Raising the Lambda timeout only extends the maximum time a function can run before being terminated; it does nothing to reduce the time required to initialize a new execution environment. A cold start comprises downloading the code, starting the runtime, and running initialization code—all of which happen before the handler executes, regardless of the configured timeout. In fact, a longer timeout could worsen user-perceived latency by allowing slow executions to continue, but it cannot prevent the initial invocation from suffering a multi-second init duration.

  • ✗

    Move the function behind an Application Load Balancer to improve warm-up behavior.

    Why it's wrong here

    Placing the function behind an Application Load Balancer merely changes how HTTP requests are routed to Lambda; the ALB itself does not perform any warm-up or pre-initialization of Lambda execution environments. Cold starts originate from the Lambda service when it must create a new sandbox and load the function package—a process that is independent of the front-end service. ALBs do offer connection keep-alive, but that only affects network-level connections, not the lifecycle of the underlying compute environment, so first requests would still experience the exact same initialization latency.

  • ✗

    Increase the function memory to the maximum value and leave concurrency unchanged.

    Why it's wrong here

    Increasing the function memory allocates proportionally more CPU and can make the execution phase faster, but it does not eliminate the initial environment setup time. The cold-start penalty is driven by initialization steps such as loading the code and starting the runtime, which must occur for every new execution environment no matter how much memory is configured. Even at the maximum memory setting, a fresh container still incurs an init duration before your handler starts, so without pre-warming or provisioned concurrency the spike would still see slow first responses.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

About these practice questions

This SAA-C03 question is part of Courseiva's 935-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This SAA-C03 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SAA-C03 exam.