Courseiva

SOA-C02 Cost and Performance Optimization Practice Question

A company runs a critical application on EC2 instances in an Auto Scaling group behind an Application Load Balancer. The application requires very low latency and high availability. The SysOps administrator notices that the application experiences increased latency during traffic spikes even though the Auto Scaling group is scaling out. Which solution would MOST effectively reduce latency?

⚠ Common exam trap

A common mix-up: candidates assume scaling out (adding more instances) always reduces latency, but the question explicitly states that scaling is already occurring yet latency persists, indicating a per-instance performance bottleneck that only a larger instance type can resolve.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Change the launch configuration to use a larger instance type with more CPU and memory.

Using a larger instance type with more CPU and memory directly addresses the root cause of increased latency during traffic spikes: the existing instances are resource-constrained under load. By provisioning instances with higher compute capacity, each instance can handle more requests per second, reducing queueing delays and per-request processing time. This is a more immediate and effective solution than scaling out alone, which adds instances but does not improve the performance of each individual instance.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Change the launch configuration to use a larger instance type with more CPU and memory.

    Why this is correct

    Changing the launch configuration to a larger instance type (e.g., from t3.medium to t3.large) vertically scales each EC2 instance, providing more vCPUs and memory to process requests concurrently. This directly reduces per-request CPU queueing and memory pressure during traffic spikes, lowering latency. However, the launch configuration only applies to newly launched instances, so an instance refresh or replacement may be needed to realize the benefit on existing instances.

  • ✗

    Pre-warm the load balancer by contacting AWS Support.

    Why it's wrong here

    Contacting AWS Support to 'pre-warm' the load balancer is a legacy practice for Classic Load Balancers, not Application Load Balancers, which automatically scale their capacity based on traffic. ALB nodes scale without customer intervention, and AWS Support will not manually pre-warm them. Even if the load balancer could be pre-warmed, the latency bottleneck is typically backend EC2 CPU/memory saturation, not the ALB's connection handling, so this option does not fix the stated issue.

  • ✗

    Increase the Auto Scaling cooldown period.

    Why it's wrong here

    Increasing the Auto Scaling cooldown period sets a longer wait time after a scaling activity before another scaling action can begin. This delays the launch of additional instances when the traffic spike persists, because the Auto Scaling group cannot react quickly to continued high demand. A longer cooldown also does nothing to improve per-instance performance, so it would actually worsen latency rather than reduce it.

  • ✗

    Deploy the instances in multiple Availability Zones.

    Why it's wrong here

    Deploying instances in multiple Availability Zones enhances fault tolerance by spreading workloads across isolated infrastructure, but it does not increase the aggregate compute capacity or reduce the latency of existing instances. If all AZs run the same instance type and the total instance count remains unchanged, the CPU/memory headroom per request is identical, so user-perceived latency during a spike is unaffected. Multi-AZ is an availability control, not a performance control.

About these practice questions

This SOA-C02 question is part of Courseiva's 1,169-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

1 more way this is tested on SOA-C02

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. An application runs on EC2 instances behind an ALB. Users report slow response times. CPU utilization averages 90% during peak hours. What is the MOST effective way to improve performance?

hard
  • A.Enable detailed monitoring on CloudWatch.
  • B.Switch to a memory-optimized instance type.
  • C.Add more Security Group rules.
  • ✓ D.Increase the instance size to a larger type.

Why D: Increasing the instance size (scaling up) provides more CPU resources directly, addressing the high utilization. While Auto Scaling could add more instances, the question specifically asks about the given options, and D is the most effective among them. Detailed monitoring (A) only provides metrics, not performance improvement. Memory-optimized (B) does not help CPU-bound issues. Security group rules (C) do not affect compute performance.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This SOA-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SOA-C02 exam.