An application runs on an EC2 Auto Scaling group. Over the last month, CPU utilization averaged 8% with no sustained memory pressure, and response times are stable. The team wants to lower monthly cost without changing the application. What is the most appropriate next step for cost optimization?
Trap 1: Increase desired capacity to 2x so utilization increases and…
Increasing desired capacity to 2x simply launches additional instances, which multiplies compute costs for a workload that is already underutilizing its current instances. The hypothesis that higher utilization means 'more efficient' confuses efficiency with artificially inflating activity; spreading the same load across twice the instances would actually reduce per-instance CPU utilization further, not improve it. This action increases total infrastructure cost and fails to address the underlying overprovisioning, making it the opposite of the optimal cost-optimization move.
Trap 2: Disable Auto Scaling so the group never scales down to preserve…
Disabling Auto Scaling would remove the group's ability to terminate unnecessary instances during low-demand periods, locking in wasted spend on idle capacity. While a static size might preserve baseline performance, it sacrifices the elasticity that keeps infrastructure aligned with actual demand, and the scenario indicates consistent low utilization—so a smaller, right-sized instance type (not disabling scaling) is the better remedy. Moreover, disabling scaling also reduces fault tolerance because the group can no longer self-heal or respond to spikes without manual intervention.
Trap 3: Switch the workload to Spot instances immediately to avoid…
Spot Instances can offer significant discounts, but they come with a two-minute interruption warning when EC2 reclaims capacity, making them unsuitable for workloads without fault tolerance or checkpointing. The scenario provides no evidence that this application is interruption-tolerant or stateless, so a blanket switch immediately exposes it to availability risks. Right-sizing the On-Demand fleet based on observed utilization is a safer, more appropriate first cost-optimization step; Spot should only be considered after confirming the workload can handle interruptions.
- A
Evaluate a smaller EC2 instance type (via the Auto Scaling launch template/configuration) for the group and validate performance metrics after the change.
Right-sizing involves analyzing historical utilization metrics (e.g., CloudWatch CPU, memory) to select a less expensive instance type while still satisfying workload requirements. Changing the Auto Scaling group's launch template to a smaller instance type, then monitoring metrics such as CPU credit balance, latency, and throughput, confirms the new size can handle peak demand. This directly lowers per-instance cost without altering the number of instances needed, providing cost savings while preserving availability and performance characteristics.
- B
Increase desired capacity to 2x so utilization increases and instances become “more efficient.”
Why it fails: Increasing desired capacity to 2x simply launches additional instances, which multiplies compute costs for a workload that is already underutilizing its current instances. The hypothesis that higher utilization means 'more efficient' confuses efficiency with artificially inflating activity; spreading the same load across twice the instances would actually reduce per-instance CPU utilization further, not improve it. This action increases total infrastructure cost and fails to address the underlying overprovisioning, making it the opposite of the optimal cost-optimization move.
- C
Disable Auto Scaling so the group never scales down to preserve baseline performance.
Why it fails: Disabling Auto Scaling would remove the group's ability to terminate unnecessary instances during low-demand periods, locking in wasted spend on idle capacity. While a static size might preserve baseline performance, it sacrifices the elasticity that keeps infrastructure aligned with actual demand, and the scenario indicates consistent low utilization—so a smaller, right-sized instance type (not disabling scaling) is the better remedy. Moreover, disabling scaling also reduces fault tolerance because the group can no longer self-heal or respond to spikes without manual intervention.
- D
Switch the workload to Spot instances immediately to avoid On-Demand charges, regardless of interruption risk.
Why it fails: Spot Instances can offer significant discounts, but they come with a two-minute interruption warning when EC2 reclaims capacity, making them unsuitable for workloads without fault tolerance or checkpointing. The scenario provides no evidence that this application is interruption-tolerant or stateless, so a blanket switch immediately exposes it to availability risks. Right-sizing the On-Demand fleet based on observed utilization is a safer, more appropriate first cost-optimization step; Spot should only be considered after confirming the workload can handle interruptions.