A company is deploying a web application that uses an Application Load Balancer and an Auto Scaling group of EC2 instances. The application must be able to handle sudden spikes in traffic. Which TWO actions should the Solutions Architect take to improve scalability and reduce latency? (Choose two.)
HTTP/2 multiplexes many requests over a single TCP connection, cutting head-of-line blocking and connection overhead during sudden traffic spikes. This directly reduces latency for the web application behind the Application Load Balancer, satisfying the stem's requirement to handle bursts efficiently without adding capacity.
Why this answer
Option A is correct because enabling HTTP/2 on the Application Load Balancer allows multiplexed, concurrent requests over a single TCP connection and header compression, which reduces latency and improves throughput during traffic spikes. Option D is correct because a predictive scaling policy in the Auto Scaling group uses machine learning to forecast demand and pre-provision capacity ahead of anticipated spikes, improving responsiveness and reducing latency compared to reactive scaling alone. Option B is incorrect because increasing the default cooldown period delays subsequent scaling actions, making the group slower to react to sudden traffic increases.
Option C is incorrect because simply using larger instance types does not improve elasticity or latency during spikes and can reduce the granularity of scaling. Option E is incorrect because increasing the health check interval slows detection of unhealthy targets and does not improve scalability or latency.