A company runs a production Amazon RDS for PostgreSQL Multi-AZ DB instance (db.r5.large) with 500 GB of General Purpose SSD (gp2) storage. The application experiences intermittent latency spikes every 15 minutes. Monitoring shows that during these spikes, the ReadIOPS metric on the primary instance spikes to 5,000 IOPS (the baseline is 1,500 IOPS), and the BurstBalance drops from 100% to 20% then recovers. There is no increase in CPU or connections. The application uses connection pooling with pgBouncer on an EC2 instance. The team has verified that no long-running queries or index scans are causing the spikes. Which action is MOST likely to resolve the intermittent latency?
gp3 provides consistent baseline IOPS without burst credits, eliminating the performance variability due to credit exhaustion.
Why this answer
The latency spikes are caused by gp2 storage burst credit exhaustion. The 500 GB gp2 volume has a baseline of 1,500 IOPS, but the workload spikes to 5,000 IOPS every 15 minutes, rapidly consuming burst credits. Migrating to gp3 provides a baseline of 3,000 IOPS and 125 MB/s throughput without relying on burst credits, thus eliminating the credit exhaustion issue.
Option A (read replica) does not resolve the primary instance's write IOPS spikes. Option B (larger instance) does not address storage IOPS limitations; CPU and connections are already normal. Option D (scale storage to 1,000 GB) would increase the gp2 baseline to 3,000 IOPS and provide more burst credits, but gp3 offers a simpler, more cost-effective solution with consistent performance and no credit-based throttling.
Exam trap
Candidates often confuse gp2 burst credits with gp3's fixed performance. They may think increasing volume size alone will eliminate bursts, but gp2 always uses credits for spikes above baseline. Migrating to gp3 removes the burst mechanism entirely.