A web application uses pooled JDBC connections to an Amazon Aurora cluster using the writer endpoint. During an Aurora planned failover, monitoring shows a short spike in failed requests. The Aurora cluster writer endpoint remains the same, but many existing pooled connections briefly fail. The application retries aggressively and overloads the new writer during the transition.
Which design change will most improve application resilience during Aurora failovers without requiring application redeployment?
Trap 1: Change the Aurora cluster to Single-AZ to reduce failover events.
Changing the Aurora cluster from Multi-AZ to Single-AZ does not reduce the frequency of failover events; in fact, it increases the blast radius when an AZ failure or instance health issue occurs. Aurora storage is already replicated across three AZs, but a Single-AZ configuration means the primary DB instance has no cross-AZ standby or replacement target to promote during a failover, so availability is degraded rather than improved. This option trades away the very redundancy that allows Aurora to recover quickly, making connection disruption worse rather than smoothing it over.
Trap 2: Increase the application thread count so more requests can be…
Raising the application thread count during a connection disruption amplifies connection churn because every new thread that needs a JDBC connection will attempt to obtain one from a pool already filled with stale, broken connections from the old writer. This creates a thundering herd of retry requests that can overwhelm the newly promoted writer instance, increasing CPU and memory pressure and slowing transaction recovery. Pooled connections are not refreshed by adding threads; the application still needs a mechanism like RDS Proxy to transparently replace backend connections and reduce the stall and retry storm.
Trap 3: Pin all database traffic to a specific instance hostname instead of…
Pinning database traffic to a specific instance hostname, such as the old writer's instance endpoint, defeats the purpose of the Aurora cluster endpoint because the cluster endpoint uses DNS to automatically redirect applications to the new writer after a failover. The instance endpoint remains associated with the same physical node, which becomes read-only or unavailable after a role switch, so writes fail with read-only or connection errors. This approach directly undermines Aurora's failover design, making an already disruptive failover into an application outage.
- A
Add an RDS Proxy between the application and Aurora to manage database connections across failovers.
RDS Proxy terminates and manages client connections, while maintaining separate managed connections to the database. During a writer failover, the proxy can re-establish backend connections to the new writer, reducing failed pooled connections seen by the application and lowering retry pressure.
- B
Change the Aurora cluster to Single-AZ to reduce failover events.
Why it fails: Changing the Aurora cluster from Multi-AZ to Single-AZ does not reduce the frequency of failover events; in fact, it increases the blast radius when an AZ failure or instance health issue occurs. Aurora storage is already replicated across three AZs, but a Single-AZ configuration means the primary DB instance has no cross-AZ standby or replacement target to promote during a failover, so availability is degraded rather than improved. This option trades away the very redundancy that allows Aurora to recover quickly, making connection disruption worse rather than smoothing it over.
- C
Increase the application thread count so more requests can be served while connections reconnect.
Why it fails: Raising the application thread count during a connection disruption amplifies connection churn because every new thread that needs a JDBC connection will attempt to obtain one from a pool already filled with stale, broken connections from the old writer. This creates a thundering herd of retry requests that can overwhelm the newly promoted writer instance, increasing CPU and memory pressure and slowing transaction recovery. Pooled connections are not refreshed by adding threads; the application still needs a mechanism like RDS Proxy to transparently replace backend connections and reduce the stall and retry storm.
- D
Pin all database traffic to a specific instance hostname instead of the writer cluster endpoint.
Why it fails: Pinning database traffic to a specific instance hostname, such as the old writer's instance endpoint, defeats the purpose of the Aurora cluster endpoint because the cluster endpoint uses DNS to automatically redirect applications to the new writer after a failover. The instance endpoint remains associated with the same physical node, which becomes read-only or unavailable after a role switch, so writes fail with read-only or connection errors. This approach directly undermines Aurora's failover design, making an already disruptive failover into an application outage.