RDS Connectivity Issues: Subnet Routing in Multi-AZ Deployments
A company has deployed a multi-tier web application in a single AWS region. The architecture includes a VPC with public and private subnets across two Availability Zones. The web tier uses an Application Load Balancer (ALB) in the public subnets, and the application tier runs on EC2 instances in the private subnets. The database tier uses an Amazon RDS Multi-AZ deployment in the database subnets. The company is experiencing intermittent connectivity issues between the application tier and the database tier. The application logs show connection timeouts. The network engineer has verified that the security groups and network ACLs are correctly configured. The RDS instance is reachable from the application tier via a telnet test from one specific instance, but not consistently from all instances. What is the most likely cause of the intermittent connectivity?
Quick Answer
The answer is missing route table entries for the database subnet CIDRs in the application subnets' route tables. This is the most likely cause because when an RDS Multi-AZ deployment spans subnets in different Availability Zones, the DNS endpoint can resolve to a primary instance in a database subnet that resides in a different AZ than the application tier. If the application subnet's route table lacks an explicit route to the database subnet's CIDR block, traffic to that IP address will be dropped, causing intermittent timeouts—even though security groups and network ACLs are correctly configured. On the AWS Certified Advanced Networking Specialty ANS-C01 exam, this scenario tests your understanding of how RDS DNS resolution interacts with subnet routing in a multi-AZ architecture, a common trap where engineers focus only on security groups and forget that subnets in different AZs are separate networks requiring explicit routes. A key memory tip: RDS endpoints are DNS-based, not IP-based; always verify that every application subnet has a route to every database subnet CIDR in the RDS subnet group.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The RDS Multi-AZ failover is causing the primary instance to change, and the application is not reconnecting to the new endpoint.
The most likely cause is that the RDS Multi-AZ failover is causing the primary instance to change, and the application is not reconnecting to the new endpoint. In a Multi-AZ deployment, if a failover occurs, the DNS endpoint updates to point to the new primary, but if the application caches the resolved IP address or uses stale connections, it may experience intermittent timeouts. This explains why connectivity works from some instances (those with fresh DNS entries) but not others. Option C is incorrect because within a single VPC, the local route automatically enables routing between all subnets; missing routes would cause consistent failure, not intermittent. Option B is incorrect because network ACLs, if correctly configured, would not cause intermittent issues. Option D is incorrect because a security group misconfiguration would likely cause consistent failure from all instances, not intermittent connectivity.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
The RDS Multi-AZ failover is causing the primary instance to change, and the application is not reconnecting to the new endpoint.
Why this is correct
Correct. RDS Multi-AZ failover can cause the primary endpoint IP to change. If the application does not properly handle DNS changes or uses stale connections, it may experience intermittent connectivity after a failover.
- ✗
The network ACLs on the database subnets are blocking ephemeral ports used by the application.
Why it's wrong here
Incorrect. Network ACLs are stateless but if correctly configured, they would not block ephemeral ports intermittently. The issue is not consistent, making this unlikely.
- ✗
The database subnets are in different Availability Zones than the application subnets, and the route tables in the application subnets do not have routes to the database subnet CIDRs.
Why it's wrong here
Incorrect. In a single VPC, the local route automatically enables communication between all subnets. Missing routes to database subnet CIDRs is not possible within the same VPC.
- ✗
The security group for the database is allowing traffic only from the application tier's security group, but the application tier instances are using a different security group.
Why it's wrong here
Incorrect. If the security group for the database allowed traffic only from a different security group than the one used by the application instances, connectivity would consistently fail from all instances, not intermittently.
Visual reference
Go deeper
Related to this question
About these practice questions
Courseiva writes every ANS-C01 question from scratch — 1,621 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
Same concept, more angles
1 more way this is tested on ANS-C01
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. A company is designing a network architecture for a two-tier web application. The web tier runs on EC2 instances behind an Application Load Balancer (ALB) in public subnets. The application tier runs on EC2 instances in private subnets. The application tier needs to access an Amazon RDS for PostgreSQL database in the same private subnets. The company requires that all traffic between the ALB and web tier, as well as between web tier and application tier, remain within the AWS network and not traverse the internet. The current design uses an Internet Gateway (IGW) for public subnet internet access and a NAT Gateway for private subnet outbound internet access. The web tier instances have a default route to the IGW, and the application tier instances have a default route to the NAT Gateway. The security groups are configured correctly. However, the application tier cannot connect to the RDS database. What is the MOST likely cause?
medium- ✓ A.The application tier instances are using the RDS public DNS name instead of the private DNS name
- B.The RDS database is in a different VPC
- C.The ALB is not configured to forward traffic to the web tier
- D.The NAT Gateway is not configured with the correct route to the RDS subnet
Why A: The RDS database is in private subnets. The application tier instances are also in private subnets. They should be able to communicate within the same VPC via private IP addresses. The issue is not about internet access. The most likely cause is that the application tier instances are trying to connect to the RDS endpoint using the public DNS name, which resolves to a public IP, and the traffic is being routed to the NAT Gateway, which blocks inbound traffic from the internet (the RDS public endpoint). The application tier should use the private DNS name or the private IP address of the RDS instance. Alternatively, the security group might be misconfigured, but the question says security groups are correct. The most common mistake is using the public endpoint.
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This ANS-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the ANS-C01 exam.