Courseiva

DOP-C02 Resilient Cloud Solutions Practice Question

A company runs a critical application on EC2 instances behind an Application Load Balancer. The application uses an Amazon RDS for PostgreSQL Multi-AZ DB instance. During a recent failover test, the application experienced a 5-minute downtime. The RDS failover completed within 30 seconds. What is the most likely cause of the prolonged downtime?

⚠ Common exam trap

Candidates often assume the 5-minute downtime must be caused by the database failover itself, but the question explicitly states the failover completed in 30 seconds, so the real issue is application-side DNS caching or stale connection handling.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The application caches DNS resolutions, causing it to connect to the old writer endpoint

The most likely cause is that the application caches DNS resolutions, causing it to continue connecting to the old writer endpoint after failover. When an RDS Multi-AZ failover occurs, the DNS record for the writer endpoint is updated to point to the new primary instance, but the application's cached DNS entry still points to the old IP address. Since the old primary is now a standby and no longer accepts connections, the application experiences downtime until the DNS cache expires (typically 5–60 seconds) or the application refreshes the DNS resolution. The 5-minute downtime suggests the application uses a long DNS TTL or a custom caching layer that delays reconnection.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    The application caches DNS resolutions, causing it to connect to the old writer endpoint

    Why this is correct

    After a Multi-AZ failover, the RDS DNS record for the writer endpoint is updated to point to the new primary instance, but an application that caches DNS resolutions continues sending write connections to the old IP. That old IP belongs to the former primary, which is now promoted to standby (or replaced) and will not accept write connections, so all new database requests fail until the TTL expires or the application is restarted. This directly causes the five-minute outage because the application never re-resolves the endpoint until the cached entry times out.

  • ✗

    The RDS Multi-AZ failover took longer than expected due to a large transaction log

    Why it's wrong here

    Multi-AZ failover is designed to complete in about 60-120 seconds because the standby continuously receives synchronous replication; it does not need to recover a large transaction log before promoting itself. While a very large transaction can briefly delay certain aspects of failover, Amazon RDS handles this by rolling back incomplete transactions and promoting an already synchronized standby, so log size is not a likely cause of a five-minute outage. The failure is more plausibly in the client's reconnection behavior than in the database's internal failover duration.

  • ✗

    The Application Load Balancer health checks marked all instances as unhealthy during the failover

    Why it's wrong here

    The ALB health checks evaluate the health of EC2 targets, not the RDS database; during a database failover the instances themselves remain running and the ALB would keep routing traffic to them unless the application health check endpoint explicitly depends on DB connectivity. Even if the health checks did mark instances unhealthy, that would return HTTP 503 earlier and not explain why the failover itself took five minutes - it would be a side effect of the application losing DB access, while the actual failure was the app's inability to re-establish a database connection. So this cannot be the root cause of the outage.

  • ✗

    The application was using read replicas for writes, which failed during failover

    Why it's wrong here

    RDS read replicas are separate read-only endpoints and do not participate in Multi-AZ failover; they are never a valid target for write operations. An application using read replicas for writes would immediately receive a read-only error from the database engine, not experience an outage specifically during failover. The Multi-AZ standby is the failover target, and writes are only accepted on the primary endpoint, so this option misidentifies both the purpose of read replicas and the failover architecture.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

About these practice questions

One of 1,298 original DOP-C02 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DOP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DOP-C02 exam.