DOP-C02 Incident and Event Response Practice Question
A DevOps team is troubleshooting an application that occasionally throws 'Connection reset by peer' errors when connecting to an RDS MySQL instance. The errors are intermittent and seem to correlate with high traffic. Which TWO steps should the team take to diagnose the issue?
⚠ Common exam trap
DOP-C02 often tests the tendency to jump to configuration changes (like increasing max_connections) instead of first diagnosing via logs and performance tools.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Check the RDS error log for messages about connection timeouts or aborted connections.
Option A is correct because the RDS error log records server-side events such as 'Aborted connection' and connection timeout messages that directly explain why the server is resetting client connections under load. Option D is correct because Performance Insights captures database load (DBLoad) broken down by wait events and top SQL, letting the team pinpoint the bottleneck causing intermittent resets during high traffic. Option B is not a diagnostic step and blindly raising max_connections may not address the root cause. Option C is a high-availability configuration change, not a troubleshooting action, and Multi-AZ does not resolve connection resets. Option E is unlikely to be the cause since security group misconfiguration would produce consistent connection failures, not intermittent resets correlated with traffic.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Check the RDS error log for messages about connection timeouts or aborted connections.
Why this is correct
The RDS error log is the first place to look because it records 'Aborted connection' entries with explicit reasons such as 'Got timeout reading communication packets' or 'Client has exceeded the max_user_connections' when a reset occurs. These entries include the source host, user, and exact error code, which lets you determine whether the reset originates from the database server (e.g., idle timeout) or from the client side. This evidence directly confirms or rules out connection-level failures before you change any configuration or architecture.
- ✗
Increase the max_connections parameter in the DB parameter group.
Why it's wrong here
Raising `max_connections` only helps if the reset is the result of the database rejecting new connections when the limit is reached—but the symptom here is *intermittent resets* of existing sessions, not a steady refusal of new ones. If the actual cause is CPU saturation, lock contention, or a network path that drops packets, adding more allowed connections can make things worse by increasing concurrency and memory pressure. Without checking `DBConnections` and `DatabaseConnections` CloudWatch metrics against the limit, this change is speculative and treats a symptom as a root cause.
- ✗
Enable Multi-AZ deployment to provide a standby instance.
Why it's wrong here
Multi-AZ provides a separate standby instance solely for automatic failover during an Availability Zone outage or database maintenance, but during normal operation your application still connects to a single primary instance—so it has zero effect on connection resets caused by load, timeouts, or resource contention. In fact, failover events themselves force all existing connections to drop and require the application to reconnect with retry logic, which could *introduce* the exact kind of intermittent reset you are trying to fix. Implementing Multi-AZ is a durability/availability feature, not a connection-resilience feature.
- ✓
Enable Performance Insights to analyze database load and find bottlenecks.
Why this is correct
Performance Insights gives you a time-series view of database load correlated across waits, SQL statements, and hosts, so you can see whether the connection resets coincide with a spike in a particular wait event like `innodb/log_write` or `CPU` or with a specific saturated query. This transforms an intermittent symptom into a measurable database load pattern, and you can drill down to the exact query or session causing the contention. Unlike static parameter changes, it provides direct evidence of the bottleneck that would force connections to be aborted.
- ✗
Review the security group rules to ensure the application can connect.
Why it's wrong here
A security group is an allow-list attached to the RDS instance, and if it lacks a rule the connection would be *consistently* rejected from the start rather than failing intermittently after successful establishment. Because the application is working most of the time, the SG is demonstrably allowing traffic; the resets are more likely due to TCP-level keepalive timeouts, DB-side `wait_timeout`, or a proxy in between dropping idle sessions. Re-reviewing the rules would not provide any new data about when or why a connection is reset, so this step is a distraction from actionable diagnostics.
Visual reference
Go deeper
Related to this question
About these practice questions
Courseiva writes every DOP-C02 question from scratch — 1,298 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This DOP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DOP-C02 exam.