Courseiva

DOP-C02 Resilient Cloud Solutions Practice Question

A company has a microservices architecture running on Amazon ECS with Fargate launch type. Each service is deployed in multiple Availability Zones. The services communicate via REST APIs. Recently, a downstream service experienced a partial outage, causing upstream services to time out and leading to cascading failures. The team wants to improve resilience against such failures. Which combination of actions should the DevOps engineer take? (Choose TWO.)

⚠ Common exam trap

DOP-C02 often tests whether candidates confuse 'make the timeout longer' with resilience — longer timeouts amplify cascading failures rather than preventing them.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Implement circuit breaker patterns in the service clients.

Option B is correct because a circuit breaker in the service clients detects repeated failures from a downstream dependency and trips open, failing fast instead of letting upstream threads block on slow REST calls, which directly prevents the timeout propagation and cascading failures described. Option D is correct because moving to asynchronous communication with Amazon SQS or Amazon EventBridge decouples services: upstream services enqueue or publish events and return immediately, so a partial outage in a downstream consumer no longer blocks callers, and messages can be retried or buffered until the consumer recovers. Option A is not appropriate because increasing HTTP timeouts makes callers wait longer, consuming threads and connections and worsening cascading failures rather than containing them. Option C is wrong because removing retry logic eliminates a useful resilience mechanism for transient errors; retries should be bounded and combined with circuit breakers and backoff, not deleted. Option E is not the right fix because scaling on request count does not address a downstream dependency that is failing or slow, and could even amplify load against the impaired service.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the HTTP timeout values for all service-to-service calls.

    Why it's wrong here

    Longer timeouts keep upstream threads blocked on a degraded downstream, deepening the cascade rather than containing it. Timeouts are tuned alongside circuit breakers, but raising them alone is the right move only when a downstream is merely slow, not partially failing.

  • ✓

    Implement circuit breaker patterns in the service clients.

    Why this is correct

    Circuit breakers in service clients detect repeated downstream failures and trip open, failing fast instead of holding threads until timeout. This halts the cascade at the caller, satisfying the resilience requirement by preventing upstream exhaustion when a downstream service partially fails.

  • ✗

    Remove all retry logic from service calls.

    Why it's wrong here

    Removing retry logic strips the mechanism that lets a caller survive a transient downstream failure, so a partial outage propagates immediately as hard errors. Retries with exponential backoff and jitter are precisely the remedy here. Retry removal suits scenarios where calls are non-idempotent and duplicates cause corruption, not resilience hardening.

  • ✓

    Adopt an asynchronous communication pattern using Amazon SQS or Amazon EventBridge.

    Why this is correct

    Decoupling upstream services from the failing downstream via SQS or EventBridge removes the synchronous REST dependency, so timeouts cannot propagate. Queued or event-driven delivery lets the downstream recover at its own pace, directly satisfying the resilience requirement against partial outages and cascading failures.

  • ✗

    Configure Auto Scaling for all services based on request count.

    Why it's wrong here

    Auto Scaling handles load but does not prevent cascading failures due to downstream unavailability.

Visual reference

Client Server SYN (seq=100) SYN-ACK (seq=200, ack=101) ACK (ack=201) Connection established — data transfer begins

About these practice questions

One of 1,298 original DOP-C02 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This DOP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DOP-C02 exam.