A company runs a microservices application on Amazon ECS with Fargate. The operations team notices that some services are experiencing intermittent high latency, but CPU and memory metrics appear normal. They need to identify the root cause. Which approach should they use?
Instrumenting with the AWS X-Ray SDK is the correct approach because X-Ray traces individual requests as they traverse services, generating a trace ID that propagates across HTTP headers and AWS SDK client calls. The X-Ray console provides a service map and trace timelines, segment and subsegment views that break down latency for each downstream call, letting you pinpoint the exact service or resource causing the slowdown. You can also annotate traces with request metadata and set sampling rules to balance overhead and observability.
Why this answer
Intermittent latency with normal CPU and memory metrics points to a distributed tracing problem — the bottleneck is likely in a downstream call, a network hop, or a specific service in the request chain. AWS X-Ray traces requests end-to-end across ECS tasks, Lambda functions, and downstream services, showing exactly where time is spent. This makes it the right tool to pinpoint the root cause of intermittent latency in a microservices architecture.
Exam trap
DOP-C02 often tests the distinction between metrics, logs, and traces — candidates pick CloudWatch Logs or Prometheus because they sound comprehensive, but only X-Ray provides the per-request, cross-service causality needed to diagnose intermittent latency.
How to eliminate wrong answers
Option A is wrong because CloudWatch Logs Insights can query logs for slow requests, but it requires the application to already log timing data and does not automatically correlate latency across service boundaries — it is reactive and manual rather than a tracing solution. Option B is wrong because Amazon Managed Service for Prometheus collects metrics, and metrics alone cannot explain why a specific request was slow — they show aggregates, not per-request causality. Option C is wrong because CloudWatch Synthetics canaries measure endpoint response times from the outside but cannot tell you which internal service or call caused the latency, so they detect the symptom rather than diagnose the cause.