An API backed by Lambda returns high p95 latency after deployment. Which two telemetry sources are most useful first?
CloudWatch provides critical metrics like `Duration` and `Init Duration` for Lambda functions, directly revealing execution and cold start times. Analyzing the p95 percentile of these metrics pinpoints specific latency bottlenecks. Furthermore, detailed CloudWatch Logs offer granular insights into the function's internal execution flow, external service calls, and potential code-level inefficiencies contributing to high latency.
Why this answer
CloudWatch Lambda duration and init duration metrics directly measure the time your function spends executing and initializing, which are the primary drivers of p95 latency. Logs can reveal cold starts, timeouts, or inefficient code paths that cause high latency. These are the most immediate telemetry sources to identify performance bottlenecks in the Lambda function itself.
Exam trap
The trap here is that candidates often overlook the combination of CloudWatch metrics and X-Ray traces, mistakenly thinking that only one telemetry source (like CloudWatch logs) is sufficient, or they confuse billing data with performance monitoring.