Courseiva

CCNA Troubleshooting and Optimization Questions

75 of 179 questions · Page 2/3 · Troubleshooting and Optimization · Answers revealed

76
Multi-Selecthard

An API backed by Lambda returns high p95 latency after deployment. Which two telemetry sources are most useful first?

Select 2 answers
A.AWS Billing console only
B.CloudWatch Lambda duration/init duration/logs
C.S3 Inventory reports
D.X-Ray traces across API Gateway and Lambda
AnswersB, D

CloudWatch provides critical metrics like `Duration` and `Init Duration` for Lambda functions, directly revealing execution and cold start times. Analyzing the p95 percentile of these metrics pinpoints specific latency bottlenecks. Furthermore, detailed CloudWatch Logs offer granular insights into the function's internal execution flow, external service calls, and potential code-level inefficiencies contributing to high latency.

Why this answer

CloudWatch Lambda duration and init duration metrics directly measure the time your function spends executing and initializing, which are the primary drivers of p95 latency. Logs can reveal cold starts, timeouts, or inefficient code paths that cause high latency. These are the most immediate telemetry sources to identify performance bottlenecks in the Lambda function itself.

Exam trap

The trap here is that candidates often overlook the combination of CloudWatch metrics and X-Ray traces, mistakenly thinking that only one telemetry source (like CloudWatch logs) is sufficient, or they confuse billing data with performance monitoring.

77
MCQmedium

A developer notices that an Amazon RDS for MySQL DB instance's CPU utilization is consistently above 90% during peak hours. Which AWS service can the developer use to analyze the database queries and identify the root cause?

A.AWS X-Ray
B.Amazon CloudWatch Logs
C.Amazon RDS Performance Insights
D.AWS Trusted Advisor
AnswerC

Amazon RDS Performance Insights is the dedicated service for monitoring and analyzing database performance by visualizing database load. It collects and presents key performance metrics, including Average Active Sessions (AAS), wait events, and the full text of top SQL queries, allowing developers to quickly identify bottlenecks and understand *why* a database is slow. Its interactive dashboard provides granular insights into the database engine's activity, pinpointing specific queries, users, or hosts consuming the most resources and enabling targeted optimization.

Why this answer

Amazon RDS Performance Insights is the correct service for analyzing database queries and identifying the root cause of high CPU utilization. Performance Insights provides a database performance analysis feature that offers a visual dashboard and query-level metrics, enabling developers to pinpoint resource-intensive queries. AWS X-Ray is used for tracing application requests, not database queries.

CloudWatch Logs is for log data storage and monitoring, not query-level analysis. Trusted Advisor provides cost and security recommendations, not database performance insights. Therefore, Performance Insights is the appropriate tool for this scenario.

78
Multi-Selectmedium

A DynamoDB table shows throttling on one partition key value. Which two signs point to a hot partition problem?

Select 2 answers
A.Most traffic targets the same partition key
B.The table has point-in-time recovery enabled
C.Consumed capacity is uneven despite total table capacity being available
D.CloudTrail is enabled in all regions
AnswersA, C

DynamoDB distributes data across partitions based on the partition key. When a disproportionate amount of read or write traffic targets a small subset of partition key values, those specific partitions become 'hot.' Each partition has a maximum throughput limit, typically 3000 RCU and 1000 WCU. Exceeding this limit on a single partition, even if the overall table capacity is sufficient, results in throttling requests directed at that hot partition.

Why this answer

A hot partition occurs when a single partition key value receives a disproportionate share of read/write traffic, causing throttling on that partition even if the table's total provisioned capacity is not fully utilized. This imbalance means the partition's capacity is exhausted while other partitions remain underutilized, leading to request throttling for that specific key.

Exam trap

The trap here is that candidates confuse overall table capacity with partition-level capacity, assuming throttling only happens when total consumed capacity exceeds provisioned capacity, rather than recognizing that uneven key distribution can cause throttling on a single partition.

79
MCQmedium

A developer is monitoring an AWS Lambda function that is triggered by an Amazon SQS queue. The function's CloudWatch metrics show a high number of throttles. The function has a reserved concurrency of 10 and the SQS queue has a large backlog of messages. The function processes each message in about 2 seconds and has a timeout of 60 seconds. Which action will most effectively reduce the throttles and increase throughput?

A.Increase the reserved concurrency of the Lambda function to 50
B.Increase the batch size in the SQS event source mapping to 100
C.Increase the function timeout to 120 seconds
D.Decrease the reserved concurrency to 5
AnswerA

Increasing the reserved concurrency for a Lambda function dedicates a specific number of concurrent execution slots exclusively to that function. This action guarantees that the function can scale up to 50 simultaneous invocations, preventing it from being throttled by the account's general unreserved concurrency pool. For an SQS-triggered Lambda, this directly enables more parallel processing of messages, significantly improving throughput and reducing the backlog in the queue.

Why this answer

The high throttles indicate that the Lambda function's reserved concurrency of 10 is insufficient to handle the incoming messages from the SQS queue. By increasing reserved concurrency to 50, you allow more concurrent executions, which reduces throttling and increases throughput. The function's 2-second processing time and 60-second timeout are not the bottleneck; the concurrency limit is.

Exam trap

The trap here is that candidates may think increasing batch size or timeout will help, but they overlook that the root cause is the reserved concurrency cap, which directly limits the number of concurrent executions and is the primary driver of throttles.

How to eliminate wrong answers

Option B is wrong because increasing the batch size to 100 would cause the function to receive more messages per invocation, but with a reserved concurrency of 10, the function can only process 10 batches concurrently, so throttles would persist and latency could increase due to longer processing per batch. Option C is wrong because increasing the timeout to 120 seconds does not address the concurrency limit; the function already completes in 2 seconds, so a longer timeout has no effect on throttles. Option D is wrong because decreasing reserved concurrency to 5 would reduce the number of concurrent executions, worsening throttles and decreasing throughput.

80
Multi-Selecthard

Which TWO are best practices for optimizing DynamoDB performance? (Choose two.)

Select 2 answers
A.Use SQS to decouple write-heavy workloads and handle spikes.
B.Use partition keys with high cardinality to distribute traffic evenly.
C.Provision maximum write capacity units to handle any spike.
D.Use Scan operations instead of Query for retrieving data.
E.Enable strongly consistent reads for all read operations.
AnswersA, B

SQS acts as a buffer, decoupling the producer (application) from the consumer (DynamoDB). During write-heavy workloads or sudden traffic spikes, SQS queues the requests, allowing the application to continue processing without being throttled by DynamoDB's provisioned capacity. A separate worker process can then consume messages from SQS at a controlled rate, ensuring DynamoDB's write capacity units (WCUs) are not exceeded and operations are processed reliably. This prevents throttling errors and improves overall system resilience.

Why this answer

Using SQS to decouple write-heavy workloads allows DynamoDB to absorb traffic spikes by buffering writes in a queue, preventing throttling and enabling batch processing. This pattern, often called 'queue-based load leveling,' ensures that DynamoDB's provisioned capacity is not overwhelmed by sudden bursts, improving overall system resilience and cost efficiency.

Exam trap

The trap here is that candidates often confuse 'handling spikes' with over-provisioning capacity (Option C) instead of using decoupling patterns like SQS, or they mistakenly believe that Scan operations are acceptable for frequent data retrieval, ignoring the cost and performance penalties.

81
MCQeasy

A developer is troubleshooting an AWS Lambda function that is timing out. The function processes S3 events and writes to DynamoDB. The average execution time is 5 seconds, but the function times out after 3 seconds. What is the most likely cause?

A.The S3 bucket is not configured to send event notifications.
B.DynamoDB write capacity is insufficient.
C.The Lambda function timeout is set to 3 seconds.
D.The Lambda function concurrency limit is exceeded.
AnswerC

The Lambda function timeout configuration directly dictates the maximum duration an invocation is allowed to run before the Lambda service forcibly terminates it. If the function's code, including any synchronous downstream calls or complex processing, requires more than 3 seconds to complete, it will inevitably result in a timeout error. The default timeout is often 3 seconds, and increasing this value is the primary solution when a function consistently fails to finish within its allocated execution window.

Why this answer

The function's average execution time is 5 seconds, but it times out at 3 seconds — this mismatch directly indicates the Lambda function's configured timeout value is set to 3 seconds, which is lower than the actual runtime. The timeout setting is a per-function configuration (default 3 seconds, max 15 minutes) that terminates execution when exceeded. Since the workload legitimately needs ~5 seconds, the fix is to increase the timeout to a value above the observed runtime.

Exam trap

DVA-C02 often tests the confusion between a function timing out due to its own configured timeout versus downstream service throttling — candidates see 'DynamoDB' in the stem and jump to capacity, ignoring that the timeout value (3s) exactly matches the default and is below the stated 5s runtime.

How to eliminate wrong answers

Option A is wrong because missing S3 event notifications would prevent the function from being invoked at all, not cause it to time out mid-execution. Option B is wrong because insufficient DynamoDB write capacity would produce throttling errors (ProvisionedThroughputExceededException) or retries, not a hard timeout at exactly 3 seconds. Option D is wrong because exceeding the concurrency limit results in throttled invocations (429 TooManyRequestsException) or queued events, not a function-level timeout during execution.

82
Multi-Selectmedium

Which THREE are valid methods to handle application configuration in AWS? (Choose three.)

Select 3 answers
A.AWS CloudFormation template parameters
B.AWS Secrets Manager
C.AWS IAM roles
D.Lambda environment variables
E.AWS Systems Manager Parameter Store
AnswersB, D, E

AWS Secrets Manager is a dedicated service designed for securely storing, managing, and retrieving sensitive credentials and other secrets, such as database passwords, API keys, and OAuth tokens. Applications can programmatically fetch these secrets at runtime, ensuring that sensitive configuration data is never hardcoded or exposed in plain text within application code or configuration files. It also offers automatic rotation capabilities to enhance security.

Why this answer

AWS Secrets Manager is a valid method for handling application configuration because it securely stores and manages sensitive configuration data such as database credentials, API keys, and other secrets. It supports automatic rotation of secrets, fine-grained access control via IAM policies, and integrates with AWS services like RDS, Redshift, and Lambda. This makes it ideal for managing dynamic configuration values that require high security and lifecycle management.

Exam trap

The trap here is that candidates often confuse IAM roles with configuration storage, thinking that roles can hold configuration data, when in fact roles only define permissions and cannot store key-value pairs or secrets.

83
MCQhard

A developer notices that an AWS Lambda function, which uses Amazon RDS Proxy to connect to an Aurora MySQL database, is experiencing increased latency and occasional connection timeouts. The function is configured with a reserved concurrency of 100 and is deployed in a VPC. The RDS Proxy's maximum connections is set to 1000. CloudWatch metrics show that the DatabaseConnections metric for the proxy is consistently at 1000. What is the most likely cause of the increased latency and timeouts?

A.The Lambda function is not reusing database connections properly, exhausting the proxy connection pool
B.The RDS Proxy target group is not configured with the correct DB instance
C.The Lambda function's execution role is missing the rds-db:connect permission
D.The VPC does not have a NAT Gateway for outbound traffic
AnswerA

Lambda functions are inherently stateless and often short-lived. Without explicit connection pooling implemented within the Lambda function's code (e.g., by declaring the connection object in a global scope), each new invocation will attempt to establish a fresh connection to the RDS Proxy. This rapid creation of new client connections, especially under high concurrency, quickly exhausts the limited connection pool managed by the RDS Proxy, leading to connection failures and increased latency as requests wait for available connections.

Why this answer

The RDS Proxy's DatabaseConnections metric is consistently at 1000, which equals the proxy's maximum connections setting. This indicates the proxy connection pool is fully saturated. When all connections are in use, new connection requests from Lambda invocations must wait, causing increased latency, and if the wait exceeds the timeout, connection timeouts occur.

The most likely cause is that the Lambda function is not reusing database connections (e.g., not using connection pooling or keeping connections open across invocations), exhausting the pool.

Exam trap

The trap here is that candidates may focus on the reserved concurrency (100) versus proxy max connections (1000) and assume the numbers are fine, missing that the real issue is connection reuse per invocation, not the total count.

How to eliminate wrong answers

Option B is wrong because if the target group were misconfigured, the proxy would fail to connect to the database entirely, not just experience latency and timeouts while the connection pool is full. Option C is wrong because missing the rds-db:connect permission would cause immediate authentication failures (e.g., 'Access denied') for all connection attempts, not gradual pool exhaustion. Option D is wrong because Lambda functions in a VPC use Elastic Network Interfaces (ENIs) for outbound traffic to RDS Proxy within the same VPC; a NAT Gateway is only needed for internet-bound traffic, not for connecting to RDS Proxy in the same VPC.

84
MCQmedium

A company's application uses Amazon S3 to store user-uploaded images. Users report that recently uploaded images are sometimes not immediately available for viewing. The application uses S3 Event Notifications to trigger a Lambda function that processes images and stores metadata in DynamoDB. What is the MOST likely cause of the delay?

A.Lambda function has a cold start that adds several seconds to processing time.
B.S3 is eventually consistent for new object writes, so the object may not be immediately available.
C.S3 Event Notifications may have a slight delay, and the application polls for the processed image before the notification triggers Lambda.
D.DynamoDB has insufficient read capacity causing throttling on metadata retrieval.
AnswerC

S3 Event Notifications are delivered asynchronously and on a best-effort basis, meaning there can be an inherent, variable delay between an object being uploaded and the corresponding Lambda function being invoked. If the application immediately polls for the *processed* image after the initial upload, it creates a race condition where the polling might occur before the S3 event has triggered the Lambda function to process the image, or before the processing itself has completed and the processed image is stored. This asynchronous nature and potential latency in event delivery are a common cause for such perceived delays.

Why this answer

S3 Event Notifications are delivered asynchronously and can take seconds to minutes to reach Lambda, so if the application polls for the processed image or metadata immediately after upload, it may query before Lambda has run. This race condition between the upload and the notification-driven processing is the most likely cause of the intermittent delay.

Exam trap

DVA-C02 often tests whether candidates still believe S3 is eventually consistent for new writes (it has been strongly consistent since 2020) and whether they understand that event notifications are asynchronous, so the trap is blaming cold starts or DynamoDB throttling instead of the notification latency and polling race.

How to eliminate wrong answers

Option A is wrong because Lambda cold starts add hundreds of milliseconds to a few seconds, not the kind of noticeable delay described, and they would not cause 'sometimes not immediately available' behavior tied to polling. Option B is wrong because since December 2020 S3 provides strong read-after-write consistency for new object PUTs, so the object is immediately readable — the old eventual-consistency model no longer applies. Option D is wrong because DynamoDB read throttling would produce errors or retries, not a delay in image availability, and the scenario describes metadata retrieval rather than capacity exhaustion.

85
MCQeasy

A developer is troubleshooting a slow Amazon RDS MySQL database query. The query is frequently executed and takes 5 seconds to complete. Which AWS service should the developer use to analyze the query performance?

A.AWS CloudTrail
B.Amazon RDS Performance Insights
C.Amazon CloudWatch Logs
D.AWS X-Ray
AnswerB

Performance Insights is purpose-built for this scenario: it visualizes database load as Average Active Sessions, breaks load down by SQL statement, wait event, host, and user, and lets the developer pinpoint exactly which query is consuming the most database time and why.

Why this answer

Amazon RDS Performance Insights is the correct service to analyze query performance on an RDS MySQL database. It provides a dashboard that visualizes database load and helps identify the top SQL statements, wait events, and users consuming the most resources. This allows the developer to pinpoint the slow query and understand its impact.

Exam trap

DVA-C02 often tests the confusion between monitoring services like CloudWatch and specialized database performance tools like Performance Insights, leading candidates to choose CloudWatch Logs for query analysis.

How to eliminate wrong answers

Option A is wrong because AWS CloudTrail records API activity and management events, not database query performance. Option C is wrong because Amazon CloudWatch Logs can capture database logs, but it does not provide the detailed performance analysis and visualization that Performance Insights offers; it would require manual parsing and analysis. Option D is wrong because AWS X-Ray is used for tracing requests in distributed applications, not for analyzing database query performance.

86
MCQmedium

An AWS Lambda function processes messages from an Amazon SQS queue and writes results to an Amazon DynamoDB table. The function is configured with a reserved concurrency of 5 and a batch size of 10. CloudWatch metrics show high throttling and a growing queue backlog. The function's execution time averages 1 second per message. What is the MOST effective action to reduce throttling while improving throughput?

A.Increase the reserved concurrency to 20.
B.Increase the batch size to 100.
C.Decrease the reserved concurrency to 2.
D.Increase the provisioned write capacity of the DynamoDB table.
AnswerA

Increasing reserved concurrency allows Lambda to scale and invoke more function instances concurrently. This directly reduces throttling and allows the function to process more messages from the SQS queue simultaneously, improving throughput and reducing backlog.

Why this answer

The Lambda function is throttling because its reserved concurrency of 5 limits it to 5 concurrent executions. With a batch size of 10 and 1-second execution time, the function can process at most 5 * 10 = 50 messages per second. Increasing reserved concurrency to 20 allows 20 concurrent executions, raising throughput to 200 messages per second, which directly reduces throttling and clears the backlog.

Exam trap

The trap here is that candidates may confuse Lambda throttling with downstream resource throttling (like DynamoDB) and choose to increase write capacity, or they may think increasing batch size alone will solve the problem without considering the concurrency bottleneck.

How to eliminate wrong answers

Option B is wrong because increasing batch size to 100 would cause each invocation to process more messages, but with only 5 concurrent executions, the function would still be limited to 5 invocations at a time, and the 1-second execution time per message would scale linearly, likely causing timeouts or increased latency without addressing the root cause of throttling. Option C is wrong because decreasing reserved concurrency to 2 would reduce throughput to 20 messages per second, worsening throttling and backlog. Option D is wrong because increasing DynamoDB write capacity addresses potential write throttling from DynamoDB, but the CloudWatch metrics show Lambda throttling, not DynamoDB throttling; the bottleneck is Lambda concurrency, not the database.

87
MCQmedium

A developer is troubleshooting an application that uses Amazon ElastiCache for Redis to improve performance. The application periodically experiences high latency during peak hours. The developer checks the ElastiCache metrics and sees that the 'Evictions' metric is consistently high and the 'CacheHitRate' metric is low. The cluster has a single node with a cache.t3.small instance type. Which action will most likely improve the cache hit rate and reduce latency?

A.Scale up to a larger node type (e.g., cache.t3.medium) to increase available memory.
B.Enable cluster mode and distribute data across multiple shards to reduce memory pressure.
C.Change the eviction policy to 'allkeys-lfu' to better manage which keys are evicted.
D.Add a read replica for the Redis cluster to offload read traffic.
AnswerA

Scaling up to a larger node type directly increases the available RAM for the Redis instance. This additional memory allows the cache to store more data, significantly reducing the frequency of key evictions caused by memory pressure. Consequently, the cache hit rate improves, as more requested data is found in cache, leading to lower latency and better application performance by minimizing database lookups.

Why this answer

The high 'Evictions' and low 'CacheHitRate' metrics indicate that the Redis node is running out of memory, forcing it to evict keys to make room for new data. Scaling up to a larger node type (cache.t3.medium) increases the available memory, allowing more data to be cached and reducing evictions, which directly improves the cache hit rate and reduces latency.

Exam trap

The trap here is that candidates may focus on optimizing eviction policies or adding replicas, but the core issue is insufficient memory capacity, which only scaling up can resolve.

How to eliminate wrong answers

Option B is wrong because enabling cluster mode and distributing data across multiple shards does not increase the total memory per node; it only partitions data, and if the total memory across shards is insufficient, evictions will still occur. Option C is wrong because changing the eviction policy to 'allkeys-lfu' only changes which keys are evicted (least frequently used) but does not address the root cause of insufficient memory; evictions will continue at the same rate. Option D is wrong because adding a read replica offloads read traffic but does not increase the primary node's memory, so evictions and low cache hit rate will persist on the primary node.

88
Multi-Selecteasy

A developer is using AWS X-Ray to trace requests through a microservices application. The developer notices that some traces are incomplete. Which TWO actions can help ensure complete traces?

Select 2 answers
A.Use the X-Ray SDK to instrument the application code.
B.Open port 2000 on the security groups for TCP traffic.
C.Deploy the X-Ray daemon as a centralized service in a separate instance.
D.Install the CloudWatch agent on all instances.
E.Ensure the X-Ray daemon is running on all EC2 instances.
AnswersA, E

The X-Ray SDK must be integrated directly into the application code because it is what creates trace data in the first place. For supported web frameworks, middleware or interceptors automatically capture incoming HTTP requests, generate a trace ID, manage segments and subsegments, and propagate the X-Amzn-Trace-Id header to downstream services. The SDK then sends completed segments to the local X-Ray daemon over UDP port 2000 for eventual upload to the X-Ray API. Without this code-level instrumentation, a request never becomes a trace, regardless of daemon status or network configuration.

Why this answer

Option A is correct because the X-Ray SDK must be used to instrument the application code so that it emits segment data and propagates trace headers across service calls; without instrumentation, downstream services cannot contribute subsegments and traces remain incomplete. Option E is correct because the X-Ray daemon must be running on every EC2 instance (or equivalent compute) to receive UDP segment data from the SDK on port 2000 and forward it to the X-Ray API; a missing daemon on any instance means those segments are dropped, producing incomplete traces. Option B is not correct because opening port 2000 in a security group is not required for the daemon to receive local UDP traffic from the SDK on the same host, and it does not by itself fix incomplete traces.

Option C is not correct because the X-Ray daemon is designed to run locally on each instance (or as a sidecar/daemon set), not as a single centralized service, so centralizing it would not ensure all instances' segments are captured. Option D is not correct because the CloudWatch agent collects metrics and logs, not X-Ray trace segments, so it does not contribute to complete X-Ray traces.

Exam trap

The trap is thinking that network configuration (port 2000) or centralized daemon deployment solves incomplete traces, when the real requirements are SDK instrumentation and a running daemon on every instance.

89
MCQhard

A Lambda function using a Kinesis event source repeatedly retries one bad record and blocks progress in the shard. Which feature helps isolate failed records after retry limits?

A.Increase memory to 10 GB only
B.Disable batch processing
C.Configure failure handling with bisect batch on error and an on-failure destination where supported
D.Convert the stream to an S3 bucket
AnswerC

Configuring `ReportBatchItemFailures` (often referred to as "bisect batch on error" in the console) for a Kinesis event source allows the Lambda function to return a partial success, indicating which specific records within a batch failed. Lambda then automatically retries only the failed records, potentially splitting the batch further to isolate the problematic items. Combining this with an on-failure destination, such as an SQS queue or SNS topic, ensures that records that ultimately cannot be processed are sent to a dead-letter queue for analysis and manual intervention, preventing them from indefinitely blocking the stream processing.

Why this answer

Lambda's Kinesis event source mapping supports a 'bisect batch on error' feature that splits a failed batch into two smaller batches, allowing the bad record to be isolated and retried separately. Additionally, configuring an on-failure destination (e.g., an SQS queue or SNS topic) sends the record to a dead-letter destination after the retry limit is exhausted, preventing the shard from blocking progress.

Exam trap

The trap here is that candidates often think increasing memory or disabling batch processing will solve the blocking issue, but they fail to recognize that only explicit failure handling with bisect and a dead-letter destination can isolate and remove the bad record without manual intervention.

How to eliminate wrong answers

Option A is wrong because increasing memory to 10 GB only allocates more CPU and memory to the function, but does not address the root cause of a single bad record blocking the shard; it does not provide any mechanism to isolate or skip failed records. Option B is wrong because disabling batch processing (setting batch size to 1) would still cause the same blocking behavior—each record would be processed individually, but a persistent bad record would still be retried indefinitely, blocking the shard. Option D is wrong because converting the stream to an S3 bucket is not a direct replacement for Kinesis event processing; S3 does not support the same record-level retry and failure handling semantics, and this would require a complete architectural change, not a simple configuration fix.

90
Multi-Selecthard

A company is using AWS CodePipeline for CI/CD. The pipeline has a build stage using AWS CodeBuild, and a deploy stage using AWS CodeDeploy. The deployment is failing with 'Error: Health checks failed'. Which TWO steps should the developer take to troubleshoot this issue? (Select TWO.)

Select 2 answers
A.Verify that the target group's health check path and port are correctly configured.
B.Check the S3 bucket where the build artifacts are stored.
C.Check the CodeDeploy deployment logs for detailed error messages.
D.Check the CodeBuild build logs for errors.
E.Increase the number of EC2 instances in the Auto Scaling group.
AnswersA, C

CodeDeploy's health check failures during an in-place or blue/green deployment are most often caused by the target group's health check path returning a non-2xx response or the port not matching what the application actually listens on, so confirming this configuration directly addresses the failure signal.

Why this answer

Options A and C are correct. Verifying the target group's health check path and port (A) ensures that the load balancer's health check matches the application's actual endpoint, which is a common cause of health check failures. Checking the CodeDeploy deployment logs (C) provides detailed error messages from the deployment process, which can pinpoint why the health checks are failing.

Option B (checking S3 bucket) is not directly related to health check failures, as artifacts are typically stored correctly if the build succeeded. Option D (checking CodeBuild logs) is irrelevant because the build stage succeeded. Option E (increasing instances) does not address the root cause of health check failures.

91
MCQhard

A company runs a critical application on Amazon EC2 instances behind an Application Load Balancer (ALB). The application experiences intermittent errors where some requests return HTTP 503 (Service Unavailable) errors. The developers have verified that the application code is healthy and the EC2 instances pass health checks. The ALB health check is configured to hit a specific endpoint (/health) with a healthy threshold of 2 and an unhealthy threshold of 2. The health check interval is 30 seconds, and the timeout is 5 seconds. The application's /health endpoint sometimes takes up to 6 seconds to respond due to a dependency on a third-party service. The developers want to minimize the 503 errors without changing the application code. Which action should the developer take?

A.Increase the health check timeout to 10 seconds to accommodate the slow /health endpoint.
B.Decrease the unhealthy threshold to 1 so that instances are marked unhealthy after one failed health check.
C.Increase the deregistration delay to 300 seconds to allow connections to drain.
D.Decrease the health check interval to 10 seconds to detect health changes faster.
AnswerA

The /health endpoint can take six seconds, exceeding the five-second timeout, so the ALB marks instances unhealthy and returns 503s. Raising the timeout to ten seconds lets the health check succeed, keeping instances in service without code changes.

Why this answer

Increasing the health check timeout to 10 seconds allows the /health endpoint to respond within the timeout period, preventing the ALB from marking the instance as unhealthy due to a slow response. Since the endpoint sometimes takes up to 6 seconds, a 5-second timeout is too short, causing health checks to fail intermittently. By increasing the timeout, the health checks will succeed, and the ALB will not remove the instance from service, thus reducing 503 errors.

Exam trap

The trap is that candidates might think decreasing the unhealthy threshold or interval would help, but those actions would make the situation worse; the key is to match the timeout to the application's response time.

How to eliminate wrong answers

Option B is wrong because decreasing the unhealthy threshold to 1 would make the ALB mark instances unhealthy faster, potentially increasing 503 errors if a single health check fails. Option C is wrong because increasing the deregistration delay affects connection draining during instance deregistration, not health check failures; it does not address the root cause. Option D is wrong because decreasing the health check interval to 10 seconds would make health checks more frequent, but with a 5-second timeout, the slow endpoint would still cause failures; it might even increase the number of failed health checks.

92
Multi-Selectmedium

Which THREE factors should a developer consider when designing a stateless application on AWS? (Choose 3)

Select 3 answers
A.Avoid storing data on the local file system of the instances
B.Store session state in a shared external datastore like ElastiCache
C.Store session state in the instance memory for low latency
D.Use sticky sessions on the load balancer to maintain session affinity
E.Use a shared database like Amazon DynamoDB for persistent data
AnswersA, B, E

A stateless application must not persist session or transactional data on an instance's local disk, because Auto Scaling can terminate or replace that instance at any time, and any data written only to its EBS root or instance store volume is permanently lost with it.

Why this answer

A stateless application should not store session state locally, so option A is correct. Session state should be stored in an external shared datastore like ElastiCache (option B) or a shared database like DynamoDB (option E). Storing state in instance memory (C) or using sticky sessions (D) would introduce statefulness, which is not desired in a stateless architecture.

93
MCQeasy

A developer is troubleshooting a web application that intermittently returns HTTP 504 errors. The application runs on EC2 instances behind an Application Load Balancer. What is the most likely cause of these errors?

A.The target group is using an HTTPS health check but the instances only support HTTP.
B.The load balancer's cross-zone load balancing is disabled.
C.The load balancer idle timeout is set too low, and the application takes longer than the timeout to respond.
D.The security group for the EC2 instances is missing an inbound rule for the load balancer.
AnswerC

The load balancer idle timeout specifies the maximum duration the load balancer will wait for a response from a registered target before closing the connection. If the backend application takes longer to process a request and send a response than this configured timeout, the load balancer will terminate the connection. This action directly results in an HTTP 504 Gateway Timeout error being returned to the client, indicating a lack of timely response.

Why this answer

HTTP 504 (Gateway Timeout) errors from an Application Load Balancer indicate that the load balancer successfully connected to the target (EC2 instance) but the target did not respond within the configured idle timeout period. The default idle timeout is 60 seconds, and if the application's processing time exceeds this value, the load balancer terminates the connection and returns a 504. Option C directly addresses this mismatch between the load balancer timeout and the application response time.

Exam trap

The trap here is that candidates often confuse HTTP 504 (Gateway Timeout) with HTTP 502 (Bad Gateway) or health check failures, leading them to select options related to security groups or health check mismatches instead of the correct idle timeout configuration.

How to eliminate wrong answers

Option A is wrong because HTTPS health checks require the target to support HTTPS; if the instances only support HTTP, the health check would fail and the instances would be marked unhealthy, leading to 503 errors (not 504). Option B is wrong because disabling cross-zone load balancing affects traffic distribution across Availability Zones, not the timeout behavior that causes 504 errors. Option D is wrong because a missing inbound security group rule for the load balancer would prevent the load balancer from establishing connections to the instances, resulting in 502 errors or health check failures, not intermittent 504 timeouts.

94
Multi-Selecteasy

A developer is using an Amazon SQS queue with a Lambda function as a consumer. Messages are being sent to the queue but the Lambda function is not processing them. Which THREE of the following are possible causes?

Select 3 answers
A.The SQS queue has a dead-letter queue configured.
B.The SQS queue policy denies access to the Lambda function.
C.The Lambda function's execution role does not have sqs:ReceiveMessage permission.
D.The SQS queue has a rate limit that prevents Lambda from polling.
E.The event source mapping between SQS and Lambda is disabled.
AnswersB, C, E

An SQS queue policy is a resource-based policy that defines who can access the queue and what actions they can perform. If this policy contains an explicit Deny statement for the sqs:ReceiveMessage action (or other relevant polling actions) for the Lambda function's execution role, the Lambda service will be unable to poll messages from the queue. This explicit denial takes precedence over any Allow statements in the Lambda's IAM role, effectively blocking access.

Why this answer

The SQS queue policy is a resource-based policy that controls which principals (like Lambda's execution role) can perform actions on the queue. If the policy explicitly denies the Lambda function's access, the function will not be able to poll or delete messages from the queue, even if its own execution role grants those permissions.

Exam trap

The trap here is that candidates often overlook resource-based policies (like SQS queue policies) and focus only on the Lambda execution role, assuming that if the role has permissions, the integration will work, but the queue policy can independently deny access.

95
MCQeasy

A developer is deploying a serverless application using AWS CloudFormation. The stack creation fails with the error 'The following resource(s) failed to create: [MyLambdaFunction]'. The developer checks the CloudWatch logs but finds no logs for the Lambda function. What is the most likely reason?

A.The Lambda function code has a syntax error that prevents creation.
B.The Lambda function was never invoked.
C.The Lambda function's IAM role does not have permission to write to CloudWatch Logs.
D.The CloudFormation template has a syntax error.
AnswerD

A syntax error in the CloudFormation template, such as a missing required property for the Lambda resource, can cause creation failure. The error message and lack of logs are consistent with this.

Why this answer

A CloudFormation template syntax error, such as a missing required property (e.g., 'Handler' or 'Runtime'), can cause the Lambda function resource to fail creation. The error message indicates the specific resource failed, and the absence of CloudWatch logs is expected because the function was never created. Option A is incorrect because code syntax errors do not prevent resource creation; they affect invocation.

Option B is incorrect because the function not being invoked is a consequence, not the cause. Option C is incorrect because missing CloudWatch Logs permissions do not prevent creation; they only affect logging during invocation.

Exam trap

The key trap is that candidates may assume the absence of logs means the function was never invoked or that a code error existed. However, the error explicitly states the resource failed to create, so the function was never deployed. The missing logs are a direct consequence of the creation failure, not a separate issue.

The root cause is often a CloudFormation template error, such as a missing required property (e.g., 'Handler' or 'Runtime'), which causes the Lambda resource to fail creation.

How to eliminate wrong answers

Option A is wrong because a syntax error in the Lambda function code would not prevent the resource from being created; CloudFormation would still create the function, but invocation would fail, and logs would appear (if permissions allow). Option B is wrong because the error message states the resource failed to create, meaning the function was never successfully created, so it cannot be invoked; the absence of logs is not due to lack of invocation but due to creation failure. Option D is wrong because a CloudFormation template syntax error would cause a different error (e.g., 'Template validation error') and would prevent the entire stack from being parsed, not just a single resource creation failure.

96
MCQeasy

An application running on Amazon EC2 instances behind an Application Load Balancer (ALB) is experiencing intermittent 503 errors. The EC2 instances are in an Auto Scaling group. What is the MOST likely cause?

A.The SSL certificate on the ALB has expired.
B.The target group health checks are failing.
C.The ALB DNS name is not resolving.
D.The security group for the ALB is blocking traffic.
AnswerB

When all registered instances within an Application Load Balancer (ALB) target group fail their configured health checks, the ALB marks them as unhealthy and stops routing traffic to them. If there are no healthy targets remaining in any associated target group, the ALB cannot fulfill incoming client requests. Consequently, the ALB returns an HTTP 503 Service Unavailable error, indicating that while the load balancer itself is operational, it has no available backend resources to process the request.

Why this answer

The intermittent 503 errors indicate that the ALB temporarily has no healthy targets to forward requests to. When target group health checks fail, the ALB marks instances as unhealthy and stops routing traffic to them, causing a 503 response if all instances are unhealthy. This aligns with the Auto Scaling group potentially launching new instances that haven't passed health checks yet, or existing instances failing health checks due to application overload or misconfiguration.

Exam trap

The trap here is that candidates often confuse 503 errors with SSL or DNS issues, but 503 specifically indicates the ALB is reachable and functioning but has no healthy targets to serve the request.

How to eliminate wrong answers

Option A is wrong because an expired SSL certificate on the ALB would cause TLS handshake failures (e.g., 502 Bad Gateway or connection errors), not intermittent 503 errors; the ALB would still route traffic to healthy targets. Option C is wrong because if the ALB DNS name were not resolving, clients would receive a DNS resolution failure (NXDOMAIN) or timeout, not an HTTP 503 error from the ALB. Option D is wrong because if the security group for the ALB were blocking traffic, clients would receive a timeout or connection refused error, not an HTTP 503 response; the ALB would not be reachable at all.

97
MCQmedium

A developer invokes a Lambda function from the AWS CLI and receives the response shown in the exhibit. The output file contains an error message. What is the MOST likely cause of the FunctionError field being set to 'Unhandled'?

A.The function's execution role does not have permission to write to CloudWatch Logs.
B.The invocation request payload exceeded the 6 KB limit for synchronous invocation.
C.The function code threw an uncaught exception.
D.The function timed out before completing execution.
AnswerC

When a Lambda function's code encounters an error that is not explicitly handled by the developer (e.g., via a try-catch block), it results in an uncaught exception. The Lambda runtime then terminates the execution, sets FunctionError: Unhandled in the response, and returns a StatusCode 200 to the invoker, indicating the service successfully processed the invocation request. The specific error message from this exception is then captured and written to the output file specified in the AWS CLI command.

Why this answer

The FunctionError field set to 'Unhandled' in a Lambda invocation response indicates that the function's code raised an uncaught exception during execution, so option C is correct. AWS Lambda returns this value when the runtime catches an error that the handler did not handle, and the error details appear in the response payload. Option A would typically cause logging failures rather than an 'Unhandled' FunctionError, and could even prevent error details from being logged.

Option B is incorrect because the synchronous invocation payload limit is 6 MB, not 6 KB, and exceeding it produces a request validation error, not an 'Unhandled' FunctionError. Option D is incorrect because a timeout produces a different error type (for example, 'Task timed out after X seconds') rather than the generic 'Unhandled' value.

98
MCQhard

A developer is troubleshooting an application that uses Amazon SQS. Messages are being sent to a dead-letter queue (DLQ) after the maximum receive count is exceeded. The consumer processes messages but sometimes fails. The developer wants to ensure that messages are retried immediately after a failure, without waiting for the visibility timeout. Which solution should the developer implement?

A.Configure a delay queue so that messages are not immediately visible after failure.
B.After a failure, call ChangeMessageVisibility with a timeout of 0 to make the message immediately available for reprocessing.
C.Delete the message from the queue and re-send it after processing failure.
D.Increase the visibility timeout to allow more time for processing.
AnswerB

When a consumer receives a message from an SQS queue, the message becomes invisible to other consumers for the duration of the queue's visibility timeout. If processing fails, calling `ChangeMessageVisibility` with a `VisibilityTimeout` of 0 seconds explicitly overrides the current visibility timeout. This action immediately makes the message visible again in the queue, allowing another consumer (or the same consumer on a subsequent poll) to receive and attempt to process it without delay, thus enabling immediate reprocessing.

Why this answer

The correct solution is to call ChangeMessageVisibility with a timeout of 0 after a failure. This makes the message immediately available for reprocessing by other consumers, effectively retrying without waiting for the visibility timeout. Option A (delay queue) introduces a delay, which is counterproductive.

Option C (delete and re-send) loses original message attributes and is inefficient. Option D (increase visibility timeout) would delay retries, not accelerate them.

99
MCQmedium

A developer is working on a serverless application that uses AWS Lambda functions to process user uploads. The uploads are stored in an S3 bucket, and each upload triggers a Lambda function that resizes images and stores metadata in DynamoDB. Recently, users have reported that some images are not being resized. The developer checks the CloudWatch logs and sees that the Lambda function is invoked, but it fails with a timeout error after 15 seconds for a few large images. The function has a timeout of 15 seconds and a memory of 512 MB. The image sizes vary from 1 MB to 50 MB. The developer wants to handle large images without increasing the timeout significantly, as that would increase costs. The function is CPU-bound during image processing. Which solution should the developer implement?

A.Increase the memory allocated to the Lambda function to 3008 MB, which also increases CPU power, allowing faster processing within the same timeout.
B.Split the large images into smaller chunks before uploading to S3, then reassemble them after processing.
C.Increase the Lambda function timeout to 5 minutes to accommodate large images.
D.Use AWS Step Functions to orchestrate the image processing workflow, allowing longer timeouts for individual steps.
AnswerA

This is the most effective solution for CPU-bound tasks in AWS Lambda. Increasing the memory allocated to a Lambda function directly scales its proportional share of CPU power, network bandwidth, and disk I/O. For computationally intensive operations like image processing, providing more CPU resources allows the function to complete the task significantly faster, often within the existing timeout, and can even reduce overall costs by decreasing the total execution duration.

Why this answer

Option A is correct because AWS Lambda allocates CPU power proportionally to memory, so raising memory from 512 MB to 3008 MB gives the function roughly 6x more CPU, letting the CPU-bound image resizing finish within the existing 15-second timeout without extending the timeout and thus without increasing per-invocation duration costs. This directly addresses the timeout caused by CPU-bound processing of large images. Option B is impractical because chunking and reassembling images adds complexity and does not speed up the CPU-bound resize work.

Option C would allow more time but increases cost and duration, which the developer explicitly wants to avoid. Option D does not solve the underlying CPU bottleneck and Step Functions cannot extend a single Lambda invocation beyond its configured timeout.

100
MCQeasy

A developer notices that an RDS MySQL instance's CPU utilization is consistently above 80% during peak hours. Which AWS service can be used to analyze the database queries and identify the root cause?

A.RDS Performance Insights
B.RDS Enhanced Monitoring
C.AWS X-Ray
D.CloudWatch Logs Insights
AnswerA

RDS Performance Insights is specifically designed to help developers understand database load by visualizing active sessions and identifying the top SQL queries, hosts, or users consuming database resources. It provides a graphical dashboard that correlates database load with specific SQL statements, wait events, and user activity, making it ideal for pinpointing the exact queries or processes contributing to high CPU utilization on an RDS MySQL instance. This tool directly answers the need to identify *what* is causing the CPU spike, not just that a spike occurred.

Why this answer

RDS Performance Insights is purpose-built for database performance analysis: it collects and visualizes database load (Average Active Sessions) broken down by SQL statement, wait event, user, and host, so you can pinpoint which query is driving CPU above 80%. It retains up to 2 years of performance data and requires no agent installation on the DB instance. This directly answers the question of identifying the root-cause query.

Exam trap

The trap here is confusing OS-level monitoring (Enhanced Monitoring) with database-level query analysis — candidates who only read 'CPU utilization' pick Enhanced Monitoring, but the question's key phrase is 'analyze the database queries.'

How to eliminate wrong answers

Option B is wrong because Enhanced Monitoring reports OS-level metrics (CPU, memory, disk, network) at up to 1-second granularity but does not break down activity by SQL statement, so it cannot identify the offending query. Option C is wrong because AWS X-Ray traces requests through distributed applications (Lambda, API Gateway, microservices) and has no visibility into RDS MySQL internals or SQL execution. Option D is wrong because CloudWatch Logs Insights queries log data (e.g., slow query logs) but requires you to have already enabled and shipped those logs, and it does not provide the built-in database-load visualization or wait-event breakdown that Performance Insights offers out of the box.

101
Multi-Selecthard

A company has a REST API deployed on Amazon API Gateway with a Lambda integration. The API is experiencing high latency. Which TWO actions would help diagnose the issue?

Select 2 answers
A.Use AWS X-Ray to trace requests.
B.Enable detailed CloudWatch Logs for the API Gateway stage.
C.Increase the Lambda function memory.
D.Change the API integration type from Lambda to HTTP.
E.Add Amazon CloudFront in front of API Gateway.
AnswersA, B

AWS X-Ray is a distributed tracing service that provides an end-to-end view of requests as they flow through various services, including API Gateway and Lambda. It generates a service map and detailed trace data, showing the latency incurred at each hop. This allows developers to precisely identify which component – be it the API Gateway's processing, the Lambda function's execution, or a downstream dependency called by Lambda – is contributing most significantly to the overall API latency, making it an ideal diagnostic tool.

Why this answer

Option A is correct because AWS X-Ray traces the full request path through API Gateway, the Lambda integration, and downstream calls, exposing where latency is introduced (e.g., Lambda cold starts, slow downstream calls) via trace segments and subsegments. Option B is correct because enabling detailed CloudWatch Logs (execution logging and access logging) for the API Gateway stage captures per-request latency metrics such as integrationLatency and responseLatency, helping pinpoint whether the delay is in API Gateway or the Lambda backend. Option C is not a diagnostic action; increasing Lambda memory changes CPU allocation and could improve performance but does not help identify the root cause of latency.

Option D is not a diagnostic step and changing the integration type alters architecture rather than measuring where latency occurs. Option E is also a remediation/architecture change (caching and edge termination) rather than a way to diagnose the source of the high latency.

Exam trap

Candidates often confuse performance optimization actions (like increasing Lambda memory or adding CloudFront) with diagnostic actions. The question specifically asks for actions to 'diagnose' the issue, not to resolve it.

102
MCQhard

Refer to the exhibit. A CloudFormation stack creation failed. What is the most likely cause of the failure?

A.The IAM role's trust policy does not allow Lambda to assume the role.
B.The Lambda function name conflicts with an existing function.
C.The Lambda function code has a syntax error.
D.The Lambda execution role does not have the required permissions to write to CloudWatch Logs.
AnswerA

The trust policy of an IAM role specifies which entities are allowed to assume that role. For a Lambda function to successfully execute, its associated execution role must have a trust policy that explicitly lists 'lambda.amazonaws.com' as a principal. If this service principal is missing or incorrectly configured, the AWS Lambda service cannot assume the role, leading to a CloudFormation stack creation failure as the function cannot be properly provisioned.

Why this answer

The correct answer is A: the IAM role's trust policy does not allow Lambda to assume the role. In CloudFormation, when a stack provisions a Lambda function with an execution role, the role's trust policy must include lambda.amazonaws.com as a principal with sts:AssumeRole; if that trust relationship is missing or misconfigured, role assumption fails and the stack creation fails. Option B is unlikely because Lambda function names only need to be unique within an account and Region, and CloudFormation would typically surface a naming conflict differently.

Option C is not the most likely cause because a code syntax error usually causes invocation failures after deployment, not stack creation failure. Option D would cause runtime logging failures, but the role could still be assumed and the function created, so it would not typically fail the stack at creation time.

103
MCQhard

A developer is troubleshooting an AWS Elastic Beanstalk environment that is failing health checks. The environment runs a web application on Tomcat. The developer checks the logs and finds no errors. What is the most likely cause of the health check failure?

A.The application's health check URL is returning a non-200 status code.
B.The security group for the instances does not allow traffic from the load balancer.
C.The application is throwing exceptions that are not logged.
D.The application is listening on a port other than 80.
AnswerA

Elastic Beanstalk environments rely on health checks, typically performed by the associated Load Balancer, to determine the operational status of application instances. If the configured health check URL, often the root path "/", consistently returns a non-200 HTTP status code, the Load Balancer will mark the instance as unhealthy. This leads to the instance being removed from the target group, preventing traffic, and can cause Elastic Beanstalk to report a "Degraded" or "Severe" environment health status, triggering instance replacement or environment instability.

Why this answer

The most likely cause is that the application's health check URL is returning a non-200 status code. Elastic Beanstalk uses the load balancer to perform health checks against a configurable path (default: /). If the application responds with any status other than 200 OK, the load balancer marks the instance as unhealthy, even if the application logs show no errors.

This is a common misconfiguration where the health check endpoint is not implemented or returns an unexpected status.

Exam trap

The trap here is that candidates assume health check failures are always due to network or infrastructure issues (security groups, ports) rather than application-level misconfigurations like a missing or incorrect health check endpoint.

How to eliminate wrong answers

Option B is wrong because if the security group blocked traffic from the load balancer, the instances would be unreachable entirely, not just failing health checks, and the logs would likely show connection timeouts or refused connections. Option C is wrong because unlogged exceptions would still typically result in a non-200 response or an error page, which would be reflected in the health check status; the question states logs show no errors, making this unlikely. Option D is wrong because Elastic Beanstalk configures the load balancer to forward traffic to the correct port (e.g., 8080 for Tomcat), and the health check is sent to that same port; listening on a different port would cause a connection failure, not a health check failure with no errors in logs.

104
MCQhard

A company runs a stateful web application on EC2 instances in an Auto Scaling group. Users report that their session data is lost when instances are replaced during scaling events. What is the best solution to preserve session state?

A.Use ElastiCache as a centralized session store.
B.Enable sticky sessions on the Application Load Balancer.
C.Store sessions in the Application Load Balancer.
D.Use an S3 bucket to store session data.
AnswerA

ElastiCache, particularly Redis, provides an extremely fast, in-memory data store that is ideal for managing user session data. By centralizing sessions here, any EC2 instance can retrieve or update a user's session state with very low latency, ensuring a consistent experience even if the user's subsequent requests are routed to a different application instance. This approach effectively decouples the session state from individual application servers, making the web application highly scalable, resilient to instance failures, and easier to manage in an auto-scaling environment.

Why this answer

ElastiCache provides a centralized, in-memory session store that is external to the EC2 instances. This ensures session data persists independently of the instance lifecycle, so when an instance is replaced during a scaling event, the new instance can retrieve the session from ElastiCache, preserving user state. This is the best solution because it decouples session state from compute resources, aligning with the stateless application pattern recommended for Auto Scaling groups.

Exam trap

The trap here is that candidates often confuse sticky sessions (option B) with session persistence, not realizing that sticky sessions only maintain request routing to the same instance, not the session data itself when the instance is replaced.

How to eliminate wrong answers

Option B is wrong because sticky sessions (session affinity) only route a user to the same instance, but they do not preserve session data when that instance is terminated and replaced; the session is still lost. Option C is wrong because the Application Load Balancer does not store session data; it only forwards requests and can manage cookies for stickiness, but the session state itself must be stored elsewhere. Option D is wrong because S3 is an object store with higher latency and is not designed for low-latency, frequent read/write operations required for session management; it would introduce unacceptable performance overhead and is not a session store.

105
MCQmedium

A developer is optimizing a DynamoDB table for a gaming leaderboard. The table stores player scores and is read-heavy. Queries often fetch the top 10 scores. Which indexing strategy best reduces RCU consumption?

A.Create a sparse index on player ID.
B.Use a local secondary index on score.
C.Enable DynamoDB Accelerator (DAX) for caching.
D.Create a global secondary index with score as the sort key.
AnswerD

A Global Secondary Index (GSI) has its own independent partition key and sort key, allowing for different access patterns than the base table. By creating a GSI with a common partition key (e.g., a static value like "LEADERBOARD") and 'score' as the sort key, all player scores can be grouped and efficiently queried. This enables a `Query` operation on the GSI, ordered by 'score' in descending order, to retrieve the top N results with minimal RCU consumption, scaling independently from the base table.

Why this answer

A global secondary index (GSI) with score as the sort key allows efficient retrieval of the top 10 scores by querying the index in descending order, reading only the required items. This minimizes read capacity unit (RCU) consumption compared to scanning the base table, as each query reads exactly 10 items (or fewer) rather than consuming RCUs for a full table scan or filtering large result sets.

Exam trap

The trap here is that candidates often confuse local secondary indexes (LSIs) with global secondary indexes (GSIs), not realizing that LSIs are tied to the base table's partition key and cannot efficiently retrieve global top scores across all partitions.

How to eliminate wrong answers

Option A is wrong because a sparse index on player ID would not help retrieve top scores; it only indexes items where player ID is present, and querying by player ID does not sort by score. Option B is wrong because a local secondary index (LSI) on score is constrained to the same partition key as the base table, requiring a full partition scan to get top scores across all partitions, which consumes more RCUs. Option C is wrong because DynamoDB Accelerator (DAX) is an in-memory cache that reduces latency and read load, but it does not change the underlying query pattern or RCU consumption for fetching top scores; the base table or index still needs to be queried, and DAX caches results after the first read, not reducing RCUs for the initial query.

106
MCQhard

A developer attached the IAM policy above to an IAM user. The user reports being unable to list objects in the bucket 'my-bucket' using the AWS CLI command 'aws s3 ls s3://my-bucket/'. What is the most likely reason?

A.The IAM policy does not allow the s3:GetObject action on the bucket.
B.The IAM policy resource for s3:ListBucket should include the bucket and objects.
C.The IAM policy is missing the s3:ListAllMyBuckets action.
D.The IAM policy does not include the s3:GetBucketLocation action.
AnswerD

The AWS Command Line Interface (CLI) and SDKs often perform an implicit s3:GetBucketLocation API call to determine the region of an S3 bucket. This action is crucial for the client to correctly route subsequent S3 API requests to the appropriate regional endpoint, especially if the region is not explicitly specified in the command or configuration. Without this permission, the CLI or SDK may fail to connect to the bucket, even if s3:ListBucket is allowed.

Why this answer

The `aws s3 ls s3://my-bucket/` command requires the `s3:GetBucketLocation` permission to determine the bucket's region before listing its contents. Without this action, the CLI fails with an error like 'An error occurred (AccessDenied) when calling the GetBucketLocation operation', even if `s3:ListBucket` is granted. Option D correctly identifies this missing permission as the root cause.

Exam trap

The trap here is that candidates often focus on the `ListBucket` permission and overlook the prerequisite `GetBucketLocation` call, assuming the CLI only needs the list action for the `ls` command.

How to eliminate wrong answers

Option A is wrong because `s3:GetObject` is not required for listing objects; it is needed for downloading objects, not for the `ls` command. Option B is wrong because the resource for `s3:ListBucket` should be the bucket ARN (`arn:aws:s3:::my-bucket`), not the bucket and objects; specifying objects in the resource would incorrectly restrict the action. Option C is wrong because `s3:ListAllMyBuckets` is only needed for the `aws s3 ls` command without a bucket argument (listing all buckets), not for listing objects in a specific bucket.

107
MCQhard

A developer is troubleshooting performance issues in an application that uses Amazon ElastiCache for Redis. The application experiences periodic latency spikes during peak hours. The developer checks CloudWatch metrics and sees that the 'Evictions' metric is consistently high and the 'CacheHitRate' metric is low. The cluster uses a single cache.t3.small node. Which action will most likely improve the cache hit rate and reduce latency?

A.Increase the number of replicas
B.Enable cluster mode and add more shards
C.Increase the TTL of cached items
D.Use a larger instance type
AnswerD

Upgrading to a larger instance type directly addresses memory constraints by providing a substantial increase in the node's available RAM. This vertical scaling approach immediately expands the cache's capacity, allowing more items to be stored without triggering eviction policies. It is a straightforward, non-disruptive operational change for ElastiCache, often involving a brief failover, and effectively resolves memory-related performance issues without requiring any application code modifications.

Why this answer

The symptoms—high evictions and low cache hit rate—indicate that the single cache.t3.small node is running out of memory. Using a larger instance type increases the available memory, allowing more data to be cached, reducing evictions, and improving the cache hit rate. This directly addresses the root cause of memory pressure without changing the cluster architecture or data expiration behavior.

Exam trap

The trap here is that candidates often confuse scaling out (adding replicas or shards) with scaling up (increasing instance size), but for a single-node cluster suffering from memory exhaustion, the most direct and effective solution is to increase memory capacity, not to add replicas or change the cluster mode.

How to eliminate wrong answers

Option A is wrong because increasing the number of replicas does not increase the total memory capacity of the cluster; replicas are read-only copies that improve read scalability and fault tolerance, but they share the same memory limit as the primary node, so evictions and cache hit rate remain unchanged. Option B is wrong because enabling cluster mode and adding more shards distributes data across multiple nodes, which can increase total memory, but it requires application changes to support sharding and is more complex than simply scaling up the instance size; the immediate, simplest fix for a single-node cluster under memory pressure is to increase memory. Option C is wrong because increasing the TTL of cached items only delays their expiration, but if the cache is already full and evicting items due to memory pressure, longer TTLs will not prevent evictions—they may even worsen the problem by keeping stale data in memory longer.

108
MCQmedium

A company's application uses Amazon DynamoDB as its database. The application reads the same item multiple times per second and occasionally sees stale data. The DynamoDB table uses the default eventually consistent reads. What should the developer change to ensure strongly consistent reads?

A.Increase the read capacity units of the table.
B.Use DynamoDB Accelerator (DAX) to cache the item.
C.Set the ConsistentRead parameter to true in the GetItem call.
D.Use DynamoDB transactions for all read operations.
AnswerC

Setting the ConsistentRead parameter to true in a GetItem API call explicitly instructs DynamoDB to perform a strongly consistent read. This ensures that the data returned reflects all successful write operations that completed before the read request was initiated, providing the most up-to-date version of the item. While this guarantees data freshness, it may incur slightly higher latency and consume more Read Capacity Units compared to an eventually consistent read.

Why this answer

DynamoDB's default read consistency model is eventually consistent, which can return stale data if an item is updated shortly before the read. By setting the `ConsistentRead` parameter to `true` in the `GetItem` call, the developer forces a strongly consistent read, ensuring the response reflects the most recent write. This directly addresses the stale data issue without changing throughput or adding caching.

Exam trap

The trap here is that candidates often confuse throughput scaling (Option A) or caching (Option B) with consistency guarantees, or mistakenly think transactions (Option D) are required for strong consistency, when in fact a simple parameter change on the read operation is the correct and minimal fix.

How to eliminate wrong answers

Option A is wrong because increasing read capacity units (RCUs) only affects throughput and cost, not the consistency model; eventually consistent reads still return stale data regardless of RCU count. Option B is wrong because DynamoDB Accelerator (DAX) is an in-memory cache that improves read performance but does not guarantee strong consistency; it can serve stale data from its cache. Option D is wrong because DynamoDB transactions are designed for atomic, isolated multi-item operations (using `TransactGetItems` or `TransactWriteItems`), not for ensuring single-item strong consistency; using transactions for simple reads adds unnecessary overhead and cost.

109
Multi-Selecthard

A developer is troubleshooting an EC2 instance that is unreachable via SSH. The instance is in a public subnet with a security group that allows inbound SSH from 0.0.0.0/0. Which THREE are possible causes? (Choose 3.)

Select 3 answers
A.The network ACL associated with the subnet is blocking inbound SSH.
B.The SSH key pair used to launch the instance is incorrect.
C.The instance is in the 'stopped' state.
D.The instance does not have an IAM role with the necessary permissions.
E.The instance does not have a public IPv4 address.
AnswersA, C, E

Network ACLs are stateless and operate at the subnet level, so even if the security group allows inbound SSH (TCP 22) from 0.0.0.0/0, a subnet’s NACL with an explicit deny rule for inbound port 22 will block the traffic before it reaches the instance. This satisfies the stem’s constraint that the instance is unreachable despite permissive security group rules.

Why this answer

Option A is correct because a network ACL is a stateless subnet-level firewall, so even if the security group allows inbound TCP port 22 from 0.0.0.0/0, an NACL rule denying inbound SSH (and the corresponding outbound ephemeral ports) will block the connection. Option C is correct because an EC2 instance in the 'stopped' state has no running OS or network stack, so SSH cannot be served regardless of subnet, security group, or IP configuration. Option E is correct because an instance in a public subnet is only reachable from the internet if it has a public IPv4 address (or an Elastic IP); without one, external SSH clients have no routable destination.

Option B is not a valid cause here because an incorrect SSH key pair would cause authentication failure after the TCP connection is established, not make the instance unreachable. Option D is not a valid cause because IAM roles govern AWS API permissions, not inbound SSH access to the instance's operating system.

Exam trap

DVA-C02 often tests the difference between security groups (stateful) and NACLs (stateless), and the requirement for a public IP for internet reachability — candidates may overlook the NACL or public IP and instead blame IAM roles or key pairs, which do not affect network connectivity.

110
Multi-Selecteasy

Which TWO services can be used to store and retrieve application configuration data in AWS? (Choose 2)

Select 2 answers
A.AWS CloudTrail
B.Amazon Simple Queue Service (SQS)
C.AWS Systems Manager Parameter Store
D.Amazon DynamoDB
E.AWS AppConfig
AnswersC, E

AWS Systems Manager Parameter Store provides secure, hierarchical storage for configuration data management and secrets management. It allows you to store data such as passwords, database connection strings, and application settings as parameter values, which can be encrypted. Applications can then programmatically retrieve these values at runtime, making it an ideal, centralized service for managing and distributing application configuration parameters across different environments and services.

Why this answer

AWS Systems Manager Parameter Store (Option C) is a managed service specifically designed to store and retrieve application configuration data, such as database connection strings, passwords, and license keys. It integrates with AWS KMS for encryption and supports hierarchical parameter paths, making it ideal for configuration management without custom code.

Exam trap

The trap here is that candidates often select DynamoDB (Option D) because it can store key-value data, but the question asks for services 'used to store and retrieve application configuration data'—DynamoDB is a general-purpose database, not a dedicated configuration service, and AWS offers purpose-built services (Parameter Store and AppConfig) that are the correct answers.

111
Multi-Selectmedium

Users receive AccessDenied when downloading SSE-KMS encrypted S3 objects cross-account. Which two policies may need changes?

Select 2 answers
A.CloudFront cache policy
B.S3 bucket/object access policy or IAM policy
C.KMS key policy allowing decrypt to the caller
D.Route 53 resolver rule policy
AnswersB, C

Cross-account SSE-KMS downloads require the caller's IAM policy to permit s3:GetObject and the bucket policy to grant access to that principal. Both identity-based and resource-based permissions must align, so these two policies may need changes.

Why this answer

When accessing SSE-KMS encrypted S3 objects cross-account, the S3 bucket policy or the IAM policy must explicitly grant the s3:GetObject permission to the caller. Additionally, the KMS key policy must allow the kms:Decrypt action for the caller's AWS account or IAM role, because SSE-KMS uses a customer master key (CMK) to encrypt the object, and decryption requires KMS permissions. Without both policies, the request fails with AccessDenied even if the S3 permissions are correct.

Exam trap

The trap here is that candidates often assume only the S3 bucket policy needs updating, forgetting that SSE-KMS adds a second authorization layer via KMS key policies, which must explicitly allow the decrypt action for the cross-account caller.

112
MCQmedium

A developer is using AWS CodeBuild to build a Java application. The build fails with 'OutOfMemoryError'. Which configuration change would most likely resolve this issue?

A.Use a custom AMI with more memory.
B.Enable detailed CloudWatch logs for the build.
C.Increase the compute type to a larger instance size.
D.Split the build into multiple CodeBuild projects.
AnswerC

Increasing the compute type is the direct and appropriate method to provide more memory to an AWS CodeBuild project. CodeBuild offers various compute types, such as BUILD_GENERAL1_SMALL, MEDIUM, LARGE, and 2XLARGE, which correspond to different underlying EC2 instance sizes with varying CPU and memory allocations. Selecting a larger compute type directly scales up the resources, including RAM, available to the build container, thereby resolving memory-related build failures.

Why this answer

CodeBuild allows configuring the compute type (e.g., BUILD_GENERAL1_MEDIUM, BUILD_GENERAL1_LARGE) which determines the memory and CPU available. An 'OutOfMemoryError' indicates insufficient memory, so increasing the compute type to a larger instance size provides more memory and resolves the issue. Option A is incorrect because CodeBuild does not support custom AMIs; it uses managed container images.

Option B is incorrect because enabling detailed CloudWatch logs only increases logging, not memory. Option D is incorrect because splitting the build adds complexity and does not address the memory shortage in the current build.

113
MCQmedium

A DynamoDB application receives ProvisionedThroughputExceededException during predictable daily peaks. The workload is not cacheable. What should be changed?

A.Enable S3 Transfer Acceleration
B.Use on-demand capacity or configure autoscaling/scheduled scaling for the table
C.Disable CloudWatch metrics
D.Move all reads to strongly consistent mode
AnswerB

Utilizing DynamoDB's on-demand capacity mode automatically scales read and write throughput to accommodate varying workloads without requiring capacity planning. Alternatively, configuring autoscaling for the table dynamically adjusts provisioned read and write capacity units (RCUs/WCUs) based on actual traffic patterns or target utilization, preventing throttling errors during peak loads. Scheduled scaling further allows for pre-planned capacity adjustments for predictable traffic spikes, ensuring the application always has sufficient resources.

Why this answer

The ProvisionedThroughputExceededException indicates that the table's read/write capacity is insufficient during peak loads. Since the workload is predictable but not cacheable, the correct solution is to either switch to on-demand capacity mode, which automatically scales to handle any traffic level, or configure auto scaling with scheduled scaling to match the predictable peaks. This directly addresses the capacity shortfall without requiring application changes.

Exam trap

The trap here is that candidates may think disabling CloudWatch metrics reduces overhead or that strongly consistent reads improve reliability, but both actions either remove monitoring or increase capacity consumption, making the throttling worse.

How to eliminate wrong answers

Option A is wrong because S3 Transfer Acceleration is a feature for speeding up uploads to S3 over long distances, not for DynamoDB throughput issues. Option C is wrong because disabling CloudWatch metrics would remove visibility into table performance and prevent monitoring of throttling events, making troubleshooting harder. Option D is wrong because strongly consistent reads consume more read capacity units than eventually consistent reads, which would worsen the throughput problem instead of solving it.

114
MCQmedium

A company runs a web application on EC2 instances in an Auto Scaling group. The application needs to store session state. The architecture must be highly available and scalable. Which solution should the developer choose?

A.Use sticky sessions on the Application Load Balancer
B.Store session data in an S3 bucket
C.Use Amazon ElastiCache for Redis to store session state
D.Store session data in the instance's ephemeral storage
AnswerC

Amazon ElastiCache for Redis provides a highly performant, scalable, and fully managed in-memory data store perfectly suited for externalizing session state. Its extremely low-latency read and write capabilities ensure a responsive user experience, while its built-in replication and automatic failover mechanisms guarantee high availability and durability of session data. By storing sessions externally, any EC2 instance in an Auto Scaling group can access any user's session, enabling seamless horizontal scaling and robust fault tolerance without session loss.

Why this answer

Amazon ElastiCache for Redis is a fully managed, highly available, and scalable in-memory data store that is purpose-built for session state management. It supports replication, automatic failover, and horizontal scaling, making it the ideal choice for a highly available and scalable session store across an Auto Scaling group of EC2 instances.

Exam trap

DVA-C02 often tests the difference between sticky sessions (which are not a session store) and a centralized session store like ElastiCache, causing candidates to pick sticky sessions as a scalability solution.

How to eliminate wrong answers

Option A is wrong because sticky sessions on an ALB only pin a user to a single instance; if that instance fails or is replaced by Auto Scaling, the session is lost, and it does not provide a shared, scalable session store. Option B is wrong because S3 is object storage with high latency and is not designed for the low-latency, high-throughput read/write access pattern required for session state. Option D is wrong because ephemeral instance storage is tied to a single EC2 instance and is lost on stop/terminate, making it unsuitable for highly available, scalable session management.

115
MCQhard

A developer is using AWS CodeBuild to build a Docker image and push it to Amazon ECR. The build fails with the error 'no basic authentication credentials'. The build project has an IAM role with the AmazonEC2ContainerRegistryPowerUser policy. What is the most likely cause?

A.The build project is not configured to use a VPC that can reach ECR.
B.The build environment does not have Docker installed.
C.The IAM role does not have sufficient permissions to push to ECR.
D.The buildspec does not include the pre_build step to authenticate with ECR.
AnswerD

The 'no basic authentication credentials' error occurs specifically when the Docker client attempts to push to ECR without first authenticating; the buildspec must include a pre_build phase command that runs 'aws ecr get-login-password | docker login' to obtain a temporary token, and omitting this step is the classic root cause of this exact error message.

Why this answer

AWS CodeBuild does not automatically authenticate to Amazon ECR. The buildspec must include a pre_build phase that runs 'aws ecr get-login-password' piped to 'docker login' (or uses the ECR credential helper) to obtain temporary credentials. The IAM role's AmazonEC2ContainerRegistryPowerUser policy grants the necessary permissions, but the Docker client still needs explicit authentication before pushing.

Exam trap

DVA-C02 often tests the misconception that granting an IAM policy like AmazonEC2ContainerRegistryPowerUser is sufficient for Docker to push to ECR, when in fact the buildspec must explicitly authenticate the Docker client.

How to eliminate wrong answers

Option A is wrong because CodeBuild runs in an AWS-managed environment with internet access by default, and ECR is reachable via public endpoints; VPC configuration is not required for ECR access unless private endpoints are mandated. Option B is wrong because the error is specifically about authentication credentials, not a missing Docker binary — CodeBuild's standard images include Docker, and a missing Docker install would produce a 'command not found' error. Option C is wrong because AmazonEC2ContainerRegistryPowerUser provides full push/pull permissions to ECR; the error 'no basic authentication credentials' indicates the Docker client was never authenticated, not that the IAM role lacks permissions.

116
MCQmedium

A developer is using AWS CloudFormation to deploy a stack that includes an S3 bucket and a Lambda function. The stack fails with the error 'The following resource(s) failed to create: [MyBucket]'. What is the most likely cause?

A.The S3 bucket name is already taken.
B.The stack's VPC configuration is incorrect.
C.The S3 bucket policy is malformed.
D.The Lambda function code is invalid.
AnswerA

S3 bucket names must be globally unique across all AWS accounts and regions; if the specified BucketName property is already claimed by another account, CloudFormation's CreateBucket call fails immediately with a naming conflict, which is the most common cause of this specific error.

Why this answer

The correct answer is A: the S3 bucket name is already taken. S3 bucket names must be globally unique across all AWS accounts and regions, so if the requested name already exists, CloudFormation fails to create the MyBucket resource with exactly this kind of 'failed to create' error. The other options do not fit: a VPC configuration issue would typically affect resources like Lambda or EC2 networking, not S3 bucket creation; a malformed bucket policy would usually fail during policy attachment or update rather than initial bucket creation; and invalid Lambda code would cause the Lambda resource to fail, not MyBucket.

117
MCQhard

A developer notices that an Amazon RDS for MySQL DB instance's CPU utilization is consistently above 90% during peak hours. The application uses read-heavy workloads. Which action would MOST effectively reduce CPU load without major architectural changes?

A.Implement an in-memory cache layer with Amazon ElastiCache.
B.Migrate the database to Amazon Aurora with auto-scaling.
C.Increase the DB instance size to a larger instance type.
D.Create a Multi-AZ deployment and use the standby for read queries.
AnswerA

Implementing an in-memory cache layer with Amazon ElastiCache (e.g., Redis or Memcached) is an effective strategy for reducing CPU load on a read-heavy database. By caching frequently accessed data, ElastiCache intercepts read requests before they reach the database, serving them much faster from memory. This significantly decreases the number of queries the RDS instance needs to process, directly lowering its CPU utilization and improving overall application responsiveness. It requires application-level changes to interact with the cache, but these are typically manageable.

Why this answer

For a read-heavy workload causing sustained high CPU on RDS MySQL, offloading repeated read queries to an in-memory cache like ElastiCache (Redis or Memcached) removes the majority of read traffic from the database engine, directly reducing CPU utilization without re-architecting the application. This is the least invasive and most effective fix for read-dominated load.

Exam trap

The trap here is confusing RDS Multi-AZ standby with a read replica — many candidates believe the standby can serve read traffic, but it is a passive failover node, making option D a classic distractor.

How to eliminate wrong answers

Option B is wrong because migrating to Aurora with auto-scaling is a major architectural change (data migration, endpoint changes, cost implications) and auto-scaling in Aurora primarily targets read replicas, not the writer's CPU — it does not satisfy the 'without major architectural changes' constraint. Option C is wrong because vertically scaling the instance only buys temporary headroom; it does not address the root cause of repeated identical reads and is more expensive than caching. Option D is wrong because a Multi-AZ standby in RDS is a passive failover target and cannot serve read queries — that is a common misconception; only read replicas serve reads.

118
MCQeasy

A developer is deploying a new version of a Lambda function and wants to roll back immediately if errors are detected. Which deployment strategy should the developer use?

A.Use AWS CodeDeploy with a canary deployment configuration
B.Use an EC2 rolling update strategy
C.Use AWS CodeDeploy with a linear deployment configuration
D.Use an immutable update strategy
AnswerA

Using AWS CodeDeploy with a canary deployment configuration is the optimal strategy for safely deploying new Lambda function versions. This approach shifts a small, configurable percentage of traffic (e.g., 10%) to the new version for a specified 'bake time,' allowing real-world monitoring via CloudWatch alarms. If errors or performance issues are detected during this period, CodeDeploy automatically rolls back to the previous stable version, minimizing impact and ensuring application stability.

Why this answer

The correct option is A: AWS CodeDeploy with a canary deployment configuration. CodeDeploy supports Lambda deployments with traffic-shifting configurations, and a canary deployment shifts a small percentage of traffic first (e.g., 10% for 5 minutes) so that CloudWatch alarms can trigger an automatic rollback if errors are detected, satisfying the requirement to roll back immediately. Option B is wrong because EC2 rolling updates apply to EC2 instances, not Lambda functions.

Option C is less suitable because a linear configuration shifts traffic in equal increments over time, which delays error detection compared to a canary's initial small traffic slice. Option D is wrong because immutable updates are an EC2/Elastic Beanstalk strategy, not a Lambda deployment strategy.

119
Multi-Selecthard

A developer is using AWS X-Ray to trace a Lambda function that calls DynamoDB and SQS. Some traces show errors. Which TWO actions should the developer take to diagnose the issue?

Select 2 answers
A.Examine the trace details for exception messages.
B.Verify that the Lambda function's IAM role has permissions for X-Ray.
C.Check the X-Ray service map for error edges.
D.Disable X-Ray sampling to capture all requests.
E.Enable CloudFront to cache responses.
AnswersA, C

Examining X-Ray trace details is the most direct and effective method to diagnose errors within a Lambda function. Each trace provides granular information for segments and subsegments, including full stack traces, precise exception messages, error codes, and specific HTTP status codes for downstream calls. This allows a developer to pinpoint the exact failure point, whether it's within the Lambda's code logic or an issue with an external service call, such as a DynamoDB throttling exception or a permission denied error.

Why this answer

To diagnose errors in existing X-Ray traces, the developer should examine trace details (Option A) to see exception messages and stack traces for each segment, and check the service map (Option C) for error edges that indicate which service interactions failed. Option B is about enabling X-Ray permissions, which is a prerequisite for tracing but does not help diagnose errors in traces that are already captured. Option D (disabling sampling) is unnecessary because X-Ray captures errors by default regardless of sampling.

Option E (CloudFront caching) is unrelated to trace diagnostics.

Exam trap

A common trap is selecting Option B, thinking that ensuring X-Ray permissions is a diagnostic step. However, missing permissions would prevent traces from being sent at all; since traces are present, the focus should be on analyzing the existing trace data (details and service map) to find error causes.

120
Drag & Dropmedium

Drag and drop the steps to authenticate a user using Amazon Cognito User Pools in the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

First create the user pool and app client, then authenticate to receive tokens, and use tokens for authorization.

121
MCQeasy

A developer is troubleshooting an AWS CloudFormation stack creation failure. The stack creation failed with the error: 'Resource creation cancelled'. What does this error typically indicate?

A.The IAM user does not have permission to create the resource.
B.Another resource in the stack failed, causing a rollback.
C.The template has a syntax error.
D.The resource type is not supported by CloudFormation.
AnswerB

When a CloudFormation stack creates multiple resources, these operations often occur in parallel or a defined sequence. If one resource fails to create or update successfully, CloudFormation's default behavior is to initiate a rollback of the entire stack to its previous stable state. During this rollback, any resources that were in the process of being created or updated, but had not yet failed themselves, will have their operations terminated, resulting in a "Resource creation cancelled" status. This indicates an orchestrated termination rather than an individual resource failure.

Why this answer

'Resource creation cancelled' is CloudFormation's generic message when a resource's creation is aborted because another resource in the same stack failed and CloudFormation initiated a rollback. By default, stack creation uses rollback-on-failure, so once any resource fails, all in-progress creations are cancelled and successfully created resources are deleted. The error is a symptom of a sibling resource failure, not a problem with the cancelled resource itself.

Exam trap

DVA-C02 often tests whether candidates recognize that 'Resource creation cancelled' is a downstream symptom of rollback, not a root-cause error — candidates incorrectly blame IAM or syntax instead of finding the first resource that actually failed.

How to eliminate wrong answers

Option A is wrong because missing IAM permissions produce an 'Access Denied' or 'not authorized to perform' error on the specific resource, not 'Resource creation cancelled'. Option C is wrong because template syntax errors are caught during validation and return 'Template format error' or 'YAML not well-formed' before any resource provisioning begins. Option D is wrong because an unsupported resource type yields 'Resource type X is not supported' or 'Invalid resource type' during template validation, not a cancellation message.

122
MCQhard

A developer is troubleshooting an Amazon API Gateway REST API that returns 504 Gateway Timeout errors for certain requests. The backend is a Lambda function that performs a resource-intensive operation that occasionally takes up to 30 seconds. API Gateway has a default integration timeout of 29 seconds. The developer cannot reduce the execution time. What should the developer do to resolve the timeout issue?

A.Increase the API Gateway integration timeout to 30 seconds.
B.Refactor the Lambda function to use asynchronous invocation, return a 202 immediately, and have the client poll for results.
C.Enable API Gateway caching to avoid repeated calls.
D.Use multiple Lambda functions to parallelize processing.
AnswerB

Refactoring the Lambda function to use asynchronous invocation, returning a 202 immediately, and having the client poll for results is the correct approach for long-running operations. This pattern decouples the synchronous API Gateway request from the extended backend processing, allowing API Gateway to respond promptly with a 202 Accepted status. The Lambda function can then trigger an asynchronous workflow (e.g., via SQS, SNS, or directly invoking another Lambda asynchronously) and store results for the client to retrieve later through a separate polling mechanism, effectively bypassing the 29-second timeout.

Why this answer

It decouples the client from the long-running Lambda execution. By invoking the Lambda asynchronously, the API Gateway can return a 202 Accepted response immediately, well within the 29-second integration timeout. The client then polls a separate endpoint (e.g., using a presigned S3 URL or a DynamoDB status record) to retrieve the final result, completely sidestepping the timeout limitation.

Exam trap

The trap here is that candidates assume the integration timeout is configurable to any value, but AWS enforces a hard 29-second limit for REST APIs, making Option A technically impossible.

How to eliminate wrong answers

Option A is wrong because Amazon API Gateway has a hard maximum integration timeout of 29 seconds for REST APIs (and 30 seconds for HTTP APIs). You cannot increase it beyond that limit, so setting it to 30 seconds is not possible. Option C is wrong because caching only serves previously computed responses for identical requests; it does not reduce the execution time of a new, uncached request that still takes up to 30 seconds.

Option D is wrong because parallelizing the Lambda function does not reduce the total execution time of a single resource-intensive operation; the request still waits for all parallel tasks to complete, which can still exceed the 29-second timeout.

123
MCQmedium

A developer receives an AccessDenied error when trying to put an object into an S3 bucket using the AWS SDK. The IAM user has an attached policy that grants s3:PutObject on the bucket. What is the MOST likely cause of the error?

A.The request is being throttled by S3.
B.The object key is too long.
C.The AWS SDK version is outdated.
D.The bucket policy explicitly denies the action.
AnswerD

When evaluating permissions, AWS IAM follows a strict order of precedence where an explicit Deny statement always overrides any Allow statement. If a bucket policy contains an explicit Deny for a specific action or principal, that denial takes precedence over any Allow statement present in the requesting IAM user's or role's identity-based policy. This ensures that even if an identity policy grants permission, a resource-based policy can still block access, resulting in an AccessDenied error (HTTP 403).

Why this answer

The most likely cause is that the bucket policy explicitly denies the s3:PutObject action. IAM policies grant permissions, but S3 bucket policies can override them with an explicit deny, which takes precedence over any allow. Since the IAM user already has an attached policy allowing s3:PutObject, the only way to get an AccessDenied error is if a bucket policy explicitly denies the action.

Exam trap

The trap here is that candidates assume an IAM allow is sufficient, forgetting that S3 bucket policies can explicitly deny actions, and that explicit deny always wins over allow.

How to eliminate wrong answers

Option A is wrong because S3 throttling returns a 503 SlowDown error, not an AccessDenied error. Option B is wrong because an overly long object key would cause a 400 Bad Request error, not an AccessDenied error. Option C is wrong because an outdated SDK version might cause compatibility issues or missing features, but it would not result in an AccessDenied error; the error is a permissions issue, not a client version issue.

124
MCQmedium

A developer monitors an AWS Lambda function that processes records from an Amazon SQS queue and writes results to an Amazon DynamoDB table. CloudWatch Logs show that execution time has increased over the past week, and the function frequently times out at the 5-minute timeout. The function's code has not been changed recently. CloudWatch metrics show a high rate of DynamoDBProvisionedThroughputExceededException errors. The DynamoDB table has 5 write capacity units (WCUs). What action will MOST effectively reduce the function's execution time?

A.Increase the Lambda function's timeout to 10 minutes.
B.Increase the write capacity units (WCUs) on the DynamoDB table.
C.Increase the Lambda function's memory allocation to 3008 MB.
D.Use an Amazon SQS FIFO queue instead of a standard queue for the Lambda trigger.
AnswerB

The `DynamoDBProvisionedThroughputExceededException` directly signifies that the DynamoDB table's allocated write capacity units (WCUs) are insufficient to handle the incoming write requests. Increasing the WCUs directly addresses this bottleneck by provisioning more throughput for the table. This action allows DynamoDB to process more writes per second, eliminating throttling, reducing Lambda retries, and consequently speeding up the Lambda function's overall execution time.

Why this answer

The high rate of DynamoDBProvisionedThroughputExceededException errors indicates that the Lambda function is being throttled by DynamoDB due to insufficient write capacity. When writes are throttled, the Lambda function must retry, which increases execution time and can lead to timeouts. Increasing the WCUs on the DynamoDB table directly addresses the root cause by allowing the function to write without throttling, thereby reducing execution time.

Exam trap

The trap here is that candidates often assume increasing Lambda timeout or memory will fix performance issues, but the real bottleneck is the DynamoDB write capacity, which directly causes the throttling errors and increased execution time.

How to eliminate wrong answers

Option A is wrong because increasing the timeout to 10 minutes does not resolve the underlying throttling issue; it only masks the symptom by allowing the function to run longer while still being throttled. Option C is wrong because increasing memory allocation (up to 3008 MB) primarily improves CPU performance and network throughput, but does not fix DynamoDB throttling caused by insufficient WCUs. Option D is wrong because switching to an SQS FIFO queue does not affect DynamoDB write capacity; FIFO queues enforce message ordering and deduplication but do not reduce the throttling rate from DynamoDB.

125
MCQeasy

A developer is debugging an AWS Lambda function that is invoked by an Amazon S3 bucket notification. The function fails with an 'AccessDenied' error when trying to read an object from the same bucket. What should the developer check first?

A.Check if S3 bucket versioning is enabled.
B.Verify that the S3 bucket uses server-side encryption with AWS KMS.
C.Ensure the S3 bucket is not blocked by S3 Block Public Access.
D.Review the Lambda function's execution role for s3:GetObject permission.
AnswerD

An 'AccessDenied' error when an AWS Lambda function attempts to interact with Amazon S3 is most commonly caused by insufficient permissions defined in its IAM execution role. To successfully retrieve an object from S3, the Lambda function's execution role must have an IAM policy that explicitly grants the `s3:GetObject` action on the specific target S3 bucket and object path. This is the foundational security control for S3 object retrieval.

Why this answer

The developer should first review the Lambda function's execution role for s3:GetObject permission. The 'AccessDenied' error indicates the function lacks permission to read the object. The execution role must have an IAM policy granting s3:GetObject on the bucket or object.

This is the most direct cause.

Exam trap

The trap is overthinking encryption or public access blocks, but the most common cause of AccessDenied in Lambda is missing IAM permissions in the execution role.

How to eliminate wrong answers

Option A is wrong because S3 bucket versioning does not affect permissions; it only manages object versions. Option B is wrong because while KMS encryption can cause AccessDenied if the role lacks kms:Decrypt, the question says the error is when trying to read the object, and the most common cause is missing s3:GetObject; KMS is a secondary check. Option C is wrong because S3 Block Public Access prevents public access, but Lambda uses an IAM role, not public access, so it is not relevant.

126
MCQeasy

A developer is troubleshooting an Amazon RDS for MySQL instance that is experiencing high CPU utilization. The application performs many read operations. The developer wants to reduce the load on the database. What is the MOST effective solution?

A.Upgrade the DB instance to a larger instance class.
B.Create a read replica and direct read queries to it.
C.Enable Multi-AZ for automatic failover.
D.Purchase reserved instances to reduce costs.
AnswerB

Creating an Amazon RDS read replica asynchronously replicates data from the primary DB instance, allowing read queries to be directed to the replica. This effectively offloads read traffic from the primary instance, significantly reducing its CPU utilization and improving overall database performance. Read replicas are specifically designed to scale read-heavy workloads independently, providing a highly efficient and cost-effective solution.

Why this answer

Creating a read replica allows read queries to be directed to the replica, offloading the read workload from the primary instance and reducing its CPU utilization. Option A is incorrect because upgrading to a larger instance class increases capacity but does not specifically address read-heavy workloads efficiently and is less cost-effective than scaling reads with replicas. Option C is incorrect because Multi-AZ provides high availability and automatic failover, not performance improvement or load reduction.

Option D is incorrect because purchasing reserved instances reduces costs but does not affect CPU utilization.

127
MCQmedium

A developer is troubleshooting a slow-running Amazon RDS for MySQL query. The query performance has degraded over time. Which approach should the developer take first to identify the cause?

A.Enable Performance Insights and review the database load
B.Upgrade the DB instance to a larger instance class
C.Create a read replica to offload read traffic
D.Enable the MySQL query cache
AnswerA

Enabling Amazon RDS Performance Insights is the most effective initial step for diagnosing slow database performance. It provides a visual dashboard of database load, breaking down wait events, active sessions, and top SQL queries over time. This granular visibility allows developers to pinpoint specific bottlenecks, such as I/O waits, CPU contention, or inefficient query execution plans, before implementing any corrective actions.

Why this answer

Amazon RDS Performance Insights is the first-line diagnostic tool for identifying why a query has degraded — it visualizes database load (DBLoad) sliced by SQL statement, wait event, user, and host, showing exactly which query and which wait state is consuming resources. Enabling it is non-invasive and provides the evidence needed before making any infrastructure change. The other options are remediation steps that should only follow diagnosis.

Exam trap

DVA-C02 often tests the 'diagnose before remediate' principle — candidates are tempted to pick scaling or caching actions, but the question asks for the FIRST step, which is always observability (Performance Insights, slow query log, EXPLAIN).

How to eliminate wrong answers

Option B is wrong because upgrading the instance class is a costly remediation that may not address the root cause — if the bottleneck is a missing index or lock contention, a larger instance only masks the problem temporarily. Option C is wrong because a read replica offloads read traffic but does not diagnose a slow query; if the query is a write or the replica lags, it does not help, and it does not identify the cause. Option D is wrong because the MySQL query cache was deprecated in MySQL 5.7 and removed in 8.0, and even when available it is invalidated by any write to the table, making it ineffective for write-heavy workloads and irrelevant as a diagnostic step.

128
Multi-Selectmedium

A developer is troubleshooting a slow-running query on an Amazon RDS for MySQL database. The query is used by a reporting application and takes over 30 seconds to complete. The database is a db.r5.large instance with 200 GB of gp2 storage. Which TWO actions should the developer take to improve query performance?

Select 2 answers
A.Terminate idle connections to free up resources.
B.Review the slow query log to identify the query and its execution plan.
C.Increase the allocated storage to 500 GB to improve I/O performance.
D.Add appropriate indexes to the tables involved in the query.
E.Enable Multi-AZ deployment for better read performance.
AnswersB, D

Reviewing the slow query log is the first diagnostic step because it captures queries that exceed a specified duration, along with their execution time and connection metadata. Once identified, use EXPLAIN to analyze the execution plan, exposing table scans, missing indexes, or poor join ordering. This evidence-based approach tells you exactly which query to optimize and whether to add indexes or rewrite the query.

Why this answer

Option B is correct because the MySQL slow query log captures queries exceeding the long_query_time threshold (default 10 seconds), and reviewing it along with EXPLAIN output reveals the query's execution plan, helping pinpoint full table scans, missing indexes, or inefficient joins. Option D is correct because adding appropriate indexes on the columns used in WHERE, JOIN, and ORDER BY clauses lets MySQL satisfy the query with index lookups instead of full table scans, which is the most direct fix for a slow reporting query. Option A is not appropriate because idle connections consume minimal resources and terminating them does not address query execution inefficiency.

Option C is not appropriate because increasing gp2 storage size only raises the baseline IOPS (3 IOPS/GB) and burst balance; it does not fix a poorly optimized query and is a costly workaround. Option E is not appropriate because Multi-AZ is a high-availability feature that maintains a standby replica for failover, not a read-scaling mechanism, so it does not improve query performance.

Exam trap

DVA-C02 often tests the misconception that scaling storage or enabling Multi-AZ improves query performance, when the real fix is query-level tuning (indexes, execution plan review).

129
MCQeasy

A developer notices that an S3 bucket policy allows public read access to all objects. The bucket contains sensitive data that should only be accessible by authorized IAM users. What is the BEST way to remediate this?

A.Enable default encryption on the bucket.
B.Modify the bucket policy to remove the public statement and use IAM policies for access.
C.Enable S3 Block Public Access at the account level.
D.Enable S3 Object Ownership and use ACLs.
AnswerB

The most direct and secure solution is to modify the S3 bucket policy to remove any statements that grant public access, typically identified by "Principal: "*". Concurrently, implement specific IAM policies attached to users, groups, or roles to grant precise, least-privilege access to authorized principals. This approach directly addresses the misconfiguration, ensures granular control, and aligns with AWS security best practices for managing access to S3 resources.

Why this answer

The bucket policy currently grants public read access, which overrides any IAM-based restrictions. By removing the public statement from the bucket policy and relying solely on IAM policies, access is controlled at the user level, ensuring only authorized IAM users can read objects. This aligns with the principle of least privilege and follows AWS best practices for securing S3 data.

Exam trap

The trap here is that candidates often confuse encryption with access control, thinking that enabling encryption (Option A) will prevent unauthorized access, when in fact encryption only protects data at rest and does not affect public read permissions.

How to eliminate wrong answers

Option A is wrong because enabling default encryption only encrypts data at rest; it does not restrict access, so public read access would still be allowed. Option C is wrong because S3 Block Public Access at the account level would prevent all public access, but it is a broad, account-wide setting that may inadvertently block legitimate public access for other buckets; the question asks for the best remediation for this specific bucket, not a blanket account-level change. Option D is wrong because S3 Object Ownership and ACLs are legacy access control mechanisms that are less secure and more complex to manage than IAM policies, and they do not directly address the public read access granted by the bucket policy.

130
MCQhard

A company uses AWS CloudFormation to deploy a stack that includes an RDS MySQL instance. During an update, the stack fails with a 'DELETE_FAILED' status on a security group resource. The security group has a dependency on the RDS instance. What is the MOST likely cause?

A.The RDS instance is not fully deleted because of a deletion protection flag.
B.The security group has a rule that references itself.
C.The security group must be deleted manually before updating the stack.
D.The security group is attached to an EC2 instance outside the stack.
AnswerA

If the RDS instance has the DeletionProtection attribute set to true, CloudFormation's attempt to delete or replace that instance during the stack update silently fails, leaving the instance running and still attached to the security group's ENI, which in turn prevents the security group from being deleted since AWS will not remove a security group that is still in use by an active resource.

Why this answer

The most likely cause is that the RDS instance has deletion protection enabled, preventing it from being deleted even when CloudFormation attempts to delete it. The security group depends on the RDS instance, so if the RDS instance cannot be deleted, the security group also fails to delete, resulting in a DELETE_FAILED status. Option A correctly identifies this.

Option B is incorrect because a self-referencing rule would not cause a delete failure. Option C is incorrect because manual deletion is not required; the issue is with the RDS instance. Option D is incorrect because the security group being attached to an external EC2 instance would cause a different error, not a dependency-related failure.

131
MCQeasy

A developer is using AWS CloudFront to serve static content. Users in some geographic regions report slow load times. Which CloudFront feature can the developer use to reduce latency for these users?

A.Change the CloudFront price class to include all edge locations.
B.Create multiple origins in different regions.
C.Enable S3 Transfer Acceleration on the origin S3 bucket.
D.Use Lambda@Edge to optimize content delivery.
AnswerA

CloudFront's price classes determine which edge locations are utilized for content delivery. Selecting "Price Class All" ensures that CloudFront leverages its entire global network of edge locations, including those in regions with higher infrastructure costs. This maximizes the geographic proximity of cached content to end-users worldwide, thereby minimizing latency for static content delivery regardless of the user's physical location.

Why this answer

CloudFront's price class determines which edge locations are used to serve content. By default, CloudFront may use only a subset of edge locations (Price Class 100 or 200) to reduce costs, which can cause higher latency for users in regions not covered. Changing the price class to 'All Edge Locations' (Price Class All) ensures that CloudFront uses every global edge location, providing lower latency for all users.

Option B is incorrect because creating multiple origins in different regions addresses origin distance, not edge location coverage. Option C is incorrect because S3 Transfer Acceleration is for speeding up uploads to S3, not for improving CloudFront content delivery. Option D is incorrect because Lambda@Edge can modify content at edge locations but does not expand the set of edge locations used.

132
MCQeasy

A developer is deploying a new version of a Lambda function using an alias for blue/green deployment. Traffic is gradually shifted to the new version. During the shift, a high error rate is observed. What should the developer do to minimize impact?

A.Use the Lambda function's provisioned concurrency to pre-warm the new version.
B.Manually revert the alias to point back to the old version.
C.Configure the alias with a canary deployment and an error rate alarm for automatic rollback.
D.Delete the new version and redeploy after fixing the issue.
AnswerC

Configuring a Lambda alias with a canary deployment allows for gradual shifting of traffic to the new function version, starting with a small percentage. Integrating this with an Amazon CloudWatch error rate alarm enables automatic rollback: if the new version's error rate exceeds a predefined threshold during the canary phase, the alias automatically reverts all traffic to the stable old version. This strategy minimizes the impact of potential issues by detecting them early and automating recovery.

Why this answer

It automates the rollback process using AWS CodeDeploy's canary deployment with an Amazon CloudWatch alarm on the error rate. When the alarm triggers, CodeDeploy automatically shifts traffic back to the previous version, minimizing impact without manual intervention. This is the recommended approach for safe blue/green deployments with Lambda aliases.

Exam trap

The trap here is that candidates may think manual reversion (Option B) is the simplest fix, but the exam emphasizes automated rollback strategies (like canary deployments with alarms) as the best practice for minimizing impact during blue/green deployments.

How to eliminate wrong answers

Option A is wrong because provisioned concurrency pre-warms execution environments to reduce cold starts, but it does not address a high error rate during traffic shifting; errors are typically caused by code defects, not cold starts. Option B is wrong because manually reverting the alias is a valid fallback but is slower and error-prone compared to an automated rollback; the question asks to minimize impact, and manual reversion introduces delay and potential for human error. Option D is wrong because deleting the new version and redeploying after fixing the issue is a reactive approach that does not minimize impact during the shift; it requires manual intervention and does not provide automatic recovery.

133
MCQmedium

A developer is managing an application running on Amazon EC2 instances behind an Application Load Balancer. Users report that the application becomes unresponsive after several hours, and restarting the instance temporarily fixes the issue. The developer suspects a memory leak but cannot add custom instrumentation. Which AWS service can collect memory utilization metrics and help identify the memory leak with minimal configuration?

A.Use Amazon CloudWatch Logs agent to capture application logs.
B.Use the EC2 instance metadata service to query memory usage.
C.Install the CloudWatch agent on the EC2 instances to collect memory metrics and emit them to CloudWatch.
D.Use AWS X-Ray to trace memory allocation.
AnswerC

The unified CloudWatch agent is the correct and recommended solution for collecting detailed operating system-level metrics, including memory utilization, from EC2 instances. This agent can be configured to gather various custom metrics, such as used memory percentage, free memory, and swap usage, directly from the instance's operating system. These collected metrics are then reliably published to CloudWatch, enabling comprehensive monitoring, alarming, and dashboarding capabilities.

Why this answer

The CloudWatch agent can collect custom metrics, including memory utilization, from EC2 instances and publish them to Amazon CloudWatch. This allows the developer to monitor memory usage over time and identify a memory leak without modifying the application code. The default EC2 metrics do not include memory utilization, so the CloudWatch agent is the minimal-configuration solution for this requirement.

Exam trap

The trap here is that candidates assume EC2 automatically provides memory metrics in CloudWatch, but in reality, only CPU, network, and disk metrics are available by default; memory requires the CloudWatch agent.

How to eliminate wrong answers

Option A is wrong because the CloudWatch Logs agent captures application logs, not memory utilization metrics; logs could indirectly indicate issues but do not provide direct memory metrics needed to identify a leak. Option B is wrong because the EC2 instance metadata service provides information about the instance itself (e.g., instance ID, AMI ID) but does not expose memory utilization data; it is not a monitoring service for OS-level metrics. Option D is wrong because AWS X-Ray traces requests and identifies performance bottlenecks in distributed applications, not memory allocation or utilization; it is designed for tracing, not OS-level resource monitoring.

134
MCQmedium

A developer is troubleshooting a DynamoDB table that is experiencing high write throttling (ProvisionedThroughputExceededException) on certain days. The table has provisioned write capacity of 1000 WCU. The table has a partition key of 'user_id' which is a UUID. The table is accessed by multiple services. CloudWatch metrics show that the WriteThrottleEvents are spiking during specific hours, and the ConsumedWriteCapacityUnits often reaches 1000. What is the most likely cause of the throttling?

A.The partition key is not distributed evenly, causing a hot partition.
B.The provisioned write capacity is insufficient to handle the traffic spikes.
C.The table does not have DynamoDB Accelerator (DAX) enabled.
D.The table is configured with eventual consistency, which throttles writes.
AnswerB

This is the correct answer. DynamoDB tables operate on a provisioned throughput model, where Write Capacity Units (WCUs) must be sufficient to handle the incoming write traffic. When the rate of write requests, especially during traffic spikes, exceeds the allocated provisioned write capacity, DynamoDB will begin to throttle requests. This throttling mechanism protects the underlying infrastructure and ensures consistent performance for other requests within the provisioned limits, but it results in rejected write operations for the application.

Why this answer

The ConsumedWriteCapacityUnits consistently reaches the provisioned 1000 WCU during specific hours, and WriteThrottleEvents spike at those same times. This indicates that the provisioned capacity is insufficient to handle peak traffic, causing requests to be throttled. The partition key (UUID) is well-distributed, so a hot partition is unlikely.

Exam trap

The trap here is that candidates often assume throttling must be caused by a hot partition (Option A) when the partition key is not a UUID, but in this case the UUID ensures even distribution, so the real issue is simply insufficient capacity during traffic spikes.

How to eliminate wrong answers

Option A is wrong because the partition key is a UUID, which is inherently random and evenly distributes writes across partitions, making a hot partition improbable. Option C is wrong because DAX is an in-memory cache for reads, not writes, and does not affect write throttling or provisioned write capacity. Option D is wrong because eventual consistency applies only to reads, not writes; writes are always strongly consistent and throttling is based on write capacity, not consistency settings.

135
MCQmedium

A company is using Amazon API Gateway to expose a REST API. The API is integrated with an AWS Lambda function. Lately, the API is returning 502 Bad Gateway errors. What is the MOST likely cause?

A.The API Gateway request throttling limit has been exceeded.
B.The API Gateway API key is invalid.
C.The Lambda function is returning an unhandled exception.
D.The Lambda function's execution role does not allow API Gateway to invoke it.
AnswerC

API Gateway expects its integrated Lambda function to return a specific JSON response format, including status code, headers, and body, for successful processing and mapping. When a Lambda function encounters an unhandled exception, times out, or returns malformed output that does not conform to this expected structure, API Gateway cannot properly map this response to an HTTP response for the client. Consequently, API Gateway returns an HTTP 502 Bad Gateway error, indicating that it received an invalid response from the upstream Lambda service.

Why this answer

A 502 Bad Gateway error from API Gateway typically indicates that the backend integration (in this case, the Lambda function) returned an error response. When a Lambda function throws an unhandled exception, API Gateway receives a 200 OK with a function error payload, but it cannot parse the response into a valid HTTP response, resulting in a 502. This is distinct from throttling or permission issues, which produce different HTTP status codes.

Exam trap

The trap here is that candidates often confuse 502 errors with throttling (429) or permission issues (403/500), but the 502 specifically points to a malformed or error response from the backend integration.

How to eliminate wrong answers

Option A is wrong because exceeding API Gateway request throttling limits results in a 429 Too Many Requests error, not a 502 Bad Gateway. Option B is wrong because an invalid API key causes a 403 Forbidden error, not a 502. Option D is wrong because if the Lambda function's execution role does not allow API Gateway to invoke it, API Gateway would return a 500 Internal Server Error or a 403, not a 502.

136
MCQmedium

A developer notices that an AWS Lambda function, configured to access an Amazon RDS database in the same VPC, is timing out. The function has a 30-second timeout. CloudWatch Logs show that the function starts execution but never reaches the database. The VPC configuration includes private subnets without a NAT gateway. The RDS database is in the same VPC. What is the most likely cause of the timeout?

A.The Lambda function does not have internet access because it is in a VPC without a public IP.
B.The security group of the RDS database does not allow inbound traffic from the Lambda function's security group.
C.The Amazon RDS database is not publicly accessible and the Lambda function cannot resolve the database endpoint.
D.The VPC does not have a VPC endpoint for Amazon RDS, and the Lambda function cannot access the database through the NAT gateway.
AnswerB

For a Lambda function to successfully connect to an Amazon RDS database, the RDS instance's security group must explicitly permit inbound traffic on the database port (e.g., 3306 for MySQL, 5432 for PostgreSQL). A common best practice is to configure the RDS security group to allow inbound connections from the *security group associated with the Lambda function's ENIs*. If this rule is missing or incorrectly configured, the connection will be blocked, making this a highly probable cause of connectivity issues.

Why this answer

The Lambda function is timing out when trying to connect to the RDS database, which is in the same VPC. The most likely cause is that the RDS database's security group does not have an inbound rule allowing traffic from the Lambda function's security group on the database port (e.g., 3306 for MySQL, 5432 for PostgreSQL). Without this rule, the TCP connection attempt is silently dropped or rejected, causing the Lambda function to wait until its 30-second timeout expires.

Exam trap

The trap here is that candidates often assume the Lambda function needs internet access or a NAT gateway to communicate with an RDS database in the same VPC, overlooking the fact that security group rules are the primary control for inbound traffic within a VPC.

How to eliminate wrong answers

Option A is wrong because the Lambda function does not need internet access to reach an RDS database in the same VPC; private subnet communication within a VPC does not require a public IP or NAT gateway. Option C is wrong because the RDS database being publicly accessible is irrelevant when both resources are in the same VPC; DNS resolution of the database endpoint works via the VPC's internal DNS, and the Lambda function can resolve it without public access. Option D is wrong because a VPC endpoint for Amazon RDS is used for accessing RDS API operations (e.g., CreateDBInstance), not for database client connections (e.g., MySQL/PostgreSQL protocol), and the scenario explicitly states there is no NAT gateway, but the Lambda function does not need one to communicate within the VPC.

137
MCQhard

A web application runs on Amazon EC2 instances behind an Application Load Balancer (ALB). During rolling updates of the Auto Scaling group, users intermittently receive HTTP 502 (Bad Gateway) errors. The developer checks the ALB access logs and notices that requests are being routed to instances that are in the 'Draining' state. The ALB has connection draining enabled with a timeout of 30 seconds. The Auto Scaling group terminates instances after they are taken out of service. What is the most likely cause of the 502 errors?

A.The connection draining timeout is too short, causing the ALB to terminate connections before in-flight requests finish.
B.The health check interval is set too long, causing the ALB to consider unhealthy instances as healthy.
C.Cross-zone load balancing is disabled, so the ALB is routing requests to instances that are already draining.
D.The Auto Scaling group's minimum size is too small, causing the ALB to have no healthy targets.
AnswerA

When an EC2 instance is deregistered from an Application Load Balancer (ALB) target group, connection draining (also known as deregistration delay) begins. During this period, the ALB stops sending new requests to the instance but attempts to allow existing in-flight requests to complete. If the configured deregistration delay timeout is shorter than the time required for active requests to finish processing, the ALB will forcibly close those connections, leading to 502 Bad Gateway errors for the client, as the backend server did not return a proper response.

Why this answer

The 502 errors occur because the ALB's connection draining timeout of 30 seconds is too short to allow all in-flight requests to complete before the Auto Scaling group terminates the instances. When an instance enters the 'Draining' state, the ALB stops sending new requests but waits up to the draining timeout for existing connections to finish. If the timeout expires before requests complete, the ALB forcibly closes connections, resulting in HTTP 502 (Bad Gateway) errors for clients whose requests were still in progress.

Exam trap

The trap here is that candidates often confuse connection draining timeout with health check interval, assuming that a long health check interval causes the ALB to route to unhealthy instances, when in fact the 502 errors are caused by the ALB forcibly terminating connections before in-flight requests complete due to an insufficient draining timeout.

How to eliminate wrong answers

Option B is wrong because a long health check interval would cause the ALB to consider unhealthy instances as healthy for longer, but the issue here is that requests are being routed to instances already in the 'Draining' state, not that unhealthy instances are mistakenly considered healthy. Option C is wrong because cross-zone load balancing affects how traffic is distributed across Availability Zones, not the routing of requests to draining instances; the ALB routes to draining instances only when connection draining is active, regardless of cross-zone settings. Option D is wrong because a small minimum size would cause a lack of healthy targets, leading to 503 errors, not 502 errors; the 502 errors here are specifically tied to connection termination during draining, not insufficient capacity.

138
Multi-Selecteasy

A developer is using Amazon RDS for MySQL and notices that the database performance has degraded. The developer suspects that slow queries are the cause. Which THREE actions should the developer take to identify and address the slow queries?

Select 3 answers
A.Enable the slow query log in RDS and review the logs.
B.Increase the DB instance size to improve performance.
C.Enable Performance Insights to analyze database performance.
D.Use the RDS console to review metrics for high CPU or IOPS usage.
E.Create a read replica to offload read traffic.
AnswersA, C, D

The slow query log records every SQL statement that takes longer than the `long_query_time` threshold to execute, capturing the exact query text, execution time, lock time, and rows examined. Enabling it via the RDS parameter group (`slow_query_log=1`) is the most direct way to pinpoint which specific statements are causing the observed slowdown, allowing targeted optimization such as adding indexes or rewriting the query. This makes it the definitive first step for diagnosing slow queries at the statement level rather than relying on inferred metrics.

Why this answer

Option A is correct because enabling the MySQL slow query log on RDS captures queries exceeding long_query_time, and these logs can be downloaded or published to CloudWatch Logs for review to pinpoint offending SQL. Option C is correct because Performance Insights provides a database load view (DB load by wait event and SQL) so the developer can identify the top SQL statements and waits causing degradation. Option D is correct because RDS console CloudWatch metrics such as CPUUtilization and ReadIOPS/WriteIOPS help correlate resource saturation with suspected slow queries and confirm the bottleneck.

Option B is not appropriate as a first step because resizing the instance masks symptoms without identifying the slow queries and may not resolve inefficient SQL. Option E is also not appropriate because a read replica offloads read traffic but does not diagnose or fix the slow queries themselves.

Exam trap

DVA-C02 often tests the difference between diagnostic actions (slow query log, Performance Insights, CloudWatch metrics) and remediation/scaling actions (resize instance, read replica), so candidates who pick scaling options as 'fixes' miss the intent of identifying slow queries.

139
MCQhard

A developer notices that the Lambda function 'my-function' is not generating any logs in CloudWatch, although the function is invoked successfully. The developer runs the command above. What is the MOST likely cause?

A.The log group retention policy is set to 0 days.
B.The Lambda function is configured to log to a custom log group.
C.The Lambda function has reserved concurrency set to 0.
D.The Lambda function's execution role is missing the 'logs:CreateLogStream' and 'logs:PutLogEvents' permissions.
AnswerD

The Lambda function's execution role requires specific permissions to interact with CloudWatch Logs. The 'logs:CreateLogStream' permission allows the function to create a unique log stream for its execution environment within the designated log group, while 'logs:PutLogEvents' is essential for sending the actual log data to that stream. Without both of these critical permissions, the function cannot successfully publish any output or runtime messages to CloudWatch Logs, leading to the observation of no logs.

Why this answer

The log group exists but has 0 stored bytes, meaning no log streams have been created. This typically indicates that the Lambda function's execution role does not have permissions to create log streams and put log events. The function runs but fails silently to write logs.

140
MCQeasy

A developer runs a script that uses the AWS CLI to copy a large number of files from an on-premises server to an S3 bucket. The copy operation fails partway through with a 'RequestTimeout' error. What is the MOST efficient way to resume the copy and ensure all files are transferred?

A.Delete the S3 bucket and restart the copy operation.
B.Use the aws s3 sync command to synchronize the source directory with the S3 bucket.
C.Use the cp command with the --recursive flag to copy the remaining files.
D.Increase the --cli-read-timeout value in the AWS CLI configuration and retry the original command.
AnswerB

The aws s3 sync command is the most appropriate and efficient solution for resuming an interrupted file transfer to S3. It intelligently compares the source directory with the S3 bucket, identifying only files that are new, have changed content (based on size and modification time), or are missing from the destination. This ensures that only the necessary data is transferred, minimizing bandwidth usage and significantly reducing the time required to complete the operation.

Why this answer

The `aws s3 sync` command is the most efficient way to resume the copy because it automatically compares the source directory with the destination S3 bucket and transfers only the files that are missing or have been modified. This avoids re-uploading already transferred files, directly addressing the partial failure without manual intervention or unnecessary overhead.

Exam trap

The trap here is that candidates often confuse `cp --recursive` with `sync`, assuming both can resume a copy, but only `sync` performs a differential comparison to avoid re-uploading already transferred files.

How to eliminate wrong answers

Option A is wrong because deleting the S3 bucket and restarting the entire copy operation is extremely inefficient and unnecessary; it would re-upload all files, including those already successfully transferred. Option C is wrong because the `cp --recursive` command does not perform any comparison or state tracking; it would blindly copy all files from the source again, potentially re-uploading already transferred files and wasting time and bandwidth. Option D is wrong because increasing the `--cli-read-timeout` only extends the time the CLI waits for a response from the S3 service; it does not address the root cause of the partial failure (e.g., network interruptions or throttling) and would not resume the copy from where it left off, nor does it skip already transferred files.

141
MCQeasy

An application running on Amazon EC2 generates logs that need to be streamed to Amazon CloudWatch Logs. The developer installs and configures the CloudWatch agent. However, logs are not appearing in the log group. What is the most likely cause?

A.The EC2 instance does not have an IAM role with CloudWatch Logs write permissions.
B.The CloudWatch agent cannot be installed on Amazon Linux 2.
C.The CloudWatch agent must be configured from the AWS Management Console.
D.The log group must be created manually before the agent can send logs.
AnswerA

The CloudWatch agent, when running on an EC2 instance, requires an associated IAM role with specific permissions to interact with CloudWatch Logs. Without an IAM policy granting actions such as logs:CreateLogGroup, logs:CreateLogStream, and logs:PutLogEvents, the agent will be unauthorized to publish log events to the designated log group and stream. This is a fundamental security requirement, ensuring that AWS services only perform actions for which they have explicit authorization.

Why this answer

The CloudWatch agent requires IAM permissions to call CloudWatch Logs APIs (CreateLogGroup, CreateLogStream, PutLogEvents). On EC2, these permissions are provided via an instance profile attached to the instance. If the instance lacks an IAM role with CloudWatch Logs write permissions, the agent cannot publish logs, and they will not appear in the log group.

Exam trap

DVA-C02 often tests whether candidates know the CloudWatch agent needs an IAM role with logs permissions — candidates incorrectly blame installation, console configuration, or pre-created log groups instead of the missing instance profile.

How to eliminate wrong answers

Option B is wrong because the CloudWatch agent is fully supported on Amazon Linux 2 — it is the recommended platform. Option C is wrong because the agent is configured via a JSON config file on the instance (or via SSM), not from the AWS Management Console. Option D is wrong because the agent can create the log group automatically if the config specifies it and the IAM role has CreateLogGroup permission; manual creation is not required.

142
MCQeasy

A developer is deploying a new version of an application to Amazon ECS using the Fargate launch type. The task fails to start and the error message indicates that the task cannot pull the container image from Amazon ECR. What is the MOST likely cause?

A.The task definition family name is incorrect.
B.The task execution role lacks permissions to pull from ECR.
C.The container port is not mapped to a host port.
D.The CPU or memory limits are too low for the container.
AnswerB

The task execution role is critical for allowing the Amazon ECS agent to perform necessary actions on your behalf, including pulling container images from private repositories like Amazon ECR. If this role lacks specific permissions such as `ecr:GetDownloadUrlForLayer`, `ecr:BatchGetImage`, and `ecr:BatchCheckLayerAvailability`, the ECS agent will be unauthorized to retrieve the image layers. Consequently, the container runtime will fail to download the image, resulting in a distinct image pull error.

Why this answer

The error message indicates a failure to pull the container image from Amazon ECR. The task execution role must have permissions such as ecr:GetDownloadUrlForLayer and ecr:BatchGetImage to pull images. Option A is wrong because the task definition family name does not affect image pulling.

Option C is wrong because Fargate does not use host port mapping; container port mapping is handled automatically. Option D is wrong because insufficient CPU or memory typically results in a different error (e.g., task stuck in provisioning), not an image pull failure.

143
MCQhard

A company's DynamoDB table has a read capacity of 10,000 RCUs and receives consistent traffic. Recently, users have reported increased latency for read requests. The application uses strongly consistent reads. The developer checks CloudWatch metrics and sees that 'ConsumedReadCapacityUnits' is at 9,500 but 'ThrottledRequests' is high. What is the most likely cause?

A.The application is using eventually consistent reads but expecting strongly consistent results.
B.A hot partition is exceeding its partition-level read capacity.
C.The DynamoDB table has auto scaling enabled and is scaling down too aggressively.
D.The provisioned read capacity is too low for the traffic.
AnswerB

DynamoDB distributes provisioned capacity evenly across its underlying partitions. If a specific partition key receives a disproportionately high volume of read requests, it can exhaust its allocated share of the table's total read capacity, even if the overall table capacity is not fully utilized. This scenario, known as a 'hot partition,' causes throttling errors for requests targeting that specific partition, despite ample table-level RCUs.

Why this answer

The correct answer is B: a hot partition is exceeding its partition-level read capacity. Even though the table's total consumed read capacity (9,500 of 10,000 RCUs) is below the provisioned limit, DynamoDB distributes capacity across partitions, and a single partition can only support a maximum of 3,000 RCUs (or 1,000 WCUs); if one partition key receives disproportionate traffic, that partition throttles requests while overall table capacity remains underutilized, which matches the high ThrottledRequests with consumed capacity below the table maximum. Option D is wrong because the table-level provisioned capacity is not exhausted (9,500 < 10,000), so low capacity is not the cause.

Option A is wrong because the scenario states the application uses strongly consistent reads, and switching consistency would not explain throttling. Option C is wrong because auto scaling scaling down would reduce provisioned capacity, but the metric shows consumed capacity still below the provisioned 10,000 RCUs, and aggressive scale-down is not the typical cause of partition-level throttling.

144
MCQmedium

A company uses AWS CodePipeline with CodeBuild to test and deploy a web application. The pipeline has been failing at the deploy stage with an error: 'Access Denied'. CloudTrail shows the CodePipeline service role is making the call. What is the MOST likely cause?

A.The CodeBuild project does not have internet access.
B.The CodePipeline service role lacks permissions for the deploy action.
C.The deploy provider (e.g., ECS, S3) is not in the same AWS region.
D.The source code repository does not have the correct branch.
AnswerB

An 'Access Denied' error during the deploy stage is a classic indication that the AWS CodePipeline service role lacks the necessary IAM permissions to perform the deployment actions on the target AWS resource. For instance, if deploying to an S3 bucket, the role needs `s3:PutObject` and `s3:GetObject` permissions for the artifact. Without these explicit `Allow` statements in its policy, the service principal is unauthorized to interact with the target service, resulting in the reported access denial.

Why this answer

The error 'Access Denied' in the deploy stage, with CloudTrail showing the CodePipeline service role making the call, indicates that the IAM role assumed by CodePipeline does not have the necessary permissions to perform the deploy action against the target provider (e.g., ECS, S3, Elastic Beanstalk). CodePipeline uses its service role to invoke the deploy action, and if that role lacks the required `codedeploy:*`, `s3:PutObject`, or `ecs:UpdateService` permissions, the API call will be denied.

Exam trap

The trap here is that candidates confuse the CodeBuild service role with the CodePipeline service role, assuming the build role is responsible for deployment, when in fact CodePipeline uses its own role for the deploy action.

How to eliminate wrong answers

Option A is wrong because CodeBuild not having internet access would cause build failures (e.g., cannot download dependencies), not a deploy-stage 'Access Denied' error, and CloudTrail shows the CodePipeline service role, not CodeBuild, is making the call. Option C is wrong because deploy providers can be in different regions (cross-region actions are supported with appropriate IAM and resource policies), and the error is 'Access Denied', not a region mismatch. Option D is wrong because an incorrect source branch would cause the pipeline to fetch the wrong code or fail at the source stage, not produce an 'Access Denied' error at the deploy stage.

145
MCQhard

Messages in an SQS queue are processed successfully but later reappear and are processed again. What is the most likely configuration issue?

A.The queue uses long polling
B.The queue has a dead-letter queue
C.The messages are encrypted with SSE-SQS
D.The visibility timeout is shorter than the processing time or messages are not deleted after processing
AnswerD

If the visibility timeout is shorter than the actual time required to process a message, the message will become visible again to other consumers before the initial consumer finishes and deletes it, leading to duplicate processing. Alternatively, if a consumer successfully processes a message but fails to explicitly call the `DeleteMessage` API, the message will remain in the queue and become visible again once its timeout expires, resulting in reprocessing. Both scenarios directly explain why messages might be processed successfully but still reappear.

Why this answer

When a message is processed but not deleted from the SQS queue, or when the visibility timeout expires before processing completes, the message becomes visible again in the queue and can be consumed by another worker. This causes duplicate processing. The correct fix is to ensure the visibility timeout is set longer than the expected processing time and that the message is explicitly deleted after successful processing.

Exam trap

The trap here is that candidates may confuse message reappearance with dead-letter queue behavior, but dead-letter queues only trigger after a configurable number of receive attempts, not after a single successful processing cycle.

How to eliminate wrong answers

Option A is wrong because long polling reduces empty responses and cost by waiting for messages, but does not cause messages to reappear after processing. Option B is wrong because a dead-letter queue captures messages that have failed processing multiple times, not cause reprocessing of successfully handled messages. Option C is wrong because SSE-SQS encrypts messages at rest, which has no effect on message visibility or deletion behavior.

146
Multi-Selecteasy

Which TWO actions can help reduce Lambda cold start times? (Choose two.)

Select 2 answers
A.Increase the deployment package size.
B.Increase the memory allocated to the function.
C.Use Provisioned Concurrency.
D.Place the function in a VPC.
E.Reduce the function timeout.
AnswersB, C

Lambda allocates CPU proportionally to the amount of memory configured, so more memory means more CPU power available during initialization. This speeds up tasks like loading the runtime, unpacking code, and running static initializers, thereby shortening the cold start duration. It is a practical tuning knob, though it increases cost per invocation.

Why this answer

Option B is correct because AWS Lambda allocates CPU proportionally to the configured memory, so increasing the memory allocated to the function gives it more CPU power and speeds up initialization of the runtime and code, thereby reducing cold start duration. Option C is correct because Provisioned Concurrency pre-initializes a requested number of execution environments and keeps them warm, so invocations are served by already-initialized environments and avoid the cold start entirely. Option A is incorrect because a larger deployment package takes longer to download and unpack during initialization, which increases cold start time.

Option D is incorrect because placing the function in a VPC adds ENI creation and attachment overhead during initialization, typically worsening cold starts. Option E is incorrect because the function timeout only limits how long an invocation may run; it does not affect initialization time and can even cause failures if set too low.

Exam trap

DVA-C02 often tests the misconception that a larger deployment package or VPC placement improves cold start performance, when in reality both tend to worsen it; the correct levers are memory allocation and Provisioned Concurrency.

147
Multi-Selecthard

A company runs a serverless application on AWS using API Gateway, AWS Lambda, and DynamoDB. The application processes user uploads and stores metadata in DynamoDB. Recently, users have reported that some uploads fail with a 500 Internal Server Error. The CloudWatch Logs for the Lambda function show 'ProvisionedThroughputExceededException' errors for DynamoDB, followed by 'Task timed out after 3.00 seconds' errors. The Lambda function has a 3-second timeout and 128 MB of memory. The DynamoDB table has 5 read capacity units and 5 write capacity units. The application uses a single Lambda function that processes each upload synchronously. The company expects a steady increase in uploads. Which combination of actions should a developer take to resolve the errors and prepare for future growth? (Choose TWO.)

Select 2 answers
A.Increase the DynamoDB table's write capacity units to a higher value.
B.Switch the Lambda function to asynchronous invocation with a DLQ.
C.Modify the Lambda function to implement retries with exponential backoff on DynamoDB write operations.
D.Increase the Lambda function's timeout to 30 seconds.
E.Increase the Lambda function's reserved concurrency to 100.
AnswersA, C

The ProvisionedThroughputExceededException explicitly indicates that the DynamoDB table's allocated write capacity has been surpassed. Increasing the Write Capacity Units (WCUs) directly provisions more throughput for the table, allowing it to handle a higher volume of write requests per second without throttling. This is a direct and effective solution to prevent future throttling errors by scaling the underlying resource.

Why this answer

The errors are caused by DynamoDB throttling due to insufficient write capacity. Option A increases write capacity to handle the load. Option C implements retries with exponential backoff to handle occasional throttling without failing.

Option B would not help because the errors are from DynamoDB, not Lambda concurrency. Option D would increase latency but not solve throttling. Option E might cause duplicate processing.

148
MCQmedium

A developer is troubleshooting an AWS CloudFormation stack that failed to create. The error message says 'The following resource(s) failed to create: [MyEC2Instance]'. What is the first step the developer should take?

A.Update the stack with a new template.
B.Delete the stack and try again.
C.Review the CloudFormation template for syntax errors.
D.View the stack events in the CloudFormation console to see the specific error for the resource.
AnswerD

The CloudFormation console's "Events" tab provides a chronological log of every action taken by the stack, including resource creation attempts, status changes, and, critically, any errors encountered. When a resource fails to create, CloudFormation logs a specific CREATE_FAILED event for that resource, often including the underlying AWS service error message (e.g., "User is not authorized to perform this operation," "The specified S3 bucket already exists"). This detailed information is essential for diagnosing the exact cause of the failure.

Why this answer

When a CloudFormation stack fails to create, the error message only indicates which resource failed, not why. The first troubleshooting step is to view the stack events in the CloudFormation console, which provides detailed error messages for each resource, such as an API call failure, insufficient permissions, or a resource limit exceeded. This allows the developer to diagnose the root cause before making any changes.

Exam trap

The trap here is that candidates often jump to fixing the template or retrying the stack, overlooking that the specific error details are available in the stack events, which is the fastest path to identifying the actual cause.

How to eliminate wrong answers

Option A is wrong because updating the stack with a new template without understanding the failure reason could introduce additional errors or mask the underlying issue. Option B is wrong because deleting the stack and retrying without investigation wastes time and may repeat the same failure if the root cause (e.g., a missing parameter or IAM role) is not addressed. Option C is wrong because syntax errors in the template would typically be caught during validation before stack creation, and the error message specifically indicates a resource creation failure, not a template syntax issue.

149
MCQmedium

A company's application running on Amazon ECS Fargate is experiencing high CPU utilization. The task definition has CPU set to 256 units. What should be done to improve performance?

A.Increase the desired count of tasks.
B.Increase the CPU value in the task definition and redeploy the service.
C.Increase the memory value in the task definition.
D.Switch to EC2 launch type.
AnswerB

The "cpu" parameter in an ECS Fargate task definition directly controls the amount of virtual CPU units allocated to each running task. By increasing this value, the application container within the task gains access to more processing power, enabling it to execute computations faster and reduce its overall CPU utilization percentage. Redeploying the service ensures that all new tasks launched by ECS will utilize this updated, higher CPU allocation, directly addressing the bottleneck.

Why this answer

Increasing CPU units in the task definition and redeploying the service will allocate more CPU to the tasks. Option A is wrong because horizontal scaling can help but the root cause is insufficient CPU per task. Option C is wrong because increasing memory does not affect CPU.

Option D is wrong because changing the launch type changes billing but not CPU allocation.

150
MCQmedium

A developer is troubleshooting a Lambda function that intermittently times out. The function makes HTTP requests to an external API. The function's CloudWatch logs show 'Task timed out after 3.01 seconds'. What is the MOST likely cause?

A.The Lambda function timeout is set to 3 seconds, but the HTTP request takes longer.
B.The Lambda function has insufficient reserved concurrency causing throttling.
C.The Lambda function is not configured with a VPC and cannot reach the external API.
D.The Lambda function is not starting execution due to a missing IAM role.
AnswerA

The default timeout for an AWS Lambda function is indeed 3 seconds. When an HTTP request or any other operation within the function exceeds this configured duration, Lambda forcefully terminates the execution, logging a 'Task timed out' error. The log showing 3.01 seconds precisely indicates that the function was terminated just after exceeding its allowed execution time, making the long-running HTTP request the root cause of the observed timeout.

Why this answer

The Lambda function timeout is set to 3 seconds, and the HTTP request to the external API takes longer than that, causing the 'Task timed out after 3.01 seconds' error. Option B is incorrect because throttling due to insufficient reserved concurrency would result in a 'Rate exceeded' error, not a timeout. Option C is incorrect because if the function couldn't reach the external API due to VPC configuration, it would result in a connection error, not a timeout.

Option D is incorrect because a missing IAM role would prevent the function from executing at all, and the logs show the function started execution.

← PreviousPage 2 of 3 · 179 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Troubleshooting and Optimization questions.