DVA-C02 · domain
Troubleshooting and Optimization
This domain covers diagnosing and tuning workloads on AWS: Lambda logging and invocation failures, CodeBuild/CodePipeline build errors, ECS Fargate migrations, and event-driven processing with S3 and DynamoDB. Questions present a symptom, a console or CLI observation, and ask for the most likely cause or the best fix, so you must reason from service behavior rather than memorized limits.
Focused practice
Practice Troubleshooting and Optimization questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Troubleshooting and Optimization
Be able to inspect IAM roles, VPC egress, log group configuration, and event delivery semantics to isolate the root cause, then apply the smallest fix. The single most important skill is mapping a symptom to the specific AWS configuration that produces it.
Reading Lambda CloudWatch Logs permissions and execution role policies to explain missing log output
Diagnosing CodeBuild npm network failures caused by VPC configuration, NAT gateways, or proxy settings
Choosing ECS Fargate deployment and load balancer settings that minimize downtime during migration
Tracing S3 event notification delivery to Lambda and idempotent DynamoDB writes for duplicate records
Watch out for
Common Troubleshooting and Optimization exam traps
- ▸Assuming a successful Lambda invocation means logs exist; missing logs usually point to the execution role lacking CloudWatch Logs permissions or a misconfigured log group.
- ▸Blaming application code for intermittent npm ERR! network failures in CodeBuild when the real cause is a VPC without NAT or missing egress rules.
- ▸Treating S3 event notifications as exactly-once delivery, so retries create duplicate Lambda invocations and duplicate DynamoDB items without idempotency keys.
Question index
All Troubleshooting and Optimization questions (179)
Click any question to see the full explanation, or start a practice session above.
A company is using Amazon DynamoDB with on-demand capacity. A developer notices that write requests are being throttled during peak hours. What is the MOST effective way to resolve this issue?
Hard2An application running on Amazon ECS Fargate is experiencing intermittent connection timeouts when calling an external API. The task has a public IP and a security group that allows outbound HTTPS. What is the most likely cause?
Hard3A developer is troubleshooting an AWS Lambda function that times out when processing large files from Amazon S3. The function has a 15-minute timeout and 512 MB memory. What should the developer do to resolve this issue?
Easy4A company runs a web application on EC2 instances behind an Application Load Balancer (ALB). The application stores session data in an RDS MySQL database. Recently, users have reported that they are being logged out unexpectedly and their session data is lost. The developer investigates and finds that the RDS instance's CPU utilization spikes periodically, coinciding with the logout events. The application uses connection pooling via an RDS Proxy. The developer suspects that the session table is being dropped or truncated. After checking the application logs, the developer finds no evidence of truncation commands. The RDS instance has automated backups enabled, and the binary logs are retained for 24 hours. The developer wants to identify the root cause and prevent future occurrences. Which course of action should the developer take?
Hard5A developer is optimizing an AWS Lambda function that processes streaming data from Amazon Kinesis. The function is CPU-bound. Which TWO actions should the developer take to improve performance?
Hard6A developer is using Amazon DynamoDB for a new application. The developer wants to reduce read latency. Which design pattern should the developer use?
Easy7An application uses an Amazon SQS queue to decouple microservices. The producer is sending messages, but the consumer is not processing them. The consumer is an Auto Scaling group of EC2 instances. The SQS queue's ApproximateNumberOfMessagesVisible metric is increasing. What is the MOST likely cause?
Hard8A developer is using AWS CodePipeline to deploy a web application. The pipeline has a source stage from GitHub and a deploy stage to Elastic Beanstalk. The deploy stage fails with the error 'The S3 bucket does not allow access to the artifact'. Which THREE actions could resolve this issue?
Hard9A developer is deploying a serverless application using AWS CloudFormation. The stack creation fails with the error 'CREATE_FAILED: The following resource(s) failed to create: [MyLambdaFunction]'. The developer checks the CloudFormation events and sees 'Resource creation cancelled'. What is the most likely cause?
Medium10A web application running on EC2 instances behind an Application Load Balancer (ALB) is experiencing intermittent 503 errors. The ALB target group health checks are succeeding. Which step should the developer take FIRST to diagnose the issue?
Medium11A developer is deploying a serverless application using AWS SAM. The deployment fails with the error 'Resource creation cancelled'. What is the most likely cause?
Easy12A Lambda function reading from Kinesis is falling behind. Which two metrics/settings should be reviewed first?
Hard13A developer is deploying a CloudFormation stack and sees the event above. What should the developer do to fix the error?
Easy14An application running on Amazon ECS with Fargate is unable to pull an image from Amazon ECR. The task definition uses the 'default' task execution role. What is the most likely cause?
Easy15A developer is troubleshooting an AWS Lambda function that is triggered by an Amazon SQS queue. The function processes messages but occasionally fails. The failed messages are not being sent to the dead-letter queue (DLQ). What is the most likely reason?
Medium16A developer monitors an AWS Lambda function that processes messages from an Amazon SQS queue. CloudWatch logs show that the function's execution time has increased significantly over the past week. The function's code has not been changed recently. The function makes calls to an Amazon DynamoDB table. CloudWatch metrics show a high rate of DynamoDBProvisionedThroughputExceededException errors. The DynamoDB table has 5 read and 5 write capacity units (RCU/WCU). What is the most effective action to reduce the function's execution time?
Medium17A developer reports that an AWS Lambda function is timing out after 3 seconds. The function reads from an Amazon SQS queue. What is the most likely cause?
Easy18A developer is using Amazon DynamoDB as the database for a web application. The application experiences occasional spikes in traffic, and some write requests fail with a ProvisionedThroughputExceededException. What is the MOST cost-effective way to handle these spikes without manual intervention?
Easy19A web application runs on Amazon EC2 instances behind an Application Load Balancer (ALB). During peak hours, users report receiving HTTP 503 (Service Unavailable) errors. The developer checks Amazon CloudWatch metrics and finds that the ALB's request count is high but below the limit, and the target group's healthy host count drops to zero intermittently. The Auto Scaling group for the instances is configured with a minimum of 2, maximum of 10, and a simple scaling policy to add 2 instances when CPU utilization exceeds 70% for 5 consecutive minutes. What is the most likely cause of the 503 errors?
Hard20An application running on EC2 instances behind an Application Load Balancer (ALB) occasionally returns HTTP 503 errors. The instances are in an Auto Scaling group. Which action should be taken to resolve this issue?
Hard21Refer to the exhibit. A developer invoked a Lambda function and received this response. What does the FunctionError field indicate?
Medium22A developer notices that an AWS Lambda function configured with a VPC is timing out when trying to access an Amazon S3 bucket. The function has the necessary IAM permissions. What is the most likely cause?
Medium23A developer deploys an application on EC2 instances behind an Application Load Balancer (ALB). The application uses sticky sessions (session affinity) based on a cookie. Users report that they are intermittently logged out during their session. What is the MOST likely cause?
Medium24A company runs a production web application on EC2 instances behind an Application Load Balancer. Users report intermittent 502 errors. The developers find that the ALB access logs show 'target_response_code' of 502 for some requests. What is the MOST likely cause?
Hard25A developer is using AWS CodePipeline to automate a multi-stage pipeline. The pipeline includes a manual approval step before deploying to production. The developer wants to receive an email notification when the pipeline reaches the approval step. Which service should the developer use?
Hard26A developer is deploying an AWS Elastic Beanstalk application and notices that the environment's health is degraded because the application is returning HTTP 5xx errors. The developer wants to quickly identify the root cause without redeploying. Which action should the developer take?
Easy27A developer optimized an Amazon S3 bucket for high request rates. The bucket receives over 5,000 PUT requests per second. Recently, some requests are failing with a 503 Slow Down error. What is the most likely cause and how should the developer fix it?
Hard28A developer is using Amazon CloudFront to distribute content from an S3 bucket. The bucket is configured as an origin with Origin Access Control (OAC). Recently, some users have reported that they receive 403 Forbidden errors when accessing certain objects. The developer checks the CloudFront distribution and confirms that the OAC is set up correctly. The S3 bucket policy allows the CloudFront service principal to get objects. The developer also notes that the objects in question have been updated recently. What is the MOST likely cause of the 403 errors?
Hard29A company uses AWS CodePipeline with CodeBuild to deploy a Node.js application. The build fails intermittently with 'npm ERR! network' errors. What is the most likely cause and solution?
Hard30A developer is troubleshooting a slow Amazon RDS for MySQL database. The application experiences high latency on write operations. Which TWO actions can improve write performance?
Easy31A developer runs the AWS CLI command shown in the exhibit. The output includes 'FunctionError': 'Unhandled'. What does this indicate?
Medium32A developer is using Amazon API Gateway with a Lambda authorizer to control access to APIs. The authorizer is failing with a 500 error. The Lambda function logs show 'User: arn:aws:iam::123456789012:role/MyLambdaRole is not authorized to perform: sts:AssumeRole'. What is the most likely cause?
Hard33A web application running on Amazon EC2 instances behind an Application Load Balancer (ALB) is experiencing intermittent 503 errors. Which TWO steps should be taken to diagnose the issue?
Easy34An application running on Amazon ECS with Fargate is experiencing high latency. The application writes logs to Amazon CloudWatch Logs. Which AWS service can be used to analyze the logs to pinpoint the cause of the latency?
Medium35An application running on Amazon ECS (Fargate) uses an Application Load Balancer (ALB) with connection draining enabled. The application is experiencing intermittent 502 (Bad Gateway) errors during rolling updates of the ECS service. The developer notices that the ALB is routing requests to tasks that are in the 'Draining' state. The ECS service is configured with a deployment circuit breaker that automatically rolls back a failed deployment. What is the most likely cause of the 502 errors?
Hard36A developer monitors an AWS Lambda function that processes messages from an Amazon SQS queue. CloudWatch logs show that the function's execution time has increased significantly over the past week, and it now frequently times out at the 5-minute timeout. The function's code has not been changed recently. The function makes calls to an Amazon DynamoDB table. What is the most likely cause of the increased execution time?
Medium37A developer deployed a new version of an AWS Lambda function that is part of a serverless application. The function uses an Amazon DynamoDB table as a data store. After deployment, the developer notices that the function's latency has increased significantly for some requests. CloudWatch traces show that the increase is due to DynamoDB throttle events. The function is configured with a reserved concurrency of 100 and the DynamoDB table has 5 read capacity units (RCUs) and 5 write capacity units (WCUs). What is the most effective way to reduce the throttling while maintaining application performance?
Hard38A company runs a web application on EC2 instances behind an Application Load Balancer (ALB). Users report intermittent 503 errors. The ALB health checks are failing for a few instances, but the instances themselves are running and have healthy application processes. What is the MOST likely cause?
Medium39A developer receives an Access Denied error when trying to download an object from an S3 bucket. The developer's IAM policy is shown in the exhibit. The bucket policy also grants access. What is the MOST likely cause?
Hard40A developer is troubleshooting an AWS Lambda function that experiences high latency for the first few invocations after being idle. The function is written in Python and uses a large library (e.g., Pandas). The function connects to an RDS database in a VPC. What is the most effective way to reduce the latency for the first invocation after idle?
Hard41A developer notices that an AWS Lambda function processing S3 events is being retried frequently due to throttling errors from Amazon DynamoDB. The function writes records to a DynamoDB table and has reserved concurrency set to 100. The DynamoDB table uses on-demand capacity mode. What should the developer do to reduce retries and improve overall throughput?
Medium42A developer is troubleshooting an AWS Lambda function that is failing with an 'AccessDenied' error when trying to write to an S3 bucket. The function's execution role has the following policy. What is the most likely cause of the failure? (Policy: { 'Version': '2012-10-17', 'Statement': [ { 'Effect': 'Allow', 'Action': 's3:PutObject', 'Resource': 'arn:aws:s3:::my-bucket/*' } ] })
Easy43The exhibit shows an IAM policy attached to a Lambda function's execution role. The function writes objects to an S3 bucket that is encrypted with a KMS key (the key specified in the policy). When the function tries to write an object, it receives an access denied error. What is the MOST likely missing permission?
Hard44A developer is running a Docker container on Amazon ECS with Fargate. The container logs are not appearing in CloudWatch Logs even though the task definition has a logConfiguration specifying the awslogs driver and a log group. What is the MOST likely missing configuration?
Medium45A developer notices that an AWS Lambda function, which processes messages from an SQS queue, is taking longer than expected. The function has a reserved concurrency of 5 and a batch size of 10. The SQS queue has a large backlog. CloudWatch metrics show that the function's throttles are high. The function is idempotent and can process up to 100 messages per invocation. What is the most effective way to increase throughput without increasing reserved concurrency?
Hard46An application running on Amazon ECS Fargate is experiencing intermittent high latency and timeout errors. The application makes API calls to an external third-party service. The ECS service is configured with a target group using HTTP health checks. The ALB health check logs show occasional 503 responses. What is the MOST likely cause?
Hard47A developer notices that an S3 bucket used for static website hosting returns 403 Forbidden for anonymous requests. The bucket policy allows s3:GetObject for Principal "*". What is the most likely issue?
Easy48A company runs a monolithic application on EC2 Behind an Application Load Balancer. They want to migrate to a microservices architecture using ECS Fargate. What is the most important optimization to ensure minimal downtime during the migration?
Hard49A developer configured an S3 bucket to trigger a Lambda function on object creation. The Lambda function processes the object and then deletes it. Some objects are not being processed. What should the developer do to ensure all objects are processed?
Medium50A developer is troubleshooting an AWS Lambda function that returns timeout errors when calling an external HTTPS API. The function is configured with a 30-second timeout and runs in a VPC with a public subnet and NAT Gateway. The developer checks CloudWatch logs and sees that the function is timing out at exactly 30 seconds. What is the most likely cause?
Medium51A developer is troubleshooting performance issues in an application that uses Amazon DynamoDB as the primary data store. The application reads a large set of items using a Query operation on a Global Secondary Index (GSI). The developer notices high read latency and throttled requests on the GSI. The base table has sufficient read capacity. The GSI is projected with KEYS_ONLY. Which action would most likely reduce the latency and throttling?
Hard52A developer is troubleshooting an EC2 instance that cannot connect to the internet. The instance has a public IP address and is in a public subnet with a route to an internet gateway. The security group allows all outbound traffic. What is the most likely cause?
Easy53A developer is troubleshooting an AWS Lambda function that processes messages from an Amazon SQS queue. The function is configured with a batch size of 10 and a maximum concurrency of 5. The function frequently reports errors related to message processing timeouts. The function code is idempotent. Which combination of actions will reduce the number of timeouts and improve processing efficiency?
Hard54A company is using Amazon S3 to store large objects. Users report that uploads are slow. Which THREE actions should the developer take to optimize upload performance?
Medium55A developer notices that an EC2 instance running a web application is unreachable via its public IP. The instance passes status checks but security group rules appear correct. What should the developer check NEXT?
Easy56Refer to the exhibit. A developer created this CloudFormation template. After deployment, the stack creation fails with 'Bucket name already exists'. What should the developer do to fix the issue?
Easy57Why is the Lambda function not being invoked?
Medium58A developer is troubleshooting an AWS Lambda function that is triggered by an S3 event. The function occasionally fails with a timeout error. CloudWatch logs show that the timeout occurs during the processing of large files. The function has a memory setting of 128 MB and a timeout of 3 seconds. The developer wants to process large files without modifying the code. Which parameter should the developer adjust first?
Medium59A developer needs to trace a request across API Gateway, Lambda, and downstream AWS service calls. Which service should be enabled?
Medium60A developer is using AWS X-Ray to trace requests through a microservices application. One of the services, Service B, is not appearing in the trace map. What is the MOST likely reason?
Easy61A developer is optimizing a Node.js Lambda function that processes CSV files from S3. The function reads the entire file into memory, processes it, and writes results to DynamoDB. For large files, the function runs out of memory. What is the MOST effective optimization?
Medium62A developer is deploying a new version of an AWS Lambda function using the AWS CLI. The deployment fails with a 'ResourceConflictException' error. What is the MOST likely cause?
Medium63A developer is using AWS Lambda with a VPC configuration. The function needs to access an Amazon RDS instance in the same VPC. The function is timing out after 3 seconds. What is the MOST likely cause?
Hard64A developer is troubleshooting an AWS Lambda function that writes to an S3 bucket. The function is configured with a resource-based policy that allows the S3 service to invoke the function. However, the function fails with an access denied error when trying to write to S3. What is the MOST likely cause?
Medium65A Lambda function processing SQS messages is failing with concurrency errors. The function is configured with reserved concurrency of 5. The SQS queue has a batch size of 10. What is the most effective way to prevent throttling?
Medium66The developer invokes a Lambda function using the AWS CLI and gets the output shown. What is the most likely cause of the error?
Medium67A developer is using Amazon DynamoDB with provisioned throughput. The application is receiving ProvisionedThroughputExceededException errors. What is the BEST way to handle this error?
Easy68A developer deploys a new version of an AWS Lambda function using the AWS CLI. After deployment, the function returns stale results. What is the most likely cause?
Easy69A company runs a microservices application on Amazon ECS with Fargate. The application includes a service that processes messages from an SQS queue. The service's CPU utilization is consistently above 80%, and messages are accumulating in the queue. The service is configured with a desired count of 2 tasks and auto scaling based on CPU utilization. What should a developer do to improve message processing throughput?
Hard70A developer is debugging an issue where an Amazon S3 bucket policy is not allowing cross-account access for a user from another AWS account. The bucket policy grants access to the other account's root user. The IAM user in the other account has an IAM policy that allows s3:GetObject on the bucket. When the user tries to download an object, they get an Access Denied error. What is the most likely cause?
Medium71Match each AWS storage class to its description.
Medium72An application uses an Auto Scaling group with a launch configuration that includes a user data script to configure instances. After a scaling event, new instances launch but fail to register with the target group. The existing instances continue to work. What should the developer do to resolve this issue?
Hard73A developer is troubleshooting a slow-running query in Amazon RDS for MySQL. The query is used by a reporting dashboard. Which AWS service should the developer use to identify the bottleneck?
Easy74A developer is troubleshooting an AWS Lambda function that is invoked from an Amazon S3 bucket via event notifications. The function processes images and stores metadata in Amazon DynamoDB. The developer notices that some images are being processed multiple times, resulting in duplicate entries in DynamoDB. The S3 event notification is configured to send events to the Lambda function with the 's3:ObjectCreated:*' event type. The function uses the 'uuid' library to generate a unique ID for each image upon processing. What is the most likely cause of the duplicate processing?
Hard75The exhibit shows a CloudFormation template that creates an S3 bucket with versioning enabled. After deploying the stack, a developer uploads an object to the bucket. Later, the developer updates the object by uploading a new version. The developer wants to retrieve the original object. What is the correct way to do this?
Easy76An API backed by Lambda returns high p95 latency after deployment. Which two telemetry sources are most useful first?
Hard77A developer notices that an Amazon RDS for MySQL DB instance's CPU utilization is consistently above 90% during peak hours. Which AWS service can the developer use to analyze the database queries and identify the root cause?
Medium78A DynamoDB table shows throttling on one partition key value. Which two signs point to a hot partition problem?
Medium79A developer is monitoring an AWS Lambda function that is triggered by an Amazon SQS queue. The function's CloudWatch metrics show a high number of throttles. The function has a reserved concurrency of 10 and the SQS queue has a large backlog of messages. The function processes each message in about 2 seconds and has a timeout of 60 seconds. Which action will most effectively reduce the throttles and increase throughput?
Medium80Which TWO are best practices for optimizing DynamoDB performance? (Choose two.)
Hard81A developer is troubleshooting an AWS Lambda function that is timing out. The function processes S3 events and writes to DynamoDB. The average execution time is 5 seconds, but the function times out after 3 seconds. What is the most likely cause?
Easy82Which THREE are valid methods to handle application configuration in AWS? (Choose three.)
Medium83A developer notices that an AWS Lambda function, which uses Amazon RDS Proxy to connect to an Aurora MySQL database, is experiencing increased latency and occasional connection timeouts. The function is configured with a reserved concurrency of 100 and is deployed in a VPC. The RDS Proxy's maximum connections is set to 1000. CloudWatch metrics show that the DatabaseConnections metric for the proxy is consistently at 1000. What is the most likely cause of the increased latency and timeouts?
Hard84A company's application uses Amazon S3 to store user-uploaded images. Users report that recently uploaded images are sometimes not immediately available for viewing. The application uses S3 Event Notifications to trigger a Lambda function that processes images and stores metadata in DynamoDB. What is the MOST likely cause of the delay?
Medium85A developer is troubleshooting a slow Amazon RDS MySQL database query. The query is frequently executed and takes 5 seconds to complete. Which AWS service should the developer use to analyze the query performance?
Easy86An AWS Lambda function processes messages from an Amazon SQS queue and writes results to an Amazon DynamoDB table. The function is configured with a reserved concurrency of 5 and a batch size of 10. CloudWatch metrics show high throttling and a growing queue backlog. The function's execution time averages 1 second per message. What is the MOST effective action to reduce throttling while improving throughput?
Medium87A developer is troubleshooting an application that uses Amazon ElastiCache for Redis to improve performance. The application periodically experiences high latency during peak hours. The developer checks the ElastiCache metrics and sees that the 'Evictions' metric is consistently high and the 'CacheHitRate' metric is low. The cluster has a single node with a cache.t3.small instance type. Which action will most likely improve the cache hit rate and reduce latency?
Medium88A developer is using AWS X-Ray to trace requests through a microservices application. The developer notices that some traces are incomplete. Which TWO actions can help ensure complete traces?
Easy89A Lambda function using a Kinesis event source repeatedly retries one bad record and blocks progress in the shard. Which feature helps isolate failed records after retry limits?
Hard90A company is using AWS CodePipeline for CI/CD. The pipeline has a build stage using AWS CodeBuild, and a deploy stage using AWS CodeDeploy. The deployment is failing with 'Error: Health checks failed'. Which TWO steps should the developer take to troubleshoot this issue? (Select TWO.)
Hard91A company runs a critical application on Amazon EC2 instances behind an Application Load Balancer (ALB). The application experiences intermittent errors where some requests return HTTP 503 (Service Unavailable) errors. The developers have verified that the application code is healthy and the EC2 instances pass health checks. The ALB health check is configured to hit a specific endpoint (/health) with a healthy threshold of 2 and an unhealthy threshold of 2. The health check interval is 30 seconds, and the timeout is 5 seconds. The application's /health endpoint sometimes takes up to 6 seconds to respond due to a dependency on a third-party service. The developers want to minimize the 503 errors without changing the application code. Which action should the developer take?
Hard92Which THREE factors should a developer consider when designing a stateless application on AWS? (Choose 3)
Medium93A developer is troubleshooting a web application that intermittently returns HTTP 504 errors. The application runs on EC2 instances behind an Application Load Balancer. What is the most likely cause of these errors?
Easy94A developer is using an Amazon SQS queue with a Lambda function as a consumer. Messages are being sent to the queue but the Lambda function is not processing them. Which THREE of the following are possible causes?
Easy95A developer is deploying a serverless application using AWS CloudFormation. The stack creation fails with the error 'The following resource(s) failed to create: [MyLambdaFunction]'. The developer checks the CloudWatch logs but finds no logs for the Lambda function. What is the most likely reason?
Easy96An application running on Amazon EC2 instances behind an Application Load Balancer (ALB) is experiencing intermittent 503 errors. The EC2 instances are in an Auto Scaling group. What is the MOST likely cause?
Easy97A developer invokes a Lambda function from the AWS CLI and receives the response shown in the exhibit. The output file contains an error message. What is the MOST likely cause of the FunctionError field being set to 'Unhandled'?
Medium98A developer is troubleshooting an application that uses Amazon SQS. Messages are being sent to a dead-letter queue (DLQ) after the maximum receive count is exceeded. The consumer processes messages but sometimes fails. The developer wants to ensure that messages are retried immediately after a failure, without waiting for the visibility timeout. Which solution should the developer implement?
Hard99A developer is working on a serverless application that uses AWS Lambda functions to process user uploads. The uploads are stored in an S3 bucket, and each upload triggers a Lambda function that resizes images and stores metadata in DynamoDB. Recently, users have reported that some images are not being resized. The developer checks the CloudWatch logs and sees that the Lambda function is invoked, but it fails with a timeout error after 15 seconds for a few large images. The function has a timeout of 15 seconds and a memory of 512 MB. The image sizes vary from 1 MB to 50 MB. The developer wants to handle large images without increasing the timeout significantly, as that would increase costs. The function is CPU-bound during image processing. Which solution should the developer implement?
Medium100A developer notices that an RDS MySQL instance's CPU utilization is consistently above 80% during peak hours. Which AWS service can be used to analyze the database queries and identify the root cause?
Easy101A company has a REST API deployed on Amazon API Gateway with a Lambda integration. The API is experiencing high latency. Which TWO actions would help diagnose the issue?
Hard102Refer to the exhibit. A CloudFormation stack creation failed. What is the most likely cause of the failure?
Hard103A developer is troubleshooting an AWS Elastic Beanstalk environment that is failing health checks. The environment runs a web application on Tomcat. The developer checks the logs and finds no errors. What is the most likely cause of the health check failure?
Hard104A company runs a stateful web application on EC2 instances in an Auto Scaling group. Users report that their session data is lost when instances are replaced during scaling events. What is the best solution to preserve session state?
Hard105A developer is optimizing a DynamoDB table for a gaming leaderboard. The table stores player scores and is read-heavy. Queries often fetch the top 10 scores. Which indexing strategy best reduces RCU consumption?
Medium106A developer attached the IAM policy above to an IAM user. The user reports being unable to list objects in the bucket 'my-bucket' using the AWS CLI command 'aws s3 ls s3://my-bucket/'. What is the most likely reason?
Hard107A developer is troubleshooting performance issues in an application that uses Amazon ElastiCache for Redis. The application experiences periodic latency spikes during peak hours. The developer checks CloudWatch metrics and sees that the 'Evictions' metric is consistently high and the 'CacheHitRate' metric is low. The cluster uses a single cache.t3.small node. Which action will most likely improve the cache hit rate and reduce latency?
Hard108A company's application uses Amazon DynamoDB as its database. The application reads the same item multiple times per second and occasionally sees stale data. The DynamoDB table uses the default eventually consistent reads. What should the developer change to ensure strongly consistent reads?
Medium109A developer is troubleshooting an EC2 instance that is unreachable via SSH. The instance is in a public subnet with a security group that allows inbound SSH from 0.0.0.0/0. Which THREE are possible causes? (Choose 3.)
Hard110Which TWO services can be used to store and retrieve application configuration data in AWS? (Choose 2)
Easy111Users receive AccessDenied when downloading SSE-KMS encrypted S3 objects cross-account. Which two policies may need changes?
Medium112A developer is using AWS CodeBuild to build a Java application. The build fails with 'OutOfMemoryError'. Which configuration change would most likely resolve this issue?
Medium113A DynamoDB application receives ProvisionedThroughputExceededException during predictable daily peaks. The workload is not cacheable. What should be changed?
Medium114A company runs a web application on EC2 instances in an Auto Scaling group. The application needs to store session state. The architecture must be highly available and scalable. Which solution should the developer choose?
Medium115A developer is using AWS CodeBuild to build a Docker image and push it to Amazon ECR. The build fails with the error 'no basic authentication credentials'. The build project has an IAM role with the AmazonEC2ContainerRegistryPowerUser policy. What is the most likely cause?
Hard116A developer is using AWS CloudFormation to deploy a stack that includes an S3 bucket and a Lambda function. The stack fails with the error 'The following resource(s) failed to create: [MyBucket]'. What is the most likely cause?
Medium117A developer notices that an Amazon RDS for MySQL DB instance's CPU utilization is consistently above 90% during peak hours. The application uses read-heavy workloads. Which action would MOST effectively reduce CPU load without major architectural changes?
Hard118A developer is deploying a new version of a Lambda function and wants to roll back immediately if errors are detected. Which deployment strategy should the developer use?
Easy119A developer is using AWS X-Ray to trace a Lambda function that calls DynamoDB and SQS. Some traces show errors. Which TWO actions should the developer take to diagnose the issue?
Hard120Drag and drop the steps to authenticate a user using Amazon Cognito User Pools in the correct order.
Medium121A developer is troubleshooting an AWS CloudFormation stack creation failure. The stack creation failed with the error: 'Resource creation cancelled'. What does this error typically indicate?
Easy122A developer is troubleshooting an Amazon API Gateway REST API that returns 504 Gateway Timeout errors for certain requests. The backend is a Lambda function that performs a resource-intensive operation that occasionally takes up to 30 seconds. API Gateway has a default integration timeout of 29 seconds. The developer cannot reduce the execution time. What should the developer do to resolve the timeout issue?
Hard123A developer receives an AccessDenied error when trying to put an object into an S3 bucket using the AWS SDK. The IAM user has an attached policy that grants s3:PutObject on the bucket. What is the MOST likely cause of the error?
Medium124A developer monitors an AWS Lambda function that processes records from an Amazon SQS queue and writes results to an Amazon DynamoDB table. CloudWatch Logs show that execution time has increased over the past week, and the function frequently times out at the 5-minute timeout. The function's code has not been changed recently. CloudWatch metrics show a high rate of DynamoDBProvisionedThroughputExceededException errors. The DynamoDB table has 5 write capacity units (WCUs). What action will MOST effectively reduce the function's execution time?
Medium125A developer is debugging an AWS Lambda function that is invoked by an Amazon S3 bucket notification. The function fails with an 'AccessDenied' error when trying to read an object from the same bucket. What should the developer check first?
Easy126A developer is troubleshooting an Amazon RDS for MySQL instance that is experiencing high CPU utilization. The application performs many read operations. The developer wants to reduce the load on the database. What is the MOST effective solution?
Easy127A developer is troubleshooting a slow-running Amazon RDS for MySQL query. The query performance has degraded over time. Which approach should the developer take first to identify the cause?
Medium128A developer is troubleshooting a slow-running query on an Amazon RDS for MySQL database. The query is used by a reporting application and takes over 30 seconds to complete. The database is a db.r5.large instance with 200 GB of gp2 storage. Which TWO actions should the developer take to improve query performance?
Medium129A developer notices that an S3 bucket policy allows public read access to all objects. The bucket contains sensitive data that should only be accessible by authorized IAM users. What is the BEST way to remediate this?
Easy130A company uses AWS CloudFormation to deploy a stack that includes an RDS MySQL instance. During an update, the stack fails with a 'DELETE_FAILED' status on a security group resource. The security group has a dependency on the RDS instance. What is the MOST likely cause?
Hard131A developer is using AWS CloudFront to serve static content. Users in some geographic regions report slow load times. Which CloudFront feature can the developer use to reduce latency for these users?
Easy132A developer is deploying a new version of a Lambda function using an alias for blue/green deployment. Traffic is gradually shifted to the new version. During the shift, a high error rate is observed. What should the developer do to minimize impact?
Easy133A developer is managing an application running on Amazon EC2 instances behind an Application Load Balancer. Users report that the application becomes unresponsive after several hours, and restarting the instance temporarily fixes the issue. The developer suspects a memory leak but cannot add custom instrumentation. Which AWS service can collect memory utilization metrics and help identify the memory leak with minimal configuration?
Medium134A developer is troubleshooting a DynamoDB table that is experiencing high write throttling (ProvisionedThroughputExceededException) on certain days. The table has provisioned write capacity of 1000 WCU. The table has a partition key of 'user_id' which is a UUID. The table is accessed by multiple services. CloudWatch metrics show that the WriteThrottleEvents are spiking during specific hours, and the ConsumedWriteCapacityUnits often reaches 1000. What is the most likely cause of the throttling?
Medium135A company is using Amazon API Gateway to expose a REST API. The API is integrated with an AWS Lambda function. Lately, the API is returning 502 Bad Gateway errors. What is the MOST likely cause?
Medium136A developer notices that an AWS Lambda function, configured to access an Amazon RDS database in the same VPC, is timing out. The function has a 30-second timeout. CloudWatch Logs show that the function starts execution but never reaches the database. The VPC configuration includes private subnets without a NAT gateway. The RDS database is in the same VPC. What is the most likely cause of the timeout?
Medium137A web application runs on Amazon EC2 instances behind an Application Load Balancer (ALB). During rolling updates of the Auto Scaling group, users intermittently receive HTTP 502 (Bad Gateway) errors. The developer checks the ALB access logs and notices that requests are being routed to instances that are in the 'Draining' state. The ALB has connection draining enabled with a timeout of 30 seconds. The Auto Scaling group terminates instances after they are taken out of service. What is the most likely cause of the 502 errors?
Hard138A developer is using Amazon RDS for MySQL and notices that the database performance has degraded. The developer suspects that slow queries are the cause. Which THREE actions should the developer take to identify and address the slow queries?
Easy139A developer notices that the Lambda function 'my-function' is not generating any logs in CloudWatch, although the function is invoked successfully. The developer runs the command above. What is the MOST likely cause?
Hard140A developer runs a script that uses the AWS CLI to copy a large number of files from an on-premises server to an S3 bucket. The copy operation fails partway through with a 'RequestTimeout' error. What is the MOST efficient way to resume the copy and ensure all files are transferred?
Easy141An application running on Amazon EC2 generates logs that need to be streamed to Amazon CloudWatch Logs. The developer installs and configures the CloudWatch agent. However, logs are not appearing in the log group. What is the most likely cause?
Easy142A developer is deploying a new version of an application to Amazon ECS using the Fargate launch type. The task fails to start and the error message indicates that the task cannot pull the container image from Amazon ECR. What is the MOST likely cause?
Easy143A company's DynamoDB table has a read capacity of 10,000 RCUs and receives consistent traffic. Recently, users have reported increased latency for read requests. The application uses strongly consistent reads. The developer checks CloudWatch metrics and sees that 'ConsumedReadCapacityUnits' is at 9,500 but 'ThrottledRequests' is high. What is the most likely cause?
Hard144A company uses AWS CodePipeline with CodeBuild to test and deploy a web application. The pipeline has been failing at the deploy stage with an error: 'Access Denied'. CloudTrail shows the CodePipeline service role is making the call. What is the MOST likely cause?
Medium145Messages in an SQS queue are processed successfully but later reappear and are processed again. What is the most likely configuration issue?
Hard146Which TWO actions can help reduce Lambda cold start times? (Choose two.)
Easy147A company runs a serverless application on AWS using API Gateway, AWS Lambda, and DynamoDB. The application processes user uploads and stores metadata in DynamoDB. Recently, users have reported that some uploads fail with a 500 Internal Server Error. The CloudWatch Logs for the Lambda function show 'ProvisionedThroughputExceededException' errors for DynamoDB, followed by 'Task timed out after 3.00 seconds' errors. The Lambda function has a 3-second timeout and 128 MB of memory. The DynamoDB table has 5 read capacity units and 5 write capacity units. The application uses a single Lambda function that processes each upload synchronously. The company expects a steady increase in uploads. Which combination of actions should a developer take to resolve the errors and prepare for future growth? (Choose TWO.)
Hard148A developer is troubleshooting an AWS CloudFormation stack that failed to create. The error message says 'The following resource(s) failed to create: [MyEC2Instance]'. What is the first step the developer should take?
Medium149A company's application running on Amazon ECS Fargate is experiencing high CPU utilization. The task definition has CPU set to 256 units. What should be done to improve performance?
Medium150A developer is troubleshooting a Lambda function that intermittently times out. The function makes HTTP requests to an external API. The function's CloudWatch logs show 'Task timed out after 3.01 seconds'. What is the MOST likely cause?
Medium151Refer to the exhibit. An IAM policy is attached to a user. The user tries to download an object from s3://my-bucket/secret/config.txt. What will happen?
Hard152Which TWO actions can help reduce latency for a web application hosted on EC2 instances behind an Application Load Balancer? (Select TWO.)
Medium153A developer has deployed a serverless application using AWS SAM. After a recent update, the API Gateway endpoints return 500 errors. The Lambda function logs show no errors. What should the developer investigate first?
Medium154A developer is troubleshooting an application that uses Amazon ElastiCache for Redis to cache database query results. The application experiences high latency during cache misses. The developer notices that frequently accessed keys (hot keys) are often missing from the cache, suggesting they are being evicted. Which action should the developer take to reduce cache misses for hot keys?
Medium155A developer is troubleshooting an AWS Lambda function that processes large CSV files (up to 1 GB) uploaded to an Amazon S3 bucket. The function uses Python and the pandas library to perform data transformations. Recently, the function started timing out on large files. CloudWatch Logs show that the function's execution time is close to the 15-minute Lambda timeout, and memory utilization peaks at around 80% of the configured 3,008 MB. The function has not been modified in months. Which action will most likely resolve the timeout issue without requiring code changes?
Hard156A developer is troubleshooting a slow-performing Amazon RDS for MySQL database. Which TWO actions should the developer take to improve query performance?
Medium157A developer is troubleshooting a CloudFormation stack that fails to create. The stack includes an Auto Scaling group with a launch template. The error message says 'Value (null) for parameter groupId is invalid.' What is the MOST likely cause?
Medium158The exhibit shows the output of invoking a Lambda function from the AWS CLI. The function returned a status code of 200 but included a FunctionError field set to 'Unhandled'. What does this indicate?
Medium159A developer is using AWS Elastic Beanstalk to deploy a web application. The application is experiencing high latency. Which TWO steps should the developer take to troubleshoot and optimize the application?
Medium160A developer is troubleshooting an AWS Lambda function that processes records from an Amazon Kinesis Data Stream. The function is configured with a batch size of 100 and a parallelization factor of 1. The iterator age metric is increasing, and CloudWatch Logs show the function execution time is around 4 minutes (timeout is 5 minutes). The stream has 10 shards. What is the most cost-effective way to increase processing throughput?
Medium161A developer is troubleshooting an AWS Lambda function that is timing out. The function has a timeout of 5 seconds and is configured with 128 MB of memory. Which TWO of the following are effective ways to resolve the timeout?
Medium162A developer is running an AWS Lambda function that is triggered by Amazon S3 events. The function writes processed data to an Amazon DynamoDB table. Over time, the function's execution time has increased significantly. CloudWatch Logs show many DynamoDBProvisionedThroughputExceededException errors. The table is configured with 5 read capacity units (RCUs) and 5 write capacity units (WCUs). The function performs both reads and writes. Which optimization will MOST effectively reduce throttling errors while maintaining performance?
Hard163A developer is using Amazon S3 to host a static website. The website returns 403 Forbidden errors. The bucket policy allows public read access. What is the most likely cause?
Easy164Refer to the exhibit. A developer runs the AWS CLI command for an EC2 instance. The instance is in the 'running' state, but the application hosted on it is not reachable. What should the developer check first?
Medium165A developer ran the above CLI command to describe an EC2 instance. The instance is running but the developer cannot connect to it via SSH. Which additional step should the developer take to troubleshoot the connectivity issue?
Hard166Which TWO approaches can be used to optimize costs for an Amazon DynamoDB table with predictable read/write patterns? (Select TWO.)
Easy167A developer has deployed an AWS Lambda function that is triggered by an Amazon S3 event. The function processes image files and stores metadata in an Amazon DynamoDB table. CloudWatch metrics show that the function's error count has increased. The developer checks CloudWatch Logs and sees errors related to insufficient memory. The function is configured with 128 MB of memory. What should the developer do to resolve the errors?
Medium168A developer invoked a Lambda function and saw the above output. What is the root cause of the error?
Easy169A development team is using AWS CodeCommit as a source repository and CodeBuild for build automation. They want to trigger a build automatically whenever a pull request is created or updated in the repository. Which configuration should they use?
Medium170A developer is optimizing an API Gateway REST API that uses Lambda integration. The response times are high, and CloudWatch logs show that the Lambda function has cold starts frequently. The function is written in Java and uses a large library. What is the MOST effective optimization?
Hard171Match each AWS deployment strategy to its description.
Medium172A developer deployed a new version of a Lambda function that processes S3 events. After deployment, some S3 events are not being processed. The CloudWatch Logs show no errors. What is the most likely cause?
Medium173A developer is troubleshooting a slow-running application that uses ElastiCache for Redis as a caching layer. The application frequently reads and writes data to the cache. Which TWO actions should the developer take to improve cache performance?
Medium174A developer is troubleshooting a slow Amazon DynamoDB table. The table has a read capacity of 1000 RCU and a write capacity of 500 WCU. The application frequently reads the same item. Which TWO actions can improve read performance?
Medium175A company runs a Node.js application on AWS Elastic Beanstalk. The application is experiencing high latency. The developer suspects the database queries are slow. Which step should the developer take first to diagnose the issue?
Medium176A developer is troubleshooting a CloudFront distribution that serves static content from an S3 bucket. Users in some geographic locations report slow load times. The developer checks the CloudFront metrics and sees a high number of cache misses. What is the MOST likely cause?
Medium177An application running on Amazon EC2 instances behind an Application Load Balancer (ALB) is experiencing increased latency. The developer suspects the ALB is the bottleneck. How can the developer confirm this using CloudWatch metrics?
Medium178A developer is using Amazon CloudFront to serve static content from an S3 bucket. Users are reporting that they see outdated content. The CloudFront distribution has a default TTL of 24 hours. What is the MOST efficient way to serve updated content immediately?
Medium179A developer uses the CloudFormation template in the exhibit to create an S3 bucket. The stack creation fails with the error 'Bucket already exists'. What is the MOST likely reason?
EasyOther domains
All DVA-C02 exam domains
Frequently asked questions
- What does the Troubleshooting and Optimization domain cover on the DVA-C02 exam?
- Be able to inspect IAM roles, VPC egress, log group configuration, and event delivery semantics to isolate the root cause, then apply the smallest fix. The single most important skill is mapping a symptom to the specific AWS configuration that produces it.
- How many questions are in this domain?
- This page lists all 179 Troubleshooting and Optimization questions in the DVA-C02 question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Troubleshooting and Optimization questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.