Courseiva

CCNA Incident and Event Response Questions

75 of 183 questions · Page 2/3 · Incident and Event Response · Answers revealed

76
Multi-Selectmedium

A company uses AWS CodePipeline for CI/CD. A recent pipeline execution failed at the 'Deploy' stage with the error 'Action execution failed: Access Denied'. The pipeline uses an IAM service role. Which THREE checks should the engineer perform to resolve this?

Select 3 answers
A.Check that CloudWatch Events rule is configured to trigger the pipeline.
B.Verify that the IAM service role has sufficient permissions to perform the deploy action on the target resource.
C.Ensure the artifact store S3 bucket has a bucket policy that allows the pipeline role to access it.
D.Enable S3 event notifications to trigger the pipeline on code changes.
E.Confirm that the service role's trust policy allows CodePipeline to assume the role.
AnswersB, C, E

CodePipeline executes deploy actions by assuming a dedicated IAM service role, and that role must contain an identity-based policy granting the required API permissions (e.g., codedeploy:CreateDeployment, ecs:UpdateService, or cloudformation:CreateStack) on the target resource. If a required permission is missing or explicitly denied, the deploy stage fails with an AccessDenied error even though every other pipeline configuration is correct. You should review the pipeline execution details to identify the specific denied action, then attach an appropriate managed or inline policy to the service role.

Why this answer

The 'Access Denied' error at the Deploy stage indicates a permissions issue. The correct checks are: B) Verify that the IAM service role has sufficient permissions for the deploy action on the target resource, as the role must have the necessary IAM policies to perform deployments. C) Ensure the artifact store S3 bucket has a bucket policy that allows the pipeline role to access it, because the pipeline needs to read artifacts from the bucket.

E) Confirm that the service role's trust policy allows CodePipeline to assume the role, since the trust policy must grant the `sts:AssumeRole` permission to the CodePipeline service. Option A is incorrect because CloudWatch Events rules trigger pipeline execution but are not related to the deploy stage's access denied error. Option D is incorrect because S3 event notifications trigger the pipeline on code changes, not resolve deployment failures.

77
MCQeasy

A DevOps engineer receives an alarm that an EC2 instance's CPU utilization has exceeded 90% for 5 minutes. The engineer needs to automatically recover the instance. Which AWS service should be used to configure automatic recovery?

A.Amazon CloudWatch Alarms
B.AWS Lambda
C.AWS Systems Manager Automation
D.EC2 Auto Scaling
AnswerA

Amazon CloudWatch Alarms are the native mechanism for EC2 AutoRecovery. By configuring an alarm on the System Status Check metric (StatusCheckFailed_System) with the 'recover' action, you let EC2 automatically restart the instance on new hardware while preserving its instance ID, private IP, Elastic IP, and instance store data. This is the direct, built-in solution that requires no custom code or additional orchestration.

Why this answer

Amazon CloudWatch Alarms can be configured to trigger an EC2 instance recovery action when a metric like CPU utilization exceeds a threshold (e.g., 90% for 5 minutes). The alarm sends a signal to the EC2 service, which automatically recovers the instance by stopping it and starting it on a new underlying host, preserving the instance ID, private IP, and Elastic IP. This is the native, built-in mechanism for automatic instance recovery without requiring additional compute or orchestration services.

Exam trap

The trap here is that candidates often confuse EC2 Auto Scaling (which replaces instances) with automatic recovery (which recovers the same instance), or they overcomplicate the solution by choosing Lambda or Systems Manager when a simple CloudWatch Alarm action is the correct and native AWS mechanism.

How to eliminate wrong answers

Option B is wrong because AWS Lambda is a serverless compute service that can execute custom code in response to events, but it is not the direct service used to configure automatic EC2 instance recovery; while Lambda could be used to script a recovery, it adds unnecessary complexity and latency compared to the native CloudWatch Alarm recovery action. Option C is wrong because AWS Systems Manager Automation provides runbooks for automated remediation and operational tasks, but it is not the primary service for configuring automatic EC2 instance recovery; it would require additional setup and is not the simplest or recommended approach. Option D is wrong because EC2 Auto Scaling is designed to manage the number of instances in an Auto Scaling group based on scaling policies, not to recover a specific impaired instance; it would terminate and replace the instance rather than recover it, which changes the instance ID and associated resources.

78
Multi-Selectmedium

A company is designing an incident response strategy for its Amazon EKS cluster. Which THREE steps should be taken to ensure rapid response to a compromised pod?

Select 3 answers
A.Delete the entire namespace to ensure all resources are removed.
B.Scale down the deployment to 0 replicas.
C.Delete the pod using kubectl delete pod.
D.Use kubectl exec to gather forensic data from the pod before termination.
E.Apply a Kubernetes NetworkPolicy to deny all ingress/egress traffic to the compromised pod.
AnswersC, D, E

Deleting the pod with `kubectl delete pod <name>` immediately sends a termination signal (SIGTERM) to the container, and after the grace period, a SIGKILL, removing the pod from the cluster and stopping the malicious process. This is the primary containment step to halt active compromise; however, if the pod is managed by a ReplicaSet or Deployment, the controller will replace it, so you must also update the workload definition or scale down to prevent recreation. It is a precise operation that affects only the compromised instance.

Why this answer

Deleting the pod with `kubectl delete pod` immediately terminates the compromised container, stopping any malicious activity. This is a rapid containment step that removes the pod from the cluster without affecting the broader deployment or namespace, allowing the incident response team to investigate and remediate without unnecessary disruption.

Exam trap

The trap here is that candidates may think scaling down the deployment (Option B) is the fastest containment action, but it actually triggers a reconciliation loop that can recreate the pod or delay termination, whereas `kubectl delete pod` is the most direct and immediate way to stop a compromised container.

79
Multi-Selectmedium

A company runs a critical application on EC2 instances behind an Application Load Balancer (ALB) in an Auto Scaling group. The team wants to automate the response to an instance failure. Which THREE steps should be taken to ensure automatic recovery and notification?

Select 3 answers
A.Create a CloudWatch alarm to terminate the instance
B.Configure Auto Scaling to replace unhealthy instances
C.Configure the ALB health check to mark instances as unhealthy
D.Set up Amazon SNS notifications for Auto Scaling events
E.Create a scaling policy based on CPU utilization
AnswersB, C, D

Auto Scaling can be configured with an appropriate health check type (EC2 status checks or ELB/ALB health checks) and a termination policy to automatically detect and replace unhealthy instances. This ensures that the ASG maintains the desired instance count by terminating the unhealthy instance and launching a new one, providing a self-healing architecture for the critical application.

Why this answer

Configuring the Auto Scaling group to replace unhealthy instances ensures that when an instance fails health checks, Auto Scaling automatically terminates it and launches a new instance to maintain the desired capacity. This is the core mechanism for automated recovery in an Auto Scaling group, as it directly responds to instance failure without manual intervention.

Exam trap

The trap here is that candidates often confuse CloudWatch alarms (Option A) for instance recovery, but CloudWatch alarms are for monitoring and triggering actions like scaling policies or SNS notifications, not for directly replacing failed instances in an Auto Scaling group.

80
MCQeasy

A DevOps engineer is troubleshooting a Lambda function that times out after 3 seconds. The function makes an HTTP request to an external API. The function's timeout setting is 10 seconds. What is the most likely cause of the timeout?

A.The external API is throttling the request.
B.The HTTP client's timeout is set to a low value.
C.The Lambda function is not in a VPC.
D.The Lambda function's memory is insufficient.
AnswerB

The default inactivity timeout for many HTTP clients is exactly 3 seconds, so when the external API takes longer than that to respond, the client aborts with a timeout error before the Lambda function's own timeout limit is reached. Raising the client timeout (or configuring connection/read timeouts) to exceed the expected API response time directly resolves the issue. This matches the symptom precisely because every call that exceeds the threshold fails at the same 3-second mark.

Why this answer

The Lambda function times out after 3 seconds, which is most likely due to an HTTP client timeout explicitly set to a low value (e.g., 3 seconds) inside the function code. Even though the Lambda function's configured timeout is 10 seconds, the HTTP client's internal timeout fires first, causing the function to fail before the Lambda timeout is reached. This is the most likely cause because the symptom matches a client-side timeout rather than a server-side or infrastructure issue.

Exam trap

The trap here is that candidates assume the Lambda function's configured timeout (10 seconds) is the only timeout that matters, overlooking that application-level timeouts (like HTTP client timeouts) can fire independently and cause earlier failures.

How to eliminate wrong answers

Option A is wrong because external API throttling would typically return an HTTP 429 status code or cause a slower response, not a consistent 3-second timeout; the function would still wait for the response up to the Lambda timeout. Option C is wrong because not being in a VPC actually reduces network latency and complexity (Lambda uses the public internet by default), so it would not cause a timeout; VPC-related timeouts usually occur when the function is in a VPC without a NAT gateway or proper routing. Option D is wrong because insufficient memory would cause out-of-memory errors or slower execution, not a consistent 3-second timeout; memory affects CPU allocation but does not directly trigger a timeout at a specific second.

81
Multi-Selectmedium

A DevOps team is investigating a security incident where an unauthorized user accessed an S3 bucket. The team needs to determine what actions were taken by the user. Which TWO AWS services should be used together to investigate? (Choose TWO.)

Select 2 answers
A.S3 server access logs
B.Amazon CloudWatch metrics
C.AWS Config
D.AWS CloudTrail
E.Amazon GuardDuty
AnswersA, D

S3 server access logs record every request made to a bucket, capturing the authenticated identity, timestamp, source IP and operation performed. This directly satisfies the requirement to determine which actions the unauthorised user took, since the logs detail the specific REST operations against the bucket.

Why this answer

S3 server access logs (A) are correct because they record detailed, object-level requests made against a bucket, including the requester, bucket name, request time, action (such as REST.GET.OBJECT or REST.PUT.OBJECT), response status, and error code, which lets the team see exactly what operations the unauthorized user performed on the objects. AWS CloudTrail (D) is correct because it captures S3 data events (GetObject, PutObject, DeleteObject) and management events (such as PutBucketPolicy) as API activity with the identity of the caller, source IP address, and timestamp, providing the who-did-what audit trail needed to attribute the actions. Used together, CloudTrail identifies the principal and API calls while S3 server access logs provide the granular request-level detail for the bucket.

Amazon CloudWatch metrics (B) only provides aggregate performance and usage statistics, not per-request identity or action detail, so it cannot show what the user did. AWS Config (C) tracks resource configuration changes and compliance over time, not individual object access actions. Amazon GuardDuty (E) is a threat-detection service that generates findings about suspicious activity but does not provide the raw request-level or API-level audit records required to reconstruct the user's actions.

82
Multi-Selecthard

A company is experiencing a DDoS attack on its application hosted on AWS. The application uses an Application Load Balancer (ALB) with an Auto Scaling group of EC2 instances. The security team needs to mitigate the attack with minimal latency impact on legitimate users. Which THREE actions should the team take? (Choose THREE.)

Select 3 answers
A.Disable cross-zone load balancing on the ALB to limit the number of instances receiving traffic.
B.Enable AWS Shield Advanced on the ALB.
C.Enable connection draining (deregistration delay) on the ALB target group.
D.Configure the Auto Scaling group to scale based on the NetworkIn metric to handle the increased traffic.
E.Configure AWS WAF on the ALB to block requests based on source IP reputation or rate-based rules.
AnswersB, C, E

Shield Advanced provides always-on detection and inline mitigation at the AWS edge, absorbing large volumetric DDoS attacks before they reach the ALB. It also includes DDoS cost protection and 24/7 access to the DDoS Response Team (DRT), who can analyze attack patterns and apply mitigations directly. Enabling it on the ALB is a fundamental step for comprehensive DDoS protection.

Why this answer

AWS Shield Advanced provides enhanced DDoS mitigation for ALBs, including access to the DDoS Response Team (DRT) and financial protection against scaling costs. It operates at the network and transport layers with minimal latency, as it inspects traffic inline without introducing significant processing delay. This makes it a critical first line of defense for high-availability applications under attack.

Exam trap

The trap here is confusing scaling-based absorption (Option D) with actual mitigation, leading candidates to think handling more traffic automatically defends against DDoS, when in reality it only increases cost and resource exhaustion without blocking the attack source.

83
MCQeasy

A DevOps engineer is investigating why an Amazon ECS service is not scaling out as expected. The service has a target tracking scaling policy based on average CPU utilization. The CloudWatch alarm shows that CPU utilization has exceeded the target for several minutes, but no scaling activity has occurred. What is the most likely cause?

A.The ECS service is configured with a minimum healthy percent that prevents scaling out.
B.The ECS service does not have an IAM role that allows it to call CloudWatch.
C.The CloudWatch alarm is configured with a period that is too long.
D.The scaling policy has a cooldown period that is still in effect from a previous scaling activity.
AnswerD

A cooldown period is a duration after a scaling activity during which the scaling policy cannot trigger another action, even if the CloudWatch alarm remains in breach. In Application Auto Scaling, this prevents rapid oscillation and lets the newly added tasks take effect. If a previous scale-out or scale-in occurred recently, the outstanding cooldown will block the new scaling request, which exactly matches the symptom of the alarm being breached but no scaling activity occurring.

Why this answer

Target tracking scaling policies in Amazon ECS have a cooldown period (default 300 seconds) that prevents the policy from initiating additional scaling activities immediately after a previous scaling action. If a recent scaling activity occurred, the cooldown period would still be in effect, causing the policy to ignore the alarm even though CPU utilization has exceeded the target. This is the most likely reason no scaling activity is observed despite the alarm being triggered.

Exam trap

The trap here is that candidates often overlook the cooldown period and instead blame IAM permissions or alarm configuration, but the cooldown is a deliberate stabilization mechanism that directly explains why scaling is not occurring despite the alarm being active.

How to eliminate wrong answers

Option A is wrong because the minimum healthy percent parameter controls the minimum percentage of tasks that must remain healthy during deployments or scaling activities, but it does not prevent scaling out; it only affects how many tasks can be stopped or started at once. Option B is wrong because the ECS service does not need an IAM role to call CloudWatch; the scaling policy uses the service-linked role AWSServiceRoleForApplicationAutoScaling, which already has permissions to read CloudWatch alarms and metrics. Option C is wrong because a long CloudWatch alarm period would delay the alarm transition to ALARM state, but the question states the alarm has already exceeded the target for several minutes, meaning the period is not the issue.

84
MCQmedium

A company uses an AWS Elastic Load Balancer (ELB) to distribute traffic to EC2 instances. During an incident, some users report slow response times. The DevOps engineer suspects that one instance is unhealthy but the health check is not detecting it. What should the engineer do to improve health check accuracy?

A.Increase the health check interval to reduce load on instances
B.Configure a health check that checks a specific application endpoint
C.Decrease the health check unhealthy threshold to 2
D.Set the health check path to the root document ('/')
AnswerB

Configuring a health check against a specific application endpoint, such as /healthz or /api/v1/status, makes the ELB perform a deep health check that exercises your application's critical request path, including framework routing, connections to dependent services, and session or cache state. This directly addresses the scenario where an instance accepts TCP or HTTP requests but is not truly ready to serve application traffic. By returning a non-200 status only when application dependencies or internal state are degraded, the ELB can accurately detect the unhealthy instance and stop sending it traffic.

Why this answer

A health check that targets a specific application endpoint (e.g., /health or /api/status) validates that the application is actually functioning, not just that the instance responds on a port. This catches unhealthy instances that still accept TCP connections but fail to serve requests, which is exactly the scenario described. The other options either weaken detection or do not improve accuracy.

Exam trap

DOP-C02 often tests the difference between port-level and application-level health checks — candidates pick '/' or threshold tweaks because they sound like fixes, but only a specific application endpoint validates true health.

How to eliminate wrong answers

Option A is wrong because increasing the interval slows detection of failures, making the problem worse rather than improving accuracy. Option C is wrong because lowering the unhealthy threshold to 2 makes the check more aggressive but does not make it more accurate — it can cause false positives without validating application health. Option D is wrong because checking '/' only confirms the web server responds; it does not verify application-level health and may return 200 even when the app is degraded.

85
MCQhard

A DevOps engineer updates an ECS service via CloudFormation. The stack update fails with the message 'Resource update cancelled'. The engineer notices that the ECS service's desired count was temporarily reduced during the update. What is the most likely cause of the failure?

A.The ECS service's minimum healthy percent was set to 100, causing the desired count reduction to zero to be rejected.
B.The ECS service's target group had an unhealthy instance that prevented the deregistration.
C.The ECS service deployment circuit breaker was triggered due to a timeout.
D.The ECS service did not have the required IAM role to call ecs:UpdateService.
AnswerA

During a rolling update, CloudFormation temporarily sets the ECS service's desired count to zero to force a clean replacement of all tasks. If the service's minimumHealthyPercent is set to 100%, ECS cannot scale the current running tasks below 100% of the desired count, so the service scheduler rejects the desired count reduction and the CloudFormation update is cancelled. This is a common misconfiguration when engineers expect zero-downtime but inadvertently block the scale-in step.

Why this answer

The error 'Resource update cancelled' occurs because CloudFormation detected that the ECS service update was not progressing as expected. When the minimum healthy percent is set to 100, the deployment process cannot reduce the desired count to zero (or below the current running count) without violating the requirement that 100% of the tasks remain healthy. This causes the update to be cancelled as CloudFormation waits indefinitely for the deployment to complete, eventually timing out and rolling back.

Exam trap

The trap here is that candidates often confuse the 'Resource update cancelled' error with a permissions or circuit breaker issue, but the key clue is the temporary reduction in desired count, which directly points to a minimum healthy percent constraint that prevents the service from scaling down to zero.

How to eliminate wrong answers

Option B is wrong because an unhealthy instance in the target group would cause health check failures and potential deployment issues, but it would not directly cause a 'Resource update cancelled' error with a temporary desired count reduction; the error message specifically points to a deployment configuration issue, not a target group health issue. Option C is wrong because the deployment circuit breaker is a feature that rolls back a deployment when it detects a failure (e.g., tasks failing to start), but it would not cause the desired count to be temporarily reduced; the circuit breaker triggers after a deployment failure, not as a cause of the count reduction. Option D is wrong because a missing IAM role for ecs:UpdateService would result in an authorization error (e.g., 'AccessDenied') during the update, not a 'Resource update cancelled' error with a temporary desired count reduction; the update would fail immediately with a permissions error, not after a partial state change.

86
MCQhard

A DevOps team is configuring an Auto Scaling group for a web application behind an Application Load Balancer. The team wants to automatically replace instances that fail the health check. Which scaling policy should be used?

A.Target tracking scaling policy
B.Default health check replacement
C.Step scaling policy
D.Manual scaling
AnswerB

When the Auto Scaling group is configured with EC2 health checks or an attached Elastic Load Balancer's health checks, it automatically detects instances that become unhealthy. The default health check replacement behavior marks the failing instance as unhealthy, terminates it, and launches a replacement instance to restore desired capacity, all without manual intervention. This is the underlying mechanism that keeps the web tier healthy, making it the correct choice for replacing faulty instances.

Why this answer

The correct approach is not a scaling policy but the Auto Scaling group's built-in health check replacement functionality. When the Auto Scaling group is configured with ELB health checks, it automatically terminates and replaces instances that the Application Load Balancer marks unhealthy. No scaling policy is required for this behavior.

Exam trap

The trap is confusing scaling policies (which adjust capacity based on load metrics) with the automatic health check replacement, which is a default feature of Auto Scaling groups. Scaling policies are not involved in replacing unhealthy instances.

How to eliminate wrong answers

Option A is wrong because a target tracking scaling policy adjusts the desired capacity based on a metric (like CPU utilization or request count) to maintain a target value, not to replace individual failed instances. Option C is wrong because a step scaling policy adjusts capacity based on alarm breaches with defined step adjustments, which is designed for proactive scaling based on load changes, not for reactive replacement of unhealthy instances. Option D is wrong because manual scaling requires human intervention to adjust capacity and does not automatically replace failed instances, defeating the purpose of automated health check replacement.

87
MCQhard

A company experiences a security incident where an unauthorized user accessed an S3 bucket containing sensitive data. The DevOps team needs to identify the source IP address and user agent of the request. Which AWS service provides this information?

A.VPC Flow Logs
B.Amazon S3 server access logs
C.AWS CloudTrail
D.Amazon CloudWatch Logs
AnswerB

Amazon S3 server access logs provide detailed records for every request, including source IP address, user agent, requester, operation, and HTTP status. This is the correct service to identify the source IP and user agent of the unauthorized request.

Why this answer

Amazon S3 server access logs record detailed information about every request made to a bucket, including the requester's IP address, user agent, request time, and operation. This is the only AWS service in the list that captures the source IP and user agent for S3 object-level requests. CloudTrail records API management events but does not log data-plane GET/PUT requests by default.

Exam trap

DOP-C02 often tests the distinction between CloudTrail (API/management events) and S3 server access logs (data-plane request details including IP and user agent) — candidates frequently assume CloudTrail captures everything.

How to eliminate wrong answers

Option A is wrong because VPC Flow Logs capture IP traffic metadata at the ENI/subnet level (source/dest IP, ports, bytes) but not HTTP user agents or S3 request details. Option C is wrong because CloudTrail logs S3 management events (e.g., PutBucketPolicy) and, only if data events are explicitly enabled, object-level operations — but even then it does not capture user agent strings. Option D is wrong because CloudWatch Logs is a log aggregation service, not a source of S3 request metadata; it only contains what you route to it.

88
MCQmedium

A company uses an Auto Scaling group with a dynamic scaling policy based on the average CPU utilization of the instances. During an incident, the DevOps team notices that the Auto Scaling group is not launching new instances quickly enough to handle a traffic spike. What is a possible cause for the slow scaling response?

A.The cooldown period is set too high.
B.The health check grace period is set too low.
C.The minimum group size is set too low.
D.The launch template has a long warm-up time.
AnswerA

The cooldown period is a deliberate delay after a scaling activity before the group can launch or terminate additional instances. If set too high, it suppresses subsequent scaling actions for an extended time, so the group cannot keep pace with rapidly increasing demand. For a dynamic scaling policy, this directly prevents the Auto Scaling group from adding instances promptly, making scaling appear slow. Thus, an excessively high cooldown is a valid cause of the described problem.

Why this answer

A high cooldown period prevents the Auto Scaling group from launching additional instances immediately after a scaling activity, even if the CPU utilization remains high. During a traffic spike, if the cooldown is set too high, the group waits before responding to further scaling triggers, resulting in slow scale-out. Reducing the cooldown period allows the group to react more quickly to sustained high demand.

Exam trap

DOP-C02 often tests the confusion between cooldown (which throttles scaling actions) and health check grace period (which delays health checks) — candidates may pick the grace period thinking it affects scaling speed.

How to eliminate wrong answers

Option B is wrong because the health check grace period is the time ASG waits before checking the health of a newly launched instance; setting it too low would cause premature termination, not slow scaling. Option C is wrong because the minimum group size determines the baseline number of instances, not the speed of scaling out — a low minimum does not slow down the launch of new instances. Option D is wrong because the launch template does not have a 'warm-up time' parameter; while instance boot time can affect readiness, the question asks about the scaling response, and the cooldown is the direct throttle on scaling actions.

89
MCQhard

A company runs a critical application on Amazon ECS with Fargate launch type. The application uses an Application Load Balancer (ALB) in front. During a load test, the team notices a sudden increase in 5xx errors from the ALB, and some tasks become unhealthy. The task logs show occasional 'OutOfMemoryError' exceptions. The task definition currently has 512 CPU units and 1024 MiB memory. What should the team do to mitigate the issue while maintaining a cost-effective approach?

A.Increase the task definition CPU to 1024 units and memory to 2048 MiB.
B.Increase the task definition memory to 2048 MiB while keeping CPU at 512 units.
C.Configure the ECS service to use a rolling update with a longer health check grace period.
D.Decrease the task definition memory to 512 MiB to force garbage collection more frequently.
AnswerB

Raising the task memory to 2048 MiB while keeping CPU at 512 units is the minimal change that removes the hard memory limit causing the container's OOM kill. ECS enforces task memory as a cgroup limit, so the kernel terminates the process once the container's resident memory reaches that configured cap. This directly provides the application runtime sufficient headroom, uses the valid Fargate 0.5 vCPU / 2 GiB combination, and avoids wasting spend on CPU that was never the bottleneck.

Why this answer

The application is experiencing OutOfMemoryError, indicating the current 1024 MiB memory allocation is insufficient. Increasing memory to 2048 MiB while keeping CPU at 512 units directly resolves the memory constraint without unnecessary CPU cost. ECS Fargate allows independent scaling of CPU and memory within valid combinations, and this change maintains a cost-effective approach by only increasing the resource that is actually constrained.

Exam trap

The trap here is that candidates may assume both CPU and memory must be increased together (Option A) or that a deployment strategy change (Option C) can mitigate resource exhaustion, when in fact the root cause is a memory limit that must be raised independently.

How to eliminate wrong answers

Option A is wrong because it increases both CPU and memory, which is unnecessary and more costly; the issue is memory, not CPU, and the extra CPU units would not resolve OutOfMemoryError. Option C is wrong because a rolling update with a longer health check grace period does not address the root cause of memory exhaustion; it only delays health check failures without fixing the underlying resource shortage. Option D is wrong because decreasing memory to 512 MiB would exacerbate the OutOfMemoryError, causing more frequent failures and task crashes, not improving garbage collection behavior.

90
MCQmedium

A company uses AWS Elastic Beanstalk for its web application. After a deployment, the environment health changes to 'Severe' and the application becomes unresponsive. The DevOps team needs to quickly revert to the previous working version. What is the FASTEST way to achieve this?

A.Use the Elastic Beanstalk console to deploy the previous application version.
B.Swap environment URLs with a different environment that runs the previous version.
C.Terminate the environment and create a new one with the previous version.
D.Redeploy the same application version to the environment.
AnswerA

Redeploying the previous application version through the Elastic Beanstalk console triggers a fresh environment update that restores the last known-good build, directly reversing the faulty deployment. This is faster than rebuilding the environment or manually patching instances, satisfying the requirement to quickly revert to the previous working version.

Why this answer

Deploying the previous application version through the Elastic Beanstalk console is the fastest rollback because Beanstalk retains prior application versions and can redeploy one directly to the same environment, restoring the working state without rebuilding infrastructure. This is the standard, quickest recovery path.

Exam trap

DOP-C02 often tests the difference between a fast in-place rollback (redeploy previous version) and a blue/green swap, tempting candidates to pick URL swap even when no second environment exists.

How to eliminate wrong answers

Option B is wrong because swapping environment URLs requires a separate environment already running the previous version; if one doesn't exist, this is slower and more complex than a direct redeploy. Option C is wrong because terminating and recreating an environment is slow and disruptive, losing configuration and taking minutes to provision. Option D is wrong because redeploying the same broken version does nothing to fix the issue and would leave the environment unhealthy.

91
MCQeasy

A DevOps engineer is responsible for monitoring a production environment that uses Amazon EC2 Auto Scaling. The engineer notices that the Auto Scaling group has been launching and terminating instances frequently over the past hour. The group uses a dynamic scaling policy based on average CPU utilization. The CloudWatch alarm that triggers scaling is set to a threshold of 70% CPU for scale-out and 30% for scale-in. The engineer checks the CloudWatch metrics and sees that CPU utilization is oscillating between 40% and 60%, never reaching the thresholds. The engineer suspects that the scaling policy is not working correctly. The engineer is considering the following actions: A) Change the scaling policy to use a target tracking policy with a target value of 50% CPU utilization. B) Increase the cooldown period for the scaling policy to 300 seconds. C) Disable the scale-in policy to prevent frequent terminations. D) Use a simple scaling policy instead of a dynamic scaling policy. Which action should the engineer take?

A.Disable the scale-in policy to prevent frequent terminations.
B.Change the scaling policy to use a target tracking policy with a target value of 50% CPU utilization.
C.Use a simple scaling policy instead of a dynamic scaling policy.
D.Increase the cooldown period for the scaling policy to 300 seconds.
AnswerB

Target tracking is a dynamic scaling policy that uses a control loop to continuously compute the required capacity needed to keep the CloudWatch metric at the specified target value. Rather than reacting with binary on/off alarms, it applies proportional adjustments based on the current deviation from 50% CPU utilization, smoothing out capacity changes and preventing the overshoot/undershoot cycle that causes oscillation. This gives the Auto Scaling group a clear, single objective that balances responsiveness with stability.

Why this answer

A target tracking policy automatically adjusts the desired capacity to maintain a target utilization (e.g., 50%), which smooths out oscillations by continuously adapting to load rather than reacting to fixed thresholds. Option A is wrong because disabling scale-in could lead to over-provisioning and increased costs. Option C is wrong because simple scaling policies are more prone to causing oscillations due to step adjustments with cooldowns.

Option D is wrong because increasing the cooldown period only delays scaling actions and does not address the root cause of oscillations; target tracking is more effective.

92
MCQhard

An application running on Amazon ECS Fargate is experiencing intermittent HTTP 503 errors from the Application Load Balancer (ALB). The target group health checks are passing. Which configuration is MOST likely causing this issue?

A.The deregistration delay is set too short, causing connections to be closed before requests complete.
B.The ALB's slow start duration is too long, causing requests to be dropped.
C.The health check interval is set too low, causing targets to be marked unhealthy prematurely.
D.The ALB's circuit breaker is tripping due to high error rates.
AnswerA

The deregistration delay (connection draining) is the time ALB gives in-flight requests to finish after a target is deregistered. On ECS Fargate, when a task is stopped, the ALB stops routing new requests and, if the delay is too short, forcibly closes the connection before long-running requests complete, causing clients to receive 502/503 responses. For high-latency or streaming workloads, the delay must exceed the longest expected request duration.

Why this answer

The intermittent HTTP 503 errors from the ALB, despite health checks passing, indicate that the target (ECS Fargate task) is being deregistered while still processing active requests. A deregistration delay that is too short causes the ALB to close connections to the target before the application has finished responding, resulting in 503 errors for in-flight requests. The default deregistration delay is 300 seconds; setting it too low (e.g., 5 seconds) can cause this issue.

Exam trap

The trap here is that candidates often assume 503 errors are always caused by health check failures or overload, but the question explicitly states health checks are passing, so the issue must be related to connection draining or deregistration behavior rather than target health or load balancing algorithms.

How to eliminate wrong answers

Option B is wrong because slow start gradually increases the number of requests sent to a newly registered target, which can cause transient performance issues but does not directly cause 503 errors; it would more likely cause latency or uneven load distribution. Option C is wrong because a low health check interval would cause targets to be marked unhealthy more frequently, but the question states health checks are passing, so this is not the cause of the 503 errors. Option D is wrong because the ALB circuit breaker (a feature of AWS App Mesh or certain service meshes) is not a standard ALB feature; ALBs do not have a built-in circuit breaker that trips due to high error rates—instead, they rely on health checks and target group settings.

93
MCQmedium

A company uses AWS CloudFormation to deploy infrastructure. During an incident, a stack update fails with a stack rollback. The engineer needs to prevent the stack from rolling back on future failures and instead retain the resources for debugging. Which CloudFormation feature should the engineer use?

A.Enable drift detection on the stack
B.Use the '--disable-rollback' option when updating the stack
C.Use AWS CloudFormation StackSets to deploy the stack
D.Create a change set before updating instead of direct update
AnswerB

The `--disable-rollback` flag on `aws cloudformation update-stack` instructs CloudFormation to leave the stack in its failed state instead of reverting to the previous successful template. This is the correct way to prevent rollback because it preserves the partially updated resources—including any S3 buckets, EC2 instances, or custom resources—for root-cause analysis and manual remediation. The stack status becomes `UPDATE_FAILED`, and the original resources remain untouched until you intervene.

Why this answer

The `--disable-rollback` option (or `DisableRollback` in the CloudFormation template) instructs AWS CloudFormation to leave the stack in its current state (with the failed resources intact) instead of automatically rolling back to the last known good state. This allows engineers to retain the resources for debugging without the stack being torn down on failure.

Exam trap

The trap here is that candidates often confuse change sets (which preview changes) with the ability to prevent rollback, or mistakenly think drift detection or StackSets can alter rollback behavior, when only the `--disable-rollback` flag directly controls whether resources are retained on failure.

How to eliminate wrong answers

Option A is wrong because drift detection only identifies differences between the stack's actual deployed resources and the expected template configuration; it does not prevent rollback or retain resources after a failed update. Option C is wrong because StackSets are used to deploy stacks across multiple accounts and regions, not to control rollback behavior on a single stack update failure. Option D is wrong because a change set allows you to preview changes before updating, but it does not affect the rollback behavior; if the update fails, the stack will still roll back by default unless `--disable-rollback` is specified.

94
MCQhard

During an incident, a DevOps engineer needs to quickly revoke access to a set of IAM users who are suspected to be compromised. The users have programmatic access keys and console passwords. The engineer wants to minimize the impact on non-compromised users. Which action should the engineer take FIRST?

A.Delete the compromised IAM users.
B.Attach an IAM policy that explicitly denies all actions to the compromised users.
C.Delete the access keys of the compromised users.
D.Change the IAM password policy to require strong passwords.
AnswerB

Attaching a customer-managed IAM policy with an explicit 'Deny' effect for all actions on all resources immediately revokes both console and programmatic access for the compromised users. IAM's evaluation logic gives explicit denies precedence over any allow, so even if the user still has attached policies with broad permissions, the deny-all policy overrides them. This containment action preserves the user objects and their audit trail for forensic analysis, allowing you to safely investigate and later rotate credentials or delete the users if needed. Unlike disabling keys or changing password policies, this approach covers all access methods simultaneously.

Why this answer

Attaching an explicit Deny policy to the compromised users immediately blocks all API and console actions for those principals while leaving other users untouched, and it is reversible once the incident is resolved. This is the fastest containment step that satisfies least-impact on non-compromised users. Deleting keys or users is more destructive and slower to reverse.

Exam trap

The trap is confusing 'revoke access' with 'delete the identity' — candidates pick key deletion or user deletion, missing that an explicit Deny is the fastest, least-destructive, and most complete containment because it blocks both console and API paths.

How to eliminate wrong answers

Option A is wrong because deleting IAM users is irreversible, destroys audit trail attribution, and may break dependent resources — it is a cleanup step, not a first containment action. Option C is wrong because deleting access keys only stops programmatic access; the compromised console password would still allow console login, leaving a live attack path. Option D is wrong because strengthening the password policy affects all users globally and does nothing to revoke the already-compromised credentials.

95
MCQeasy

A company's DevOps team uses AWS Config to monitor resource compliance. They have created a custom AWS Config rule that triggers an AWS Lambda function to evaluate whether EC2 instances have the 'Environment' tag with value 'Production' or 'Staging'. The rule is set to evaluate resources on configuration changes. However, the team notices that the rule does not trigger when an EC2 instance is launched. The Lambda function's IAM role has the necessary permissions to describe EC2 instances. The CloudWatch Logs for the Lambda function show that it is not being invoked. What is the MOST likely reason?

A.The Lambda function's IAM role does not have permission to write to CloudWatch Logs.
B.The AWS Config rule is set to evaluate resources periodically, not on configuration changes.
C.The AWS Config rule is not configured to trigger on AWS::EC2::Instance resources.
D.The custom rule must be deployed using AWS CloudFormation to be active.
AnswerC

For a custom AWS Config rule to evaluate a resource on configuration changes, the rule's scope must include that resource type. If the rule is defined without AWS::EC2::Instance in its 'Resource types' scope (or if it is scoped to a specific tag/resource ID that does not match), AWS Config will not send evaluation events to the Lambda function for EC2 instances. In this scenario, the rule has the correct trigger type (change) but its scope does not cover EC2 instances, so Config never invokes the Lambda function on instance launch. This is the exact cause of the DevOps team's issue.

Why this answer

For an AWS Config custom rule to fire on resource creation, the rule must specify the resource types it evaluates (e.g., AWS::EC2::Instance) in its scope. If the scope does not include EC2 instances, Config will not invoke the rule's Lambda function when an instance is launched, which matches the symptom that the Lambda is never invoked despite having correct permissions. The trigger type (configuration change) is already stated as correct in the scenario, so the missing piece is the resource scope.

Exam trap

The trap is assuming the problem is IAM or trigger type when the scenario already rules those out; the real cause is the rule's resource-type scope, which is easy to overlook because it is configured separately from the trigger.

How to eliminate wrong answers

Option A is wrong because the scenario states the Lambda role has the necessary permissions, and a CloudWatch Logs write permission issue would still allow invocation (the function would run but fail to log), not prevent invocation entirely. Option B is wrong because the question explicitly states the rule is set to evaluate on configuration changes, so this contradicts the given facts. Option D is wrong because custom Config rules can be created via the console, CLI, or CloudFormation; CloudFormation deployment is not a requirement for the rule to be active.

96
MCQmedium

An application running on Amazon EC2 instances in an Auto Scaling group is experiencing intermittent connectivity issues. The DevOps team suspects a security group configuration problem. Which approach should the team use to analyze security group traffic and identify denied requests?

A.Use AWS Config to review security group rules
B.Check AWS CloudTrail for security group modification events
C.Enable AWS Security Hub and review the security findings
D.Enable VPC Flow Logs and query Amazon Athena
AnswerD

VPC Flow Logs capture IP traffic metadata for ENIs in the VPC, recording fields like source address, destination address, port, protocol, and the action — ACCEPT or REJECT — for each connection. By publishing flow logs to Amazon S3 and using Amazon Athena with its SerDe, you can run SQL queries to filter on action = 'REJECT' and aggregate the denied connection attempts by source IP, port, or time. This directly reveals which connections are being blocked, making it the correct way to investigate denied traffic.

Why this answer

VPC Flow Logs capture IP traffic metadata (source, destination, port, protocol, action) for ENIs, subnets, or VPCs, and publishing them to CloudWatch Logs or S3 lets you query with Amazon Athena to identify REJECT entries caused by security group or NACL rules. This is the standard AWS approach for diagnosing denied traffic at the network layer.

Exam trap

The trap is confusing 'audit who changed the rules' (CloudTrail/Config) with 'see which packets were denied' (Flow Logs) — the question asks for traffic analysis, not configuration history.

How to eliminate wrong answers

Option A is wrong because AWS Config records configuration changes and evaluates compliance — it shows what the security group rules are, not which packets were denied by them. Option B is wrong because CloudTrail logs API calls (e.g., AuthorizeSecurityGroupIngress), which tells you who changed rules but not whether traffic was blocked. Option C is wrong because Security Hub aggregates findings from services like GuardDuty and Inspector; it does not provide per-packet flow analysis of denied requests.

97
Multi-Selecteasy

A DevOps engineer is troubleshooting an AWS CodeDeploy deployment that failed. Which TWO resources should the engineer examine to identify the cause of the failure? (Choose two.)

Select 2 answers
A.EC2 instance system logs
B.CloudWatch Logs for CodeDeploy
C.S3 access logs
D.CloudTrail logs
E.CodeDeploy deployment group configuration
AnswersB, E

CloudWatch Logs for CodeDeploy is the correct source because it centralizes deployment events and error messages. When you configure a log group for the deployment group, lifecycle event execution details—such as BeforeInstall, AfterInstall, ApplicationStart, and the associated script output—are streamed into CloudWatch Logs. This lets you query and filter by deployment ID and instance ID, making it the fastest way to identify which lifecycle hook failed and why, rather than logging into individual instances.

Why this answer

AWS CodeDeploy emits detailed logs about deployment lifecycle events (e.g., BeforeInstall, ApplicationStop) to CloudWatch Logs. These logs contain error messages, script output, and status codes that directly indicate why a deployment step failed, such as a permission issue or a script syntax error. Examining CloudWatch Logs for CodeDeploy is the primary method to diagnose deployment failures.

Exam trap

The trap here is that candidates often confuse CloudTrail (API auditing) with CloudWatch Logs (application-level logging), or they mistakenly think EC2 system logs are relevant for application deployment failures, when in fact CodeDeploy-specific logs are the correct source.

98
MCQmedium

A company uses AWS Lambda functions to process incoming events from Amazon S3. The operations team notices that some events are not being processed, and there is no error in the Lambda function logs. What is the most likely cause?

A.The Lambda function has reserved concurrency set to a low value, causing throttling.
B.The S3 event notification is configured to send to an SNS topic that is not subscribed to the Lambda function.
C.The S3 bucket policy does not allow the Lambda function to be invoked.
D.The Lambda function has a timeout that is too short.
AnswerA

Reserved concurrency caps the number of concurrent executions for the function. When all slots are busy, Lambda immediately rejects new invocations with a Throttle error. Because the code never starts, no START/END/REPORT lines are emitted, so CloudWatch Logs for the function remain empty. The S3 event notification will retry for a few hours, but each throttled attempt leaves no log.

Why this answer

When a Lambda function has reserved concurrency set to a low value, it limits the number of concurrent executions allowed for that function. If incoming S3 events exceed this limit, the Lambda service throttles the invocations, causing some events to be silently dropped without generating errors in the function logs because the function never actually runs. This matches the symptom of missing events with no error logs.

Exam trap

The trap here is that candidates often assume missing events are due to permission or timeout errors, but the absence of any error logs points to throttling, where the function is never invoked and thus no logs are generated.

How to eliminate wrong answers

Option B is wrong because if the SNS topic is not subscribed to the Lambda function, the event would never reach Lambda, but the question states the Lambda function logs show no errors, implying the function is invoked for some events; the issue is about events not being processed, not about delivery failure. Option C is wrong because if the S3 bucket policy did not allow Lambda invocation, the invocation would fail with an access denied error, which would be logged in CloudTrail or appear as an error in the Lambda logs, contradicting the 'no error' condition. Option D is wrong because a timeout that is too short would cause the function to fail mid-execution and generate a timeout error in the Lambda logs, not silently drop events without any log entries.

99
MCQhard

A company uses Amazon CloudWatch Logs to collect application logs from EC2 instances. The security team requires that log data be encrypted at rest using a customer-managed AWS KMS key. The logs are currently being delivered, but they are not encrypted. What is the most likely reason?

A.The IAM role for the EC2 instance does not have kms:Encrypt permission
B.The CloudWatch Logs agent is not configured to encrypt logs
C.The KMS key is disabled
D.The KMS key policy does not allow the CloudWatch Logs service principal
AnswerD

This is the correct cause. CloudWatch Logs acts under the service principal logs.amazonaws.com when performing server-side encryption with an AWS KMS customer managed key. If the KMS key policy does not include a statement granting this principal the kms:Encrypt and kms:DescribeKey actions (and denies are absent), the service cannot encrypt the log events. In such a case, log delivery may continue to succeed but the data is stored without the intended encryption, exactly matching the symptom described.

Why this answer

For CloudWatch Logs to encrypt log data at rest with a customer-managed KMS key, the key policy must grant the CloudWatch Logs service principal the necessary permissions (kms:Encrypt, kms:Decrypt, etc.). If the key policy does not include this, CloudWatch Logs can still ingest the logs but cannot encrypt them, resulting in unencrypted logs. Option A is incorrect: the IAM role for the EC2 instance does not need kms:Encrypt permissions for server-side encryption; that is handled by the CloudWatch Logs service using the key policy.

Option B is incorrect because encryption is configured at the log group level, not in the agent. Option C would cause delivery failures, not just lack of encryption.

100
MCQmedium

A company uses AWS CodePipeline for CI/CD. During a production deployment, the pipeline fails at the 'Deploy' stage with an error: 'The deployment failed because the deployment group does not have enough capacity to handle the deployment.' The engineer checks the CodeDeploy deployment group and sees that it is configured with a minimum healthy hosts of 100% and a deployment configuration of 'CodeDeployDefault.OneAtATime'. What is the MOST likely cause?

A.The deployment configuration 'OneAtATime' is not compatible with the deployment group.
B.The target group health check is misconfigured, causing all instances to be unhealthy.
C.The CodeDeploy agent on the instances is not running.
D.The deployment group has only one instance, and the minimum healthy hosts setting prevents the deployment.
AnswerD

CodeDeploy enforces a minimum number of healthy hosts as a safety condition, and when the deployment group has exactly one instance, any positive minimum healthy host count makes an in-place deployment impossible. With a setting of 100% or 1, taking that only instance out of service to deploy to it would reduce the healthy host count to 0, violating the constraint and aborting the deployment at the very start. The error message specifically mentions that the deployment was aborted because the minimum number of healthy hosts was not met.

Why this answer

With a minimum healthy hosts of 100% and a OneAtATime deployment configuration, CodeDeploy must keep all instances healthy during the deployment. If the deployment group contains only one instance, taking it out of service to deploy would drop healthy hosts below 100%, making the deployment impossible. This is the most likely cause of the capacity error.

Exam trap

DOP-C02 often tests the interaction between deployment configuration and minimum healthy hosts — candidates focus on the deployment config alone and miss that 100% minimum healthy hosts with a single instance creates an impossible constraint.

How to eliminate wrong answers

Option A is wrong because OneAtATime is a valid deployment configuration and is compatible with any deployment group; the issue is the combination with 100% minimum healthy hosts and a single instance. Option B is wrong because a misconfigured health check would cause instances to be unhealthy, but the error specifically mentions insufficient capacity, not unhealthy targets. Option C is wrong because a stopped CodeDeploy agent would produce a different error about the agent not being available, not a capacity error.

101
MCQeasy

A DevOps engineer notices that an Amazon RDS for MySQL instance has failed over to a standby replica. The engineer needs to identify the root cause by examining metrics. Which AWS service should the engineer use to view the database load, replication lag, and failover events?

A.AWS CloudTrail
B.Amazon CloudWatch
C.VPC Flow Logs
D.AWS Trusted Advisor
AnswerB

Amazon CloudWatch is the native monitoring service for Amazon RDS, automatically collecting and storing database metrics such as CPU utilization, free memory, read/write throughput, and ReplicaLag for MySQL read replicas. These metrics are published as time-ordered data points and can be visualized on dashboards, trigger alarms via Amazon CloudWatch Alarms, and feed AWS Auto Scaling. With Enhanced Monitoring, it can also expose OS-level metrics, making it the correct source for diagnosing load and replication issues.

Why this answer

Amazon CloudWatch provides RDS metrics including DatabaseConnections, ReplicaLag, ReadIOPS, and failover-related events. CloudWatch alarms and the RDS console's monitoring graphs surface failover events and replication lag, making it the correct service for diagnosing an RDS failover root cause.

Exam trap

DOP-C02 often tests the distinction between CloudWatch (metrics and performance) and CloudTrail (API audit logs) — candidates pick CloudTrail for performance diagnosis when metrics are needed.

How to eliminate wrong answers

Option A is wrong because AWS CloudTrail records API calls and management events (who did what), not performance metrics like database load or replication lag. Option C is wrong because VPC Flow Logs capture IP traffic metadata for network troubleshooting, not database-level metrics. Option D is wrong because AWS Trusted Advisor provides best-practice checks and recommendations, not real-time database performance metrics.

102
MCQhard

A company uses an Application Load Balancer (ALB) in front of a fleet of EC2 instances. The security team reports that a specific client IP address is sending malicious requests and must be blocked immediately. The ALB's security group only allows HTTP/HTTPS from 0.0.0.0/0. What is the FASTEST way to block traffic from this IP address without affecting other traffic?

A.Create an AWS WAF web ACL with an IP set deny rule and associate it with the ALB.
B.Modify the ALB listener rules to drop requests from the client IP.
C.Update the ALB security group to add a deny rule for the client IP address.
D.Update the VPC route table to drop packets from the client IP.
AnswerA

AWS WAF provides application-layer (Layer 7) inspection and can be associated directly with an ALB. By creating a web ACL that references an IP set containing the offending client IP and setting the default or custom rule action to 'Block', requests from that IP are denied before they reach the ALB. This is the purpose-built, scalable mechanism for IP-based blocking in front of an ALB, and it can be implemented without altering routing or security group configurations.

Why this answer

AWS WAF with an IP set deny rule associated to the ALB is the fastest and most surgical way to block a specific client IP. WAF evaluates requests at the ALB before they reach targets, and an IP match condition with a block action takes effect within seconds without touching security groups or routing. This blocks only the offending IP while all other traffic continues normally.

Exam trap

DOP-C02 often tests the allow-only nature of security groups and the lack of source-IP deny in ALB listener rules — the trap is assuming security groups or listener rules can block a single IP like a firewall ACL.

How to eliminate wrong answers

Option B is wrong because ALB listener rules support only forward, redirect, and fixed-response actions — there is no native 'drop by source IP' rule, and fixed-response would still return a response rather than silently blocking. Option C is wrong because security groups are allow-only; you cannot add a deny rule, and removing 0.0.0.0/0 would break all traffic. Option D is wrong because VPC route tables route by destination CIDR, not source IP, and cannot selectively drop a single client address.

103
MCQmedium

A company's DevOps team uses AWS CodePipeline to automate deployments. A recent pipeline execution failed at the 'Deploy' stage. The engineer needs to view the detailed logs for the failed action. Which AWS service or feature should the engineer use?

A.CloudWatch Logs
B.S3 access logs
C.CodeBuild logs
D.CloudTrail
AnswerA

CodePipeline natively publishes execution-level log events to Amazon CloudWatch Logs, including stage start/end timestamps, action state transitions, and failure diagnostics. This provides a centralized, queryable repository that can be filtered with CloudWatch Logs Insights, used for metric filters and alarms, and retained beyond the console's default history. Because these logs capture the pipeline's own runtime behavior, they are the authoritative source for troubleshooting execution details.

Why this answer

AWS CodePipeline integrates with Amazon CloudWatch Logs to capture and store detailed execution logs for each pipeline action, including the 'Deploy' stage. When a deployment action fails, the engineer can view the associated logs directly from the CodePipeline console or via the CloudWatch Logs console, which provides granular error messages, timestamps, and stack traces necessary for troubleshooting. This is the designated service for accessing action-level logs in CodePipeline.

Exam trap

The trap here is that candidates may confuse CloudTrail (audit logs) with CloudWatch Logs (operational logs), or assume that CodeBuild logs cover all pipeline stages, when in fact each stage type (e.g., Deploy) has its own log destination.

How to eliminate wrong answers

Option B (S3 access logs) is wrong because S3 access logs record requests made to an S3 bucket, not the execution logs of CodePipeline actions. Option C (CodeBuild logs) is wrong because CodeBuild logs are specific to build actions within CodePipeline, not to deploy actions, which may use other providers like CodeDeploy or Elastic Beanstalk. Option D (CloudTrail) is wrong because CloudTrail records API calls made to AWS services for auditing purposes, not the detailed runtime logs of a pipeline execution.

104
MCQeasy

A company uses AWS CloudFormation to deploy infrastructure. A recent stack update failed, and the engineer needs to roll back to the previous stable state. Which CloudFormation feature should the engineer use?

A.Use AWS CloudFormation Drift Detection.
B.Use AWS CloudFormation StackSets.
C.Use the 'Rollback' action in the CloudFormation console.
D.Create a Change Set and execute it.
AnswerC

The 'Rollback' action in the CloudFormation console is the correct approach because CloudFormation maintains an automatic rollback trigger for failed update operations, and when an update enters a failed state, you can explicitly invoke 'Rollback' from the console to restore the stack to its last known good template and parameter state. This action is specifically designed to undo the partial or failed changes, reverting affected AWS resources to their pre-update configuration, and is the direct rollback mechanism within CloudFormation's lifecycle.

Why this answer

CloudFormation provides a built-in 'Rollback' action that automatically reverts a stack to its last known stable state when a stack update fails. This rollback occurs by default on failure, but if the engineer needs to manually trigger it after a failed update, they can use the 'Rollback' action in the CloudFormation console or API (e.g., `aws cloudformation rollback-stack`). This ensures the infrastructure returns to the previous working configuration without manual intervention.

Exam trap

The trap here is that candidates may confuse the 'Rollback' action with creating a Change Set or using Drift Detection, thinking they need to manually define the previous state, when in fact CloudFormation automatically preserves the previous stable state for rollback.

How to eliminate wrong answers

Option A is wrong because Drift Detection is used to identify whether a stack's actual resources have deviated from the expected template configuration, not to roll back a failed update. Option B is wrong because StackSets are designed to deploy stacks across multiple accounts and regions, not to handle rollback of a single stack update. Option D is wrong because creating and executing a Change Set is a method to propose and apply changes to a stack, but it does not automatically revert to the previous state; it would require the engineer to manually define the previous template and parameters, which is less efficient than using the built-in rollback feature.

105
MCQmedium

A company has an AWS Lambda function that processes S3 events. The function is invoked multiple times for the same S3 object, causing duplicate processing. The engineer suspects the issue is related to retries from the S3 event notification or Lambda's built-in retry behavior. What is the MOST effective way to ensure idempotent processing?

A.Modify the S3 bucket event notification configuration to use a prefix filter that excludes duplicate objects.
B.Use a DynamoDB table to store a record of processed S3 object keys and check for existence before processing.
C.Set the Lambda function's ReservedConcurrency to 1 to prevent concurrent executions.
D.Use an Amazon SQS FIFO queue as the event source and enable content-based deduplication.
AnswerB

Recording processed S3 object keys in a DynamoDB table and checking for existence before processing makes the Lambda function idempotent. Use a conditional write, such as PutItem with ConditionExpression: attribute_not_exists(partition_key), so the first invocation for a given object key proceeds and subsequent duplicates fail the condition and exit without side effects. This is the recommended pattern because it works regardless of how many duplicate events S3 delivers and does not depend on delivery semantics.

Why this answer

Storing processed S3 object keys in a DynamoDB table and checking for existence before processing ensures idempotency at the application level. This approach directly handles duplicate invocations caused by S3 event retries or Lambda's built-in retry behavior, as the function can conditionally skip processing if the key already exists in DynamoDB. It provides a durable, consistent, and scalable mechanism to prevent duplicate processing regardless of how many times the function is invoked for the same object.

Exam trap

The trap here is that candidates often confuse concurrency control (ReservedConcurrency) with idempotency, or assume SQS FIFO deduplication is a drop-in solution without realizing S3 cannot directly send events to FIFO queues.

How to eliminate wrong answers

Option A is wrong because S3 prefix filters only filter events based on object key prefixes or suffixes, not on duplicate detection; they cannot prevent multiple notifications for the same object. Option C is wrong because setting ReservedConcurrency to 1 prevents concurrent executions but does not prevent sequential duplicate invocations from retries; the function could still be invoked multiple times for the same object in sequence. Option D is wrong because using an SQS FIFO queue with content-based deduplication would require the S3 event notification to be sent to the queue, but S3 does not natively support sending events to SQS FIFO queues; it only supports standard SQS queues, and even if it did, the deduplication window is only 5 minutes, which may not cover all retry scenarios.

106
MCQhard

An application running on Amazon ECS with Fargate is experiencing increased latency. The DevOps team suspects that the task is running out of memory and swapping. Which set of CloudWatch metrics should the team examine to confirm this suspicion?

A.NetworkIn and NetworkOut
B.MemoryUtilized and MemoryReserved
C.CPUUtilization and CPUReservation
D.EphemeralStorageUtilized and EphemeralStorageReserved
AnswerB

MemoryUtilized and MemoryReserved are the correct CloudWatch ECS metrics for Fargate tasks, as they directly measure the task's memory consumption against the memory limit configured in the task definition. When MemoryUtilized consistently approaches or equals MemoryReserved, the container runtime may start swapping memory pages to disk, which degrades performance. Monitoring these two metrics together provides the earliest signal of memory pressure that leads to swapping.

Why this answer

MemoryUtilized and MemoryReserved are the CloudWatch metrics that directly track memory consumption and allocation for ECS tasks using Fargate. When a task runs out of memory, the Linux kernel’s OOM killer may terminate processes, but before that, swapping can occur if swap space is available, causing increased latency. These metrics allow the DevOps team to compare actual memory usage against the task’s memory reservation, confirming if memory pressure is the root cause of the latency.

Exam trap

The trap here is that candidates may confuse memory metrics with CPU or storage metrics, assuming any resource constraint can cause swapping, but only memory metrics directly indicate memory exhaustion and potential swapping behavior.

How to eliminate wrong answers

Option A is wrong because NetworkIn and NetworkOut measure network throughput, not memory usage or swapping, so they cannot confirm memory exhaustion. Option C is wrong because CPUUtilization and CPUReservation track CPU usage and reservation, which are unrelated to memory swapping; high CPU could cause latency but does not indicate memory pressure. Option D is wrong because EphemeralStorageUtilized and EphemeralStorageReserved measure storage usage on the Fargate ephemeral volume, not memory consumption; swapping would involve memory, not ephemeral storage.

107
MCQhard

A DevOps engineer is troubleshooting an application running on an EC2 instance. The application needs to access an Amazon RDS database using IAM database authentication. The EC2 instance is associated with an IAM role 'EC2-AppRole', and the RDS instance has a resource-based policy that allows 'DatabaseAccessRole' to connect. The engineer sees the error in the exhibit. What is the most likely cause?

A.The RDS instance does not have a resource-based policy that grants access to 'DatabaseAccessRole'.
B.The security group for the EC2 instance does not allow outbound traffic to the RDS instance.
C.The EC2 instance does not have the correct IAM instance profile attached.
D.The trust policy of the IAM role 'DatabaseAccessRole' does not allow the EC2 instance role 'EC2-AppRole' to assume it.
AnswerD

For IAM database authentication, the application must first assume the IAM role 'DatabaseAccessRole' to obtain credentials authorized to generate the RDS token; the trust policy on 'DatabaseAccessRole' must explicitly list 'EC2-AppRole' as a trusted principal. When this trust policy does not allow the EC2 instance's role to assume it, the STS AssumeRole call returns 'AccessDenied', so the application cannot acquire the token required for authentication to RDS. This exactly matches the observed error, confirming the trust policy of 'DatabaseAccessRole' is the root cause.

Why this answer

The error indicates that the EC2 instance's IAM role 'EC2-AppRole' cannot authenticate to the RDS instance. IAM database authentication requires the EC2 instance to assume a database authentication token, which is generated by calling the RDS API with the 'EC2-AppRole' credentials. However, the RDS instance's resource-based policy only allows 'DatabaseAccessRole' to connect.

For 'EC2-AppRole' to successfully authenticate, it must first assume 'DatabaseAccessRole' via a trust policy that permits the EC2 instance role to assume it. Without this trust relationship, the authentication token request fails, causing the error.

Exam trap

The trap here is that candidates often assume the error is due to missing resource-based policies or network connectivity, but the core issue is the missing trust relationship between the EC2 instance role and the database access role, which is a common misconfiguration in cross-account or cross-role IAM authentication setups.

How to eliminate wrong answers

Option A is wrong because the RDS instance does have a resource-based policy that allows 'DatabaseAccessRole' to connect, as stated in the question; the issue is that the EC2 instance role cannot assume that role. Option B is wrong because security group rules control network traffic, not IAM authentication; if the security group were blocking outbound traffic, the error would be a network timeout or connection refused, not an IAM authentication failure. Option C is wrong because the EC2 instance is already associated with the IAM role 'EC2-AppRole' (the instance profile is attached), and the error is about assuming another role, not about the instance lacking a role.

108
MCQmedium

A company uses an Auto Scaling group with a dynamic scaling policy based on a custom CloudWatch metric. After a recent deployment, the metric spikes unexpectedly, causing the Auto Scaling group to launch several EC2 instances. The operations team wants to quickly determine whether the spike was caused by a real load increase or a deployment issue. What is the MOST efficient way to investigate this?

A.Check the SNS topic that the scaling policy publishes to for notifications.
B.Use CloudWatch Logs Insights to query application logs for error patterns or deployment markers that coincide with the metric spike.
C.Use AWS CloudTrail to review API calls that modified the scaling policy.
D.Temporarily disable the scaling policy and manually increase the desired capacity to handle the load.
AnswerB

CloudWatch Logs Insights lets you run SQL-like queries against application logs stored in CloudWatch Logs, allowing you to filter for specific error codes, exception stack traces, or deployment markers that align with the exact timestamp of the metric spike. By correlating log timestamps with the scaling activity, you can identify events like a failed code release, a dependency outage, or a traffic surge that pushed the metric above the alarm threshold. This is the only option that gives access to application-level evidence needed to diagnose the upstream cause of the spike.

Why this answer

CloudWatch Logs Insights allows you to query application logs for error patterns or deployment markers (e.g., new version tags, exception stack traces) that coincide with the metric spike. This directly correlates the scaling event with application-level evidence, enabling rapid root-cause analysis without altering infrastructure or relying on indirect notifications.

Exam trap

The trap here is that candidates confuse monitoring scaling actions (SNS/CloudTrail) with diagnosing the metric's root cause, overlooking that application logs provide the direct evidence needed to distinguish real load from deployment issues.

How to eliminate wrong answers

Option A is wrong because SNS topics used by scaling policies only send notifications about scaling actions (e.g., instance launched), not the underlying cause of the metric spike; they lack application-level context. Option C is wrong because AWS CloudTrail records API calls that modify the scaling policy (e.g., PutScalingPolicy), but the metric spike itself is not an API call—it is a CloudWatch metric data point, so CloudTrail cannot show why the metric changed. Option D is wrong because temporarily disabling the scaling policy and manually increasing desired capacity is a reactive workaround that does not investigate the root cause; it masks the symptom and risks over-provisioning or missing a deployment bug.

109
MCQhard

A company runs a stateful web application on EC2 instances behind an Application Load Balancer. The application uses sticky sessions (session affinity) based on cookies. During a deployment, the DevOps engineer notices that some users are being logged out and losing session data. The deployment uses a rolling update strategy. What is the MOST likely cause?

A.The Auto Scaling group is terminating instances before the new ones are fully ready.
B.The health check interval is too long, causing the ALB to route traffic to unhealthy instances.
C.The ALB sticky session cookie is not being generated correctly.
D.The session data is stored locally on the EC2 instance, not in a shared external store.
AnswerD

Because the web application stores its session state locally—either in memory (e.g., a servlet HttpSession) or on the instance's ephemeral disk—any instance replacement destroys active sessions. During a rolling update, the Auto Scaling group terminates old instances after launching new ones, forcing users who had sessions on those instances to be logged out. Moving session data to a shared external store such as ElastiCache for Redis, DynamoDB, or a central database makes sessions independent of instance lifecycle and prevents this behavior.

Why this answer

Sticky sessions on the ALB bind a user to a specific EC2 instance via a cookie (AWSALB). If session data is stored locally on that instance, terminating the instance during a rolling update destroys the session data, logging the user out. The correct fix is to externalize session state to a shared store like ElastiCache, DynamoDB, or a database so any instance can serve any user.

Exam trap

The trap is blaming the deployment strategy or health checks when the real issue is architectural — local session storage is incompatible with elastic, stateless compute.

How to eliminate wrong answers

Option A is wrong because while premature termination could cause issues, the question specifies a rolling update strategy where instances are replaced gradually — the root cause is that session data is not shared, so even a properly executed rolling update loses sessions. Option B is wrong because a long health check interval would delay detection of unhealthy instances, not cause session loss during a rolling update; the ALB would still route to healthy instances. Option C is wrong because if the ALB sticky session cookie were not generated correctly, users would be load-balanced randomly from the start, not just during deployment — and the symptom would be inconsistent sessions, not logout during updates.

110
MCQhard

An EC2 instance is in 'running' state according to the CLI output, but the application hosted on it is unreachable. The DevOps engineer checks the security group and finds it allows inbound HTTP traffic from 0.0.0.0/0. The instance has a public IP. What is the MOST likely issue?

A.The network ACL is blocking inbound HTTP traffic.
B.The instance does not have a public IP address assigned.
C.The security group is attached to the instance but does not allow inbound HTTP.
D.The instance's OS firewall (e.g., iptables) is blocking the traffic.
AnswerD

The instance OS runs its own packet-filtering firewall—typically iptables or firewalld on Linux—which operates independently of AWS's security groups and network ACLs. Even with a permissive security group and a correct public IP, an iptables rule (e.g., default INPUT policy DROP or a REJECT rule) will drop inbound HTTP packets before they reach the web server. This is a classic scenario where all AWS-side checks pass but the instance remains unreachable, hence the CLI 'running state' does not reflect host-level firewall state.

Why this answer

The instance is running, has a public IP, and the security group allows inbound HTTP from 0.0.0.0/0. Since the security group is correctly configured, the most likely cause is a host-level firewall (e.g., iptables, Windows Firewall) blocking the traffic. Security groups operate at the hypervisor level and do not override OS-level firewalls; thus, even with permissive security group rules, the OS can still drop packets.

This is a common misconfiguration when instances are launched with restrictive default OS firewall rules.

Exam trap

DOP-C02 often tests the misconception that security groups are the only firewall layer, leading candidates to overlook OS-level firewalls when security groups are correctly configured.

How to eliminate wrong answers

Option A is wrong because a network ACL blocking inbound HTTP would typically be a less likely cause given that the security group is correctly configured and the instance is reachable at the network level (though not explicitly stated, the scenario implies the security group is the primary suspect; NACLs are stateless and would need explicit deny rules, but the question does not suggest NACL misconfiguration). Option B is wrong because the instance already has a public IP address, as stated in the question. Option C is wrong because the security group is explicitly described as allowing inbound HTTP from 0.0.0.0/0, so it does allow HTTP.

111
MCQeasy

A DevOps engineer is troubleshooting an AWS Lambda function that is intermittently timing out. The function is configured with a 3-second timeout and 128 MB memory. The function processes messages from an SQS queue. What is the most cost-effective change to reduce timeouts?

A.Increase the SQS batch size to 20
B.Increase the function timeout to 10 seconds
C.Increase the function memory to 256 MB
D.Set reserved concurrency to 10
AnswerC

Lambda allocates CPU proportionally to the configured memory, so increasing memory to 256 MB from a lower value (e.g., 128 MB) roughly doubles the available vCPU fraction, speeding up CPU-bound tasks. This directly targets the performance bottleneck and is a standard first step in Lambda performance tuning. For the stated issue, this change is more likely to reduce execution time than the other options.

Why this answer

Increasing the memory to 256 MB is the most cost-effective change because Lambda allocates CPU proportionally to memory, so doubling the memory from 128 MB to 256 MB also doubles the CPU performance. This reduces execution time, which can resolve timeouts without increasing the timeout duration, and since Lambda billing is based on compute time (GB-seconds), the total cost may stay the same or even decrease if the function finishes faster.

Exam trap

The trap here is that candidates assume increasing the timeout is the only way to fix timeouts, but AWS explicitly recommends increasing memory as a cost-effective performance tuning method because it also increases CPU, which can reduce execution time and thus avoid timeouts without increasing cost.

How to eliminate wrong answers

Option A is wrong because increasing the SQS batch size to 20 would cause the Lambda function to process more messages per invocation, increasing the workload and likely worsening timeouts rather than fixing them. Option B is wrong because increasing the function timeout to 10 seconds does not address the root cause of slow execution; it only masks the symptom and could increase costs if the function still runs longer. Option D is wrong because setting reserved concurrency to 10 limits the number of concurrent executions but does not improve the performance of a single invocation, so it would not reduce timeouts.

112
MCQmedium

A company uses AWS Systems Manager to manage a fleet of EC2 instances. During an incident, a DevOps engineer needs to execute a script on a specific instance to collect diagnostic data. The engineer does not have SSH key access. Which approach should the engineer use to execute the script?

A.Use AWS Systems Manager Run Command to execute the script.
B.Use AWS OpsWorks to run the script as a Chef recipe.
C.Use EC2 Instance Connect to SSH into the instance and run the script.
D.Use AWS Systems Manager Session Manager to open a shell and run the script.
AnswerA

AWS Systems Manager Run Command is the correct choice because it uses the SSM Agent already installed on the EC2 instance to execute a script as a one-shot, non-interactive command. It does not require SSH, opening port 22, or managing SSH keys, and it works across multiple instances concurrently with rate control and output logging. Run Command is purpose-built for exactly this ad-hoc script execution scenario, and it leverages the instance's IAM role for permissions, avoiding any need for bastion hosts or direct network access.

Why this answer

AWS Systems Manager Run Command is the correct tool because it lets you execute scripts or commands on managed EC2 instances remotely without SSH, RDP, or any inbound ports open. The instance only needs the SSM Agent and an IAM instance profile with the right permissions, which satisfies the 'no SSH key access' constraint.

Exam trap

DOP-C02 often tests the difference between Run Command (non-interactive script execution) and Session Manager (interactive shell) — candidates pick Session Manager because it also avoids SSH, but the question asks for script execution, not a shell.

How to eliminate wrong answers

Option B is wrong because AWS OpsWorks is a configuration management service (Chef/Puppet) for lifecycle management, not an ad-hoc script execution tool, and it is not the right fit for a one-off diagnostic script during an incident. Option C is wrong because EC2 Instance Connect still relies on SSH (it pushes a temporary key), and the engineer has no SSH key access — plus it requires the instance to be reachable on port 22. Option D is wrong because Session Manager opens an interactive shell session, which is not the same as executing a script non-interactively and would require the engineer to manually type or paste commands, adding friction and audit gaps.

113
MCQhard

A company runs a critical application on Amazon ECS with Fargate launch type. The application experiences intermittent connection timeouts when calling an external API. The engineer needs to capture network traffic to diagnose the issue. Which solution is most appropriate?

A.Enable detailed CloudWatch Logs for the ECS task.
B.Enable VPC Flow Logs on the ECS task's elastic network interface.
C.Enable AWS X-Ray tracing on the ECS task.
D.Run tcpdump on the EC2 instance hosting the ECS task.
AnswerB

VPC Flow Logs on the task's elastic network interface capture connection-level metadata — source and destination IP, ports, protocol, and whether the packet was accepted or rejected — for every flow through the ENI. They can be published to CloudWatch Logs or S3 and queried to correlate application failures with security group rule actions or network ACL denials. Because each Fargate task has a dedicated ENI, flow logs are the only one of these options that provides the raw network reachability data needed to diagnose connectivity issues.

Why this answer

VPC Flow Logs capture network traffic metadata at the elastic network interface (ENI) level, which can be analyzed to identify dropped packets, timeouts, and other connectivity issues. For ECS tasks with Fargate, each task gets its own ENI, so enabling VPC Flow Logs on that ENI provides network-level diagnostics without needing to run commands inside the container. Option A is wrong because CloudWatch Logs captures application logs, not network packets.

Option C is wrong because AWS X-Ray traces requests and provides latency information but does not capture raw network traffic. Option D is wrong because Fargate does not use underlying EC2 instances that the user can access; thus, running tcpdump is not possible.

114
MCQmedium

A company runs a production web application on Amazon ECS with the EC2 launch type. A DevOps engineer needs to automatically restart any ECS task that stops unexpectedly due to a container crash. The tasks are part of a service with a desired count of 4. Which configuration should the engineer implement to meet this requirement?

A.Configure the ECS service to use a daemon scheduling strategy.
B.Create a CloudWatch alarm that triggers an AWS Lambda function to restart stopped tasks.
C.Use the ECS service scheduler with a REPLICA scheduling strategy and set the desired count to 4.
D.Set the ECS service's desired count to 4 and enable service auto scaling.
AnswerC

The ECS service scheduler with a REPLICA scheduling strategy maintains the desired number of tasks by automatically launching replacement tasks if any task stops or fails. This is the default and correct approach for long-running services that require high availability. The scheduler continuously monitors the service and reconciles the actual state with the desired state, ensuring that the specified count of 4 tasks is always running.

Why this answer

The ECS service scheduler with a REPLICA scheduling strategy is designed to maintain a specified number of running tasks. When a task stops due to a container crash or other failure, the scheduler automatically launches a new task to replace it, ensuring the desired count is met. This built-in behavior provides high availability without custom automation.

Other options either do not address automatic restart or introduce unnecessary complexity.

Exam trap

The trap here is assuming that service auto scaling is responsible for replacing failed tasks, when in fact the service scheduler handles task replacement and auto scaling only adjusts the desired count based on load.

115
MCQeasy

An Amazon RDS for PostgreSQL instance is running low on storage. The DevOps engineer needs to increase the allocated storage without downtime. Which action should be taken?

A.Modify the DB instance to a larger instance class.
B.Modify the DB instance to increase the allocated storage.
C.Create a snapshot of the DB instance and restore it with larger storage.
D.Launch a new read replica with larger storage and promote it.
AnswerB

Modifying the DB instance to increase allocated storage is the correct approach because Amazon RDS for PostgreSQL supports dynamic storage scale-up without any downtime, provided the instance is in the 'Available' state and not in a pending reboot. You can increase storage either via the AWS Console, CLI, or API, and RDS immediately begins the modification while keeping the instance online. The new storage capacity is reflected after the modification is processed, and you can also enable storage autoscaling to prevent future low-storage alerts.

Why this answer

Amazon RDS supports storage autoscaling and manual storage scaling by modifying the DB instance's allocated storage. Increasing allocated storage is an online operation that does not require downtime; RDS handles the underlying volume expansion transparently. This is the correct action to increase storage without downtime.

Exam trap

DOP-C02 often tests the misconception that increasing storage requires downtime or a snapshot restore — candidates may pick snapshot/restore or read replica promotion, but RDS supports online storage scaling.

How to eliminate wrong answers

Option A is wrong because changing the instance class changes compute and memory, not storage capacity — it does not address the low storage issue. Option C is wrong because creating a snapshot and restoring to a larger instance causes downtime during the restore and cutover, which violates the no-downtime requirement. Option D is wrong because launching a read replica and promoting it also involves downtime during promotion and does not directly increase storage on the existing instance.

116
MCQmedium

A company uses AWS CloudTrail to log API events. During an incident investigation, they need to identify who deleted an S3 bucket. Which CloudTrail feature should be used to retrieve the event details quickly?

A.CloudTrail log file integrity validation
B.CloudTrail Lake
C.CloudTrail Event history
D.CloudTrail Insights
AnswerC

CloudTrail Event history is a built-in feature of the CloudTrail console that displays the last 90 days of management events across all regions, including details like the IAM user identity, event name, and timestamp. It requires no setup and can be searched or filtered directly to find a specific API call and the user who made it. This makes it the fastest and most direct way to answer the question of which user performed a given action.

Why this answer

CloudTrail Event history provides a view of the last 90 days of management events for each AWS region, allowing you to quickly search and filter by resource name (e.g., the S3 bucket name) and event name (e.g., DeleteBucket). This feature is designed for rapid retrieval of recent API activity without needing to query S3 or set up additional infrastructure, making it the fastest option for identifying who deleted an S3 bucket during an incident investigation.

Exam trap

The trap here is that candidates often confuse CloudTrail Insights (which detects unusual activity) with the ability to search for specific past events, or they overcomplicate the solution by choosing CloudTrail Lake when Event history provides the quickest and simplest retrieval for recent events.

How to eliminate wrong answers

Option A is wrong because CloudTrail log file integrity validation is a feature that uses SHA-256 hashing and digital signatures to verify that log files have not been tampered with after delivery; it does not help retrieve or search event details. Option B is wrong because CloudTrail Lake is an analytical data store for running SQL queries on historical events, but it requires configuration and ingestion time, making it slower than Event history for a quick lookup of recent events. Option D is wrong because CloudTrail Insights identifies unusual API activity and potential security threats by analyzing write management events, but it does not provide a direct searchable list of all events or the specific details of who deleted a bucket.

117
MCQhard

A Lambda function 'my-function' is invoked multiple times, but no logs appear in CloudWatch. The DevOps engineer runs the above CLI command and sees that the log group exists but 'storedBytes' is 0. What is the MOST likely cause?

A.The Lambda function is invoked too frequently, causing CloudWatch to throttle log ingestion.
B.The log group's retention policy of 7 days deletes logs immediately after creation.
C.The Lambda execution role lacks permissions to create log streams and put log events.
D.The Lambda function does not have a log group; the one shown belongs to another resource.
AnswerC

The Lambda execution role is assigned to the function via the IAM role's trust policy and must include CloudWatch Logs actions like logs:CreateLogGroup, logs:CreateLogStream, and logs:PutLogEvents on the relevant log group. Without these permissions, the Lambda runtime's logging agent cannot deliver the function's stdout/stderr output to CloudWatch Logs, causing the invocation to finish successfully while producing zero log bytes. This is the classic cause of a Lambda function that runs without any visible logs, and it also explains why no log group or log stream is created.

Why this answer

If the CloudWatch log group exists but storedBytes is 0, the Lambda function is executing but its execution role lacks the IAM permissions required to create log streams and put log events (logs:CreateLogStream, logs:PutLogEvents). Lambda creates the log group automatically on first invocation, but writing log events requires explicit permissions — without them, invocations succeed silently with no logs. This is the classic 'Lambda runs but no logs' scenario.

Exam trap

The trap is assuming an empty log group means the function isn't running or the group is wrong — but Lambda creates the group on invocation regardless of write permissions, so the real issue is IAM authorization for log-stream creation and event writes.

How to eliminate wrong answers

Option A is wrong because CloudWatch Logs does not throttle ingestion in a way that silently drops all logs — throttling produces throttling metrics and errors, not a completely empty log group with zero stored bytes. Option B is wrong because a 7-day retention policy deletes logs after 7 days, not immediately; logs would still appear for a week, and storedBytes would be non-zero during that window. Option D is wrong because the question states the log group exists and is named for the function — Lambda auto-creates /aws/lambda/<function-name>, so the group is correctly associated; the issue is write permissions, not group identity.

118
MCQhard

An application running on Amazon ECS (Fargate) experiences intermittent HTTP 503 errors. The application uses an Application Load Balancer. The ECS service has a desired count of 2. CPU and memory utilization are below 50%. What is the most likely cause?

A.The ECS service Auto Scaling is too aggressive.
B.The target group health check threshold is set too low.
C.The ALB listener rule is misconfigured.
D.The task definition has an incorrect memory hard limit.
AnswerB

The target group's unhealthy threshold is the number of consecutive failed health checks required before ALB deregisters a task and stops sending it traffic. If the threshold is set too low (e.g., 2), a single transient failure—such as a slow response during Fargate instance startup or a momentary network blip—will cause the task to be marked unhealthy and removed from rotation. While it is out of service, if other tasks are also briefly unhealthy or in a draining state, the ALB has no healthy targets and returns HTTP 503 errors until the next successful health check (possibly just seconds later), producing intermittent 503s.

Why this answer

Intermittent HTTP 503 errors from an Application Load Balancer (ALB) typically indicate that the target group health checks are failing, causing the ALB to stop routing traffic to the affected tasks. With a desired count of 2 and low CPU/memory utilization, the most likely cause is a health check threshold set too low (e.g., a low unhealthy threshold count), which makes the ALB prematurely mark tasks as unhealthy during transient issues, leading to no healthy targets and 503 responses. This aligns with the symptom of intermittent errors, as tasks may briefly fail a health check but recover quickly, yet the low threshold causes them to be deregistered.

Exam trap

The trap here is that candidates often attribute 503 errors to resource exhaustion (CPU/memory) or scaling issues, but the question explicitly states low utilization, forcing you to focus on health check configuration as the root cause of intermittent availability.

How to eliminate wrong answers

Option A is wrong because aggressive Auto Scaling would typically cause rapid scaling events, not intermittent 503 errors, and with CPU/memory below 50%, there is no resource pressure to trigger scaling. Option C is wrong because a misconfigured ALB listener rule would cause persistent routing failures (e.g., 404 or 502 errors) for specific paths or hosts, not intermittent 503 errors across all requests. Option D is wrong because an incorrect memory hard limit in the task definition would cause the ECS task to be killed (OOMKilled) or fail to start, resulting in consistent 503 errors or task failures, not intermittent ones.

119
Multi-Selecteasy

A DevOps engineer is setting up an incident response system for a critical application. The engineer needs to ensure that notifications are sent to the appropriate team when specific CloudWatch alarms trigger. Which TWO services can be used to trigger notifications based on CloudWatch alarms? (Choose TWO.)

Select 2 answers
A.AWS Systems Manager
B.Amazon Simple Notification Service (SNS)
C.Amazon Simple Queue Service (SQS)
D.AWS Chatbot
E.AWS Lambda
AnswersB, D

Amazon Simple Notification Service (SNS) is a fully managed pub/sub messaging service designed specifically to deliver messages to a large number of subscribers across multiple protocols, including email, SMS, mobile push, HTTP/S, Lambda, SQS, and AWS Chatbot. For incident response, SNS topics receive alerts from CloudWatch Alarms or EventBridge rules and fan out to all subscribed endpoints, ensuring that on-call engineers are notified reliably via their preferred channel. Its message delivery retry policies, dead-letter queues, and subscription filtering make it the core AWS-native notification service for operational alerts.

Why this answer

Amazon Simple Notification Service (SNS) is correct because CloudWatch alarm actions natively support publishing to an SNS topic, which then fans out notifications via email, SMS, or HTTP/S endpoints to the appropriate team. AWS Chatbot is correct because it integrates with SNS topics and CloudWatch alarms to deliver formatted notifications directly into team chat channels such as Slack or Amazon Chime. AWS Systems Manager is not a notification service; it handles operational tasks like patching and automation, not alarm-driven alerts.

Amazon SQS is a message queue that can receive alarm actions but does not itself send notifications to people. AWS Lambda can be invoked by an alarm but is a compute service that must be coded to send notifications, so it is not a notification service by itself.

Exam trap

DOP-C02 often tests the confusion between services that can be alarm actions (SNS, Lambda, Auto Scaling) and services that deliver human-readable notifications (SNS, Chatbot), causing candidates to pick Lambda or SQS incorrectly.

120
Multi-Selectmedium

A company uses AWS Systems Manager Patch Manager to patch its EC2 instances. After a patch window, some instances report a 'Failed' status. The DevOps engineer needs to investigate the cause. Which actions should be taken? (Choose three.)

Select 3 answers
A.Check the S3 bucket where Patch Manager logs are stored for detailed error messages.
B.Use the Systems Manager Patch Manager dashboard to view the compliance status for each patch.
C.Review CloudTrail logs for the RunCommand API calls.
D.Verify that the SSM Agent is running and up to date on the failed instances.
E.Use AWS Config to check the configuration history of the instances.
AnswersA, B, D

Patch Manager can be configured to write command output to an S3 bucket, either through the Systems Manager settings or per Run Command invocation. These logs capture the full stdout/stderr from the patching process, including detailed error messages, exit codes, and failures from the SSM Agent or patch installation tools. Therefore, checking the S3 bucket is the primary way to obtain the specific reason a patch failed on an instance.

Why this answer

Option A is correct because Patch Manager stores detailed patch execution logs, including error messages for failed patching operations, in an S3 bucket (configured as the patch log destination), which is the primary place to inspect granular failure reasons. Option B is correct because the Patch Manager dashboard in Systems Manager shows per-instance and per-patch compliance status, letting the engineer identify exactly which patches failed and on which instances. Option D is correct because patching is executed by the SSM Agent via RunCommand; if the agent is stopped, outdated, or unable to communicate with the Systems Manager endpoints, the patch operation will fail, so verifying the agent is running and current is a key troubleshooting step.

Option C is not the right focus because CloudTrail records the RunCommand API calls for auditing and API-level errors, but it does not contain the patch-level failure details needed to diagnose why a patch failed. Option E is not appropriate because AWS Config tracks resource configuration changes and compliance rules, not the execution details or error messages of Patch Manager patch jobs.

Exam trap

The trap is selecting CloudTrail or AWS Config because they sound like auditing tools — but CloudTrail only records API calls and Config only records configuration state, neither of which contains the patch execution error details needed to diagnose failures.

121
MCQmedium

A DevOps engineer notices that an EC2 instance is unresponsive and the CloudWatch alarm 'StatusCheckFailed' is in ALARM state. The instance was launched in a private subnet with no public IP. Which action should the engineer take to diagnose the issue without creating a new instance?

A.Modify the security group to allow SSH from the engineer's IP.
B.Use AWS Systems Manager Session Manager to start a session.
C.Check AWS Personal Health Dashboard for instance issues.
D.Use EC2 Serial Console to connect to the instance.
AnswerD

The EC2 Serial Console provides out-of-band server access to the instance's serial port, independent of the network stack and the SSH daemon. It works even when the guest OS is severely degraded or network interfaces are down, making it the correct first step for troubleshooting boot, kernel, and configuration problems. Access is governed through IAM and requires enabling the feature and configuring SSH/RSA keys or a password via IAM policy. Unlike the other options, this is the only one that does not rely on in-band connectivity or a functioning OS network path.

Why this answer

EC2 Serial Console provides out-of-band access to the instance console, allowing troubleshooting of OS-level issues even when network connectivity is lost. Option A is wrong because modifying the security group to allow SSH from the engineer's IP does not help if the instance is unresponsive at the OS level. Option B is wrong because Systems Manager Session Manager requires the instance to have connectivity to the Systems Manager endpoint and the agent running.

Option C is wrong because AWS Personal Health Dashboard provides service health notifications, not instance-level troubleshooting.

122
MCQmedium

A company has a serverless application using AWS Lambda functions and Amazon API Gateway. The application has been running fine, but recently users report that some requests are timing out with a 504 error. The Lambda function's timeout is set to 30 seconds, and API Gateway's integration timeout is 29 seconds. The CloudWatch logs for the Lambda function show that the function executes in under 5 seconds on average. What is the MOST likely cause of the 504 errors?

A.The Lambda function is logging too much data, causing delays in log delivery.
B.API Gateway's timeout is set to less than the Lambda function's timeout.
C.The Lambda function is experiencing concurrency limits and requests are being throttled.
D.The Lambda function's memory is too low, causing cold starts to take longer than the timeout.
AnswerC

When the Lambda function's concurrency limit is reached, synchronous invocations from API Gateway are throttled. For proxy integrations, a throttled request is not immediately rejected—the Lambda service may wait for a free concurrency slot while API Gateway's 29-second integration timeout is already counting down; if no slot frees up in time, API Gateway drops the connection and returns a 504. This explains intermittent 504s even though a 5-second function would otherwise succeed, because the delay occurs before the function code ever runs.

Why this answer

Even though the Lambda function runs in under 5 seconds, if it is throttled due to concurrency limits, the function may not be invoked until after API Gateway's 29-second integration timeout has passed. This causes API Gateway to return a 504 error. Option A is incorrect because excessive logging does not cause 504 errors; it may affect performance but not directly cause timeouts.

Option B is incorrect because the Lambda timeout (30s) is actually higher than API Gateway's integration timeout (29s), so the function could complete before API Gateway times out. Option D is incorrect because cold starts typically add only a second or two, not enough to exceed the 29-second integration timeout when average execution is under 5 seconds.

123
MCQmedium

A company uses AWS CloudTrail to monitor API activity. During an incident, they need to quickly identify any unauthorized IAM role assumption attempts. Which CloudTrail feature should be used to filter and alert on this specific event?

A.Configure VPC Flow Logs to capture traffic to the IAM endpoint.
B.Use S3 event notifications on the CloudTrail bucket for PutObject events.
C.Set up a CloudWatch Logs metric filter on the CloudTrail log group for 'AssumeRole' events.
D.Enable CloudTrail Insights to detect anomalous AssumeRole events.
AnswerD

CloudTrail Insights is the correct choice because it automatically applies machine learning to management events, including IAM AssumeRole, to establish a normal baseline and flag anomalous activity. It requires no manual filter definitions—you simply enable Insights on the trail, and it begins detecting unusual API call rates or error rates, logging them as separate Insights events. This is purpose-built for identifying abnormal role assumption patterns, such as an unexpected spike in AssumeRole calls or a new principal assuming roles outside its normal context.

Why this answer

CloudTrail Insights automatically analyzes management events to detect unusual activity, such as spikes in AssumeRole calls, without requiring manual filter configuration. This feature uses machine learning to establish a baseline and then alerts on deviations, making it ideal for quickly identifying unauthorized role assumption attempts during an incident.

Exam trap

The trap here is that candidates often assume a CloudWatch Logs metric filter (Option C) is the only way to detect specific events, but they overlook that CloudTrail Insights provides automated anomaly detection without requiring manual filter creation, which is faster during an incident.

How to eliminate wrong answers

Option A is wrong because VPC Flow Logs capture network traffic metadata (IP addresses, ports, protocols) at the VPC level, not IAM API calls or CloudTrail events; they cannot filter on specific IAM actions like AssumeRole. Option B is wrong because S3 event notifications on the CloudTrail bucket for PutObject events would trigger on every log file delivery, not on specific event types within those logs, leading to excessive noise and no filtering capability. Option C is wrong because CloudTrail logs are delivered to a CloudWatch Logs log group only if explicitly configured, and a metric filter on that log group for 'AssumeRole' events would require manual setup and ongoing maintenance, whereas the question asks for a feature that can be used quickly during an incident without pre-configuration.

124
Multi-Selecteasy

A DevOps engineer is troubleshooting an issue where an EC2 instance in a private subnet cannot reach the internet. The instance has a route to a NAT gateway. Which TWO of the following should the engineer check? (Choose TWO.)

Select 2 answers
A.The NAT gateway is in the same subnet as the instance
B.The route table of the private subnet has a route to the NAT gateway
C.The internet gateway is attached to the private subnet
D.The instance has a public IP address
E.The security group allows outbound traffic to the internet
AnswersB, E

For a private subnet to reach the internet via a NAT gateway, its route table must contain a route with a destination of 0.0.0.0/0 and a target of the NAT gateway's ID. This route tells the instance's traffic to be forwarded to the NAT gateway, which then performs source NAT using its Elastic IP. Without this specific route, the instance's outbound packets have no defined next hop to the internet, causing connectivity to fail.

Why this answer

Option B is correct because the private subnet's route table must contain a specific route (typically 0.0.0.0/0) pointing to the NAT gateway; without this route, traffic from the instance will never be forwarded to the NAT gateway even if one exists. Option E is correct because security groups are stateful and must have an outbound rule permitting the traffic (e.g., HTTPS/443 or HTTP/80) to the internet; a restrictive outbound rule will silently drop the packets before they leave the instance. Option A is incorrect because a NAT gateway must reside in a public subnet, not the same private subnet as the instance, and it is associated via the route table rather than subnet co-location.

Option C is incorrect because an internet gateway attaches to a VPC, not to a subnet, and private subnets should not have a direct route to an internet gateway. Option D is incorrect because instances in private subnets should not have public IP addresses; outbound internet access via a NAT gateway does not require the instance to have a public IP.

Exam trap

DOP-C02 often tests the misconception that a NAT gateway must be in the same subnet as the instance, or that the instance needs a public IP — both are false, and candidates who confuse NAT gateway placement with instance placement pick the wrong options.

125
MCQmedium

A Lambda function processes SQS messages but sometimes times out after 15 seconds. The function performs a database call that occasionally takes longer. What is the best way to handle this without losing messages?

A.Decrease the SQS visibility timeout to retry faster.
B.Split the batch into smaller batches using partial batch response.
C.Increase the Lambda timeout and increase the SQS visibility timeout, and add a dead-letter queue.
D.Reduce the Lambda reserved concurrency to limit invocations.
AnswerC

The correct remediation is to first raise the Lambda timeout to a value that comfortably covers the actual processing duration for a batch, then set the SQS visibility timeout to at least that same timeout so the queue does not redeliver a message while the function is still processing it. Adding a dead-letter queue to the source SQS provides a safety net: after the configured retries (maxReceiveCount), messages that still fail are diverted to the DLQ, preserving them for analysis instead of silently expiring. This combination directly resolves the timeout issue and handles residual failures cleanly.

Why this answer

The root cause is that the Lambda timeout (15s) is shorter than the occasional database call, and the SQS visibility timeout is likely too short to cover the retry window. Increasing the Lambda timeout gives the function room to finish, increasing the SQS visibility timeout prevents other consumers from picking up the message while it is still being processed, and a dead-letter queue captures messages that repeatedly fail so they are not lost. This combination preserves at-least-once delivery and prevents message loss.

Exam trap

DOP-C02 often tests the relationship between Lambda timeout and SQS visibility timeout — candidates who only increase one or forget the DLQ pick an incomplete fix.

How to eliminate wrong answers

Option A is wrong because decreasing the visibility timeout makes messages visible sooner, causing duplicate processing and potentially more timeouts, not fewer. Option B is wrong because partial batch response helps with batch-level failures but does not address the underlying timeout or the risk of message loss when the visibility timeout expires mid-processing. Option D is wrong because reducing reserved concurrency throttles invocations and can cause SQS messages to age out or hit the DLQ, but it does not fix the timeout root cause.

126
MCQhard

During a deployment, a new application version on an ECS service starts failing health checks. The previous version is still running. The deployment is a rolling update with a 200% percent start. Which ECS feature should the engineer use to automatically revert to the previous version?

A.ECS deployment circuit breaker
B.ECS service auto recovery
C.ECS managed scaling
D.CloudWatch alarm actions
AnswerA

The ECS deployment circuit breaker is a native ECS feature that continuously monitors the health of a service deployment by watching for failed health checks, crashes, or task startup failures. If it detects that the new version is unhealthy, it automatically cancels the deployment and rolls back the service to the previous stable revision, without manual intervention. This makes it the only option that directly addresses deployment failures as part of the ECS service update path.

Why this answer

(ECS deployment circuit breaker) is correct because it automatically detects failed deployments (e.g., health check failures) and triggers a rollback to the previous version. With a 200% percent start rolling update, the new version starts before the old is stopped; if health checks fail, the circuit breaker initiates a rollback. Option B (ECS service auto recovery) recovers from underlying infrastructure failures, not deployment failures.

Option C (ECS managed scaling) adjusts desired count based on load, not deployment health. Option D (CloudWatch alarm actions) can trigger rollback events but is not an ECS built-in feature; it requires custom automation. Therefore, the correct ECS feature for automatic rollback is the deployment circuit breaker.

127
MCQeasy

A DevOps team receives a CloudWatch alarm that an RDS DB instance's CPU utilization has exceeded 90% for 5 minutes. The application is experiencing latency. What is the best immediate step to mitigate the issue?

A.Analyze slow query logs and optimize queries.
B.Enable Multi-AZ deployment for failover.
C.Modify the RDS instance to a larger instance class.
D.Add a read replica to offload read traffic.
AnswerC

Modifying the RDS instance to a larger instance class is the correct immediate response because it directly adds vCPU, memory, and often dedicated EBS bandwidth to handle the current workload. Amazon RDS supports scaling instance classes with minimal downtime, especially for Multi-AZ deployments where a failover masks the restart. This scale-up action provides the compute headroom needed to bring CPU utilization back to acceptable levels and is the standard first step when a CPU alarm indicates that the instance is simply undersized for the traffic.

Why this answer

When an RDS instance's CPU utilization exceeds 90% for 5 minutes and the application is experiencing latency, the immediate step is to scale up the instance to a larger class to provide more CPU capacity. This directly addresses the resource bottleneck without requiring time-consuming analysis or architectural changes, making it the fastest mitigation for an ongoing incident.

Exam trap

The trap here is that candidates often confuse reactive scaling (immediate mitigation) with proactive optimization or architectural changes, leading them to choose slow query analysis or read replicas, which are valid but not immediate fixes for a CPU bottleneck.

How to eliminate wrong answers

Option A is wrong because analyzing slow query logs and optimizing queries is a long-term corrective action, not an immediate mitigation step during an active incident where latency is already occurring. Option B is wrong because enabling Multi-AZ deployment provides high availability and automatic failover, but does not increase CPU capacity or resolve performance issues caused by high utilization. Option D is wrong because adding a read replica offloads read traffic but does not reduce CPU utilization on the primary instance, which is the source of the latency.

128
MCQeasy

A DevOps team is configuring CloudWatch alarms for their production environment. They want to receive notifications when the CPUUtilization metric of an EC2 instance exceeds 90% for three consecutive 5-minute periods. Which combination of settings should they use?

A.Period: 5 minutes; Evaluation periods: 3; Datapoints to alarm: 3
B.Period: 5 minutes; Evaluation periods: 3; Datapoints to alarm: 1
C.Period: 5 minutes; Evaluation periods: 1; Datapoints to alarm: 3
D.Period: 5 minutes; Evaluation periods: 5; Datapoints to alarm: 3
AnswerA

With a 5-minute period, 3 evaluation periods, and 3 datapoints to alarm, the alarm enters ALARM state only when every one of the three most recent 5-minute data points breaches the threshold. This means the metric must be continuously in breach for 15 minutes, filtering out transient spikes and providing a reliable signal of a sustained problem. It is the appropriate setting for production alarms that should page responders only after a consistent degradation.

Why this answer

To alarm when CPUUtilization exceeds 90% for three consecutive 5-minute periods, you need a period of 5 minutes, evaluation periods of 3, and datapoints to alarm of 3. This means CloudWatch evaluates the metric over three consecutive periods, and all three must breach the threshold to trigger the alarm. This configuration exactly matches the requirement.

Exam trap

The trap is confusing 'datapoints to alarm' with 'evaluation periods'; candidates may think that setting datapoints to 1 still requires three periods, but it actually triggers on the first breach.

How to eliminate wrong answers

Option B is wrong because with datapoints to alarm set to 1, the alarm would trigger if any single 5-minute period exceeds 90%, not requiring three consecutive periods. Option C is wrong because with evaluation periods of 1 and datapoints to alarm of 3, it is impossible to have 3 datapoints in 1 period; this configuration is invalid. Option D is wrong because with evaluation periods of 5 and datapoints to alarm of 3, the alarm would trigger if any 3 out of 5 periods breach, not necessarily consecutive, and it evaluates over 5 periods, which is not the requirement.

129
MCQmedium

A company runs a critical web application on EC2 instances behind an Application Load Balancer (ALB) with Auto Scaling. Users report intermittent 503 errors. CloudWatch metrics show that the ALB's 'RequestCount' is normal, but 'HTTPCode_ELB_5XX_Count' spikes. The 'TargetResponseTime' metric shows occasional high latency. Which troubleshooting step should the DevOps engineer take FIRST?

A.Enable and analyze the ALB access logs stored in S3, filtering for 503 errors and correlating with target response times.
B.Increase the desired capacity of the Auto Scaling group to handle more requests.
C.Disable connection draining on the target group to prevent slow-draining instances from causing errors.
D.Review AWS CloudTrail logs for any recent configuration changes to the ALB.
AnswerA

ALB access logs record per-request target status, target IP and timing, so filtering 503s reveals whether targets returned errors or the load balancer itself failed. This directly correlates the ELB 5XX spike with backend latency, satisfying the stem's need to identify the failing tier first.

Why this answer

ALB access logs capture detailed per-request information including target IP, response time, and the specific error (e.g., 503 due to target connection errors or target timeouts). Analyzing these logs filtered for 503s and correlated with TargetResponseTime reveals whether the errors originate from unhealthy targets, connection limits, or slow responses, making it the correct first diagnostic step.

Exam trap

DOP-C02 often tests whether candidates jump to scaling or configuration changes instead of first using observability data (ALB access logs, CloudWatch metrics) to isolate whether errors originate from the ALB or the targets.

How to eliminate wrong answers

Option B is wrong because increasing Auto Scaling capacity does not address the root cause; if targets are slow or unhealthy, more instances may not help and could mask the issue. Option C is wrong because disabling connection draining would cause in-flight requests to be dropped during scale-in or deployment, likely increasing errors rather than reducing them. Option D is wrong because CloudTrail records API-level configuration changes, not per-request error details; it would not explain intermittent 503s tied to target response times.

130
MCQhard

A company runs a critical application on Amazon ECS with Fargate. The application is deployed across multiple Availability Zones and uses an Application Load Balancer (ALB) as the front-end. During a recent incident, users experienced intermittent connectivity failures. The DevOps team suspects that tasks are being stopped due to resource exhaustion. Which combination of metrics and actions should the team use to diagnose and prevent recurrence?

A.Monitor CPU and memory utilization metrics in CloudWatch; increase the task size (CPU and memory) in the task definition.
B.Set up CloudWatch Logs for the application and check for out-of-memory errors; then increase the number of tasks.
C.Monitor NetworkPacketsIn and NetworkPacketsOut metrics in CloudWatch; increase the number of tasks.
D.Monitor the ALB error metrics (5xx count) and scale the ECS service based on request count.
AnswerA

Fargate task-level CPU and memory metrics in CloudWatch reveal resource exhaustion causing task stops; raising the task definition's CPU and memory values gives tasks sufficient headroom, directly addressing the suspected exhaustion constraint and preventing recurrence.

Why this answer

CPU and memory utilization metrics in CloudWatch directly indicate resource exhaustion, which is the suspected cause of tasks being stopped. Increasing the task size (CPU and memory) in the task definition provides more resources per task, preventing the OOM killer or CPU throttling from stopping tasks, without changing the number of tasks or scaling logic.

Exam trap

The trap here is that candidates confuse horizontal scaling (increasing task count) with vertical scaling (increasing task size), assuming that adding more tasks resolves resource exhaustion when the actual issue is insufficient resources per task.

How to eliminate wrong answers

Option B is wrong because while CloudWatch Logs can show out-of-memory errors, increasing the number of tasks does not address resource exhaustion per task—it only distributes load across more tasks, which may still fail if each task is under-provisioned. Option C is wrong because NetworkPacketsIn and NetworkPacketsOut measure network throughput, not CPU or memory exhaustion; high network metrics do not cause tasks to be stopped due to resource exhaustion. Option D is wrong because ALB 5xx errors and request count scaling address load balancing and traffic spikes, not the root cause of tasks being stopped due to insufficient CPU or memory per task.

131
MCQeasy

A security engineer reviews the CloudTrail log entry above and notices that a security group was modified to allow SSH access from anywhere. The engineer wants to ensure that such changes are automatically detected and remediated in the future. What should the engineer do?

A.Configure CloudTrail to send logs to CloudWatch Logs and create a metric filter that alerts on AuthorizeSecurityGroupIngress events with 0.0.0.0/0.
B.Create an IAM policy that denies the ec2:AuthorizeSecurityGroupIngress action if the source IP is 0.0.0.0/0.
C.Create an AWS Config rule that checks security group rules and triggers an AWS Systems Manager Automation document to revoke the ingress rule.
D.Enable Amazon GuardDuty to detect and block such changes in real time.
AnswerC

This is the only option that provides automatic detection and remediation. An AWS Config managed or custom rule can evaluate each security group and mark it NON_COMPLIANT if it contains an ingress rule with 0.0.0.0/0. You can attach that rule to a Systems Manager Automation document, which uses aws:executeAwsApi to call RevokeSecurityGroupIngress and remove the offending rule, or trigger an AWS Lambda function; this remediation runs automatically on each configuration change. Thus it directly satisfies the requirement to revert unauthorized changes.

Why this answer

AWS Config can continuously evaluate security group rules against a custom or managed rule (e.g., restricted-ssh) and, upon detecting a noncompliant rule allowing 0.0.0.0/0 on port 22, trigger an AWS Systems Manager Automation document that automatically revokes the offending ingress rule. This provides both detection and remediation without manual intervention, meeting the requirement for automated detection and remediation.

Exam trap

The trap here is that candidates often confuse detection-only services (like CloudWatch alarms or GuardDuty) with services that can also perform automated remediation (like AWS Config with Systems Manager Automation), leading them to choose options that only alert but do not fix the issue.

How to eliminate wrong answers

Option A is wrong because while CloudTrail logs to CloudWatch Logs with a metric filter can alert on AuthorizeSecurityGroupIngress events with 0.0.0.0/0, this only provides notification (detection) but does not automatically remediate the change. Option B is wrong because an IAM policy that denies ec2:AuthorizeSecurityGroupIngress based on source IP 0.0.0.0/0 is not possible—IAM policies cannot inspect the contents of the API request parameters like the CIDR block; they operate on the action and resource ARN, not on the specific values of the request. Option D is wrong because Amazon GuardDuty is a threat detection service that analyzes VPC Flow Logs, DNS logs, and CloudTrail events for malicious activity, but it cannot block or remediate security group changes in real time; it only generates findings.

132
MCQmedium

A company uses AWS Systems Manager to patch EC2 instances. After a patch window, several instances are unreachable. The engineer checks the SSM Agent logs and finds no errors. What should the engineer do next to diagnose the issue?

A.Restart the SSM Agent on the affected instances.
B.Verify that the patch baseline is associated with the instances.
C.Review the IAM role attached to the instances for sufficient permissions.
D.Check if the instances have outbound internet connectivity to the SSM endpoints.
AnswerD

The SSM Agent maintains a control-plane communication channel over HTTPS to the Systems Manager regional endpoints, either through the public internet, a NAT gateway, or interface VPC endpoints. If outbound access is blocked by security groups, network ACLs, or route tables, the instance cannot register with Systems Manager, receive patch commands, or report status, yet the agent may continue running locally without logging an explicit error. Therefore, verifying that the instance can resolve and connect to endpoints like ssm.<region>.amazonaws.com on port 443 is the first diagnostic step.

Why this answer

The SSM Agent requires outbound internet connectivity to the Systems Manager endpoints (or AWS PrivateLink if configured) to communicate with the service. If the patch window or a security group change blocks this connectivity, instances become unreachable despite the agent logs showing no errors. Option A is wrong because restarting the agent does not help if the issue is network connectivity.

Option B is wrong because the patch baseline association defines which patches to apply, not connectivity. Option C is wrong because IAM permissions are likely correct since the agent logs show no errors; the issue is network-related, not permissions.

133
MCQmedium

A DevOps team uses AWS CodeDeploy to deploy an application to an Auto Scaling group. The deployment fails with an error 'The overall deployment failed because too many individual instances failed deployment'. The team checks the instance logs and finds that the 'BeforeInstall' lifecycle event script returned a non-zero exit code. What is the BEST approach to resolve this?

A.Set the 'ignoreScriptFailure' option to true in the AppSpec file and redeploy.
B.Manually run the script on an instance and then resume the deployment.
C.Fix the script error in the revision and redeploy.
D.Change the deployment configuration to 'AllAtOnce' to speed up deployment.
AnswerC

Fixing the script error in the application revision directly addresses the root cause: the failure is deterministic in the artifact, so every instance that runs the same script will encounter the same problem. A subsequent redeployment with the corrected AppSpec or associated script lets CodeDeploy re-run the lifecycle hooks on all target instances, ensuring consistent and reproducible success. This is the standard remediation because CodeDeploy deployments are immutable artifacts—changes must be made in the revision, not on live instances.

Why this answer

The deployment failed because the BeforeInstall lifecycle event script returned a non-zero exit code, which CodeDeploy interprets as a failure. The root cause is a bug or misconfiguration in the script itself. The best approach is to fix the script in the application revision and redeploy, ensuring the script exits with 0 on success.

This addresses the underlying issue rather than bypassing it.

Exam trap

The trap is looking for a quick workaround like ignoring script failures or manually intervening; candidates must recognize that CodeDeploy has no 'ignoreScriptFailure' option and that the only correct fix is to correct the script and redeploy.

How to eliminate wrong answers

Option A is wrong because 'ignoreScriptFailure' is not a valid AppSpec option in CodeDeploy; there is no such setting to ignore script failures. Option B is wrong because manually running the script on an instance does not fix the revision, and resuming the deployment will still fail on other instances or on subsequent deployments. Option D is wrong because changing the deployment configuration to AllAtOnce only affects the speed and batching of deployment, not the script failure; it may actually cause more instances to fail simultaneously.

134
MCQmedium

A company uses Amazon CloudWatch Logs to store application logs from multiple EC2 instances. The DevOps team needs to create a real-time dashboard that displays the count of ERROR-level log entries across all instances. Which combination of services should be used?

A.Amazon Athena and Amazon QuickSight
B.Amazon Kinesis Data Analytics and Amazon Elasticsearch Service
C.Amazon S3 and Amazon QuickSight
D.CloudWatch Logs Insights and CloudWatch Dashboards
AnswerD

CloudWatch Logs Insights runs SQL-like queries directly against live CloudWatch Logs data, allowing you to count matching log entries with a query such as 'stats count(*) by status' in real time. CloudWatch Dashboards can embed these query results as graph or numeric widgets, which automatically refresh (on up to a 60-second interval) to provide a live operational view. This is the native, serverless, and lowest-latency solution designed specifically for this use case.

Why this answer

CloudWatch Logs Insights allows you to query log data in real time using a query language, and you can create a query that counts ERROR-level entries across multiple log groups. CloudWatch Dashboards can then visualize the results of that query as a widget, providing a real-time dashboard. This combination is native to CloudWatch and requires no additional services or data movement.

Exam trap

DOP-C02 often tests the distinction between real-time log analysis (CloudWatch Logs Insights) and batch analysis (Athena), and candidates may incorrectly choose Athena or Elasticsearch for real-time dashboards.

How to eliminate wrong answers

Option A is wrong because Athena queries data in S3, not CloudWatch Logs directly, and QuickSight is for business intelligence dashboards, not real-time log monitoring. Option B is wrong because Kinesis Data Analytics and Elasticsearch Service would require streaming logs to Kinesis and then to Elasticsearch, adding complexity and latency, and are not the simplest solution. Option C is wrong because exporting logs to S3 and using QuickSight introduces delay and is not real-time.

135
Multi-Selectmedium

Which THREE steps should a DevOps engineer take to troubleshoot an EC2 instance that cannot be reached via SSH? (Choose three.)

Select 3 answers
A.Check the network ACL inbound rules for the subnet.
B.Verify that the corporate firewall allows SSH to the instance.
C.Create an AMI from the instance and launch a new one.
D.Check the security group inbound rules for port 22.
E.Verify that the instance has a public IP address.
AnswersA, D, E

Network ACLs are stateless, subnet-level filters that apply to all traffic crossing the subnet boundary. Even if a security group allows inbound SSH, a deny rule in the NACL inbound table for port 22 will silently drop the connection before it reaches the instance, so inspecting the subnet's inbound NACL rules is an essential first troubleshooting step.

Why this answer

Network ACLs (NACLs) are stateless firewall rules applied at the subnet level. If the inbound rule for ephemeral ports or port 22 is not explicitly allowed, SSH traffic will be dropped even if the security group permits it. Checking NACL inbound rules is a fundamental step in troubleshooting connectivity issues.

Exam trap

The trap here is that candidates often overlook network ACLs and focus only on security groups, or they mistake a recovery action (creating an AMI) for a troubleshooting step, when the correct approach is to systematically verify the layered network controls (NACLs, security groups, and public IP assignment).

136
Multi-Selectmedium

A company uses AWS CloudTrail to log API activity. The security team wants to be alerted when an IAM user creates a new access key. Which TWO steps should be taken to accomplish this? (Choose TWO.)

Select 2 answers
A.Enable CloudTrail Insights to detect unusual activity in the account.
B.Configure the CloudWatch Events rule to send a notification to an Amazon SNS topic.
C.Create an Amazon CloudWatch Events rule that matches the CreateAccessKey API call via CloudTrail.
D.Create an AWS Config rule that checks for access key creation and sends an SNS notification.
E.Use CloudWatch Logs Insights to run a query on the CloudTrail logs and set an alarm.
AnswersB, C

Configuring a CloudWatch Events (now Amazon EventBridge) rule to send a notification to an Amazon SNS topic is a correct and recommended solution for this scenario. You define an event pattern that matches the CreateAccessKey event (source: iam.amazonaws.com, eventName: CreateAccessKey) and set the target to an SNS topic, which then delivers email, text, or other notifications to subscribers. This approach provides near-real-time, event-driven alerts directly from CloudTrail, with no dependence on polling or querying. It is the most direct way to satisfy the requirement of notifying the security team immediately when an access key is created.

Why this answer

Amazon CloudWatch Events (now Events) can be configured to match specific API calls logged by CloudTrail, such as CreateAccessKey. When the rule triggers, it can invoke an SNS topic to send an alert, enabling real-time notification. This approach directly monitors the API activity without additional overhead.

Exam trap

The trap here is that candidates often confuse AWS Config rules (which check resource compliance) with CloudWatch Events (which react to API calls), leading them to choose Option D instead of the correct event-driven approach.

137
MCQeasy

A DevOps team is designing an incident response plan for a critical microservices architecture. They need to automatically collect and analyze logs from all services during an incident. Which solution should they use?

A.Stream logs to Amazon Kinesis Data Firehose and analyze with Amazon OpenSearch Service.
B.Store logs in Amazon S3 and use Amazon Athena to query them.
C.Use AWS Systems Manager Run Command to execute log collection scripts on each instance.
D.Centralize logs in Amazon CloudWatch Logs and use CloudWatch Logs Insights for real-time querying.
AnswerD

Amazon CloudWatch Logs centralizes log streams from EC2 instances, Lambda, and other AWS services via the CloudWatch agent, making logs available for query within seconds of ingestion. CloudWatch Logs Insights provides an interactive, purpose-built query engine that can search, filter, and aggregate log events across multiple log groups using a simple query language, without requiring external infrastructure. This combination supports fast, exploratory incident analysis, real-time alarming via metric filters, and full retention options—making it the most direct and operationally ready choice.

Why this answer

Amazon CloudWatch Logs provides a centralized log management service that integrates natively with AWS services. During an incident, CloudWatch Logs Insights enables real-time, ad-hoc querying and analysis of logs from all microservices without needing to set up additional infrastructure, making it the most efficient solution for incident response.

Exam trap

The trap here is that candidates often over-engineer the solution by choosing complex streaming or analytics services (like Kinesis or Athena) for real-time incident analysis, when the native CloudWatch Logs Insights service is designed specifically for this use case with minimal setup and lower latency.

How to eliminate wrong answers

Option A is wrong because Amazon Kinesis Data Firehose is a streaming data delivery service that requires additional configuration to buffer and deliver logs to Amazon OpenSearch Service, adding latency and complexity not ideal for real-time incident analysis. Option B is wrong because storing logs in Amazon S3 and querying with Athena is designed for batch analytics, not real-time querying, and incurs significant latency due to S3 eventual consistency and Athena's per-query overhead. Option C is wrong because AWS Systems Manager Run Command is a one-time or scheduled command execution tool, not a continuous log collection and analysis solution, and it requires manual intervention to trigger scripts during an incident, which violates the automated incident response requirement.

138
Multi-Selectmedium

A company uses AWS CodeDeploy to deploy a new version of an application to an Auto Scaling group of EC2 instances. During a deployment, the DevOps engineer notices that the deployment is stuck in the 'In Progress' state and some instances are failing to receive the new revision. The engineer needs to troubleshoot the issue. Which TWO actions should the engineer take to identify the cause of the failure? (Choose two.)

Select 2 answers
A.Verify that the instances have the necessary IAM permissions to access the CodeDeploy service and S3 bucket.
B.Restart the CodeDeploy agent service on all instances in the Auto Scaling group.
C.Review the deployment details in the CodeDeploy console to see which lifecycle events failed.
D.Check the CodeDeploy agent logs on the affected instances for error messages.
E.Increase the deployment configuration to deploy to more instances at once.
AnswersC, D

The CodeDeploy console provides a detailed view of each deployment, including the status of each instance and the specific lifecycle events that have succeeded or failed. This information helps pinpoint where the deployment is failing, such as during the BeforeInstall or AfterInstall hooks. It is a quick way to get an overview of the deployment progress and identify the failing stage.

Why this answer

The most direct troubleshooting steps are to check the CodeDeploy agent logs on the affected instances and to review the deployment details in the CodeDeploy console. The agent logs provide detailed error messages specific to each instance, while the console shows the overall deployment status and which lifecycle events failed. Together, they help identify whether the issue is due to script failures, network problems, or missing dependencies.

Exam trap

The trap here is assuming that IAM permissions are always the root cause, but in this scenario, the issue is isolated to some instances, making agent logs and deployment details the more precise diagnostic tools.

139
MCQhard

A company uses Amazon RDS for MySQL with Multi-AZ deployment. The primary DB instance fails, and automatic failover does not occur within the expected 1-2 minutes. The DevOps team needs to quickly restore database availability. What should the team do first?

A.Restore the latest automated snapshot to a new DB instance.
B.Modify the DB instance to change the Multi-AZ setting to enable automatic failover.
C.Connect to the standby instance directly and promote it to primary.
D.Reboot the DB instance with failover selected.
AnswerD

Rebooting the DB instance with the 'Reboot with failover' option selected forces a synchronous failover to the standby instance, typically completing in 60-120 seconds. This is the fastest method to manually initiate a failover while preserving existing data, as the standby is already in sync and becomes the new primary.

Why this answer

When automatic failover does not occur within the expected 1-2 minutes, the fastest way to manually trigger a failover is to reboot the DB instance with the 'Reboot with Failover' option selected. This forces the RDS service to promote the standby replica to the new primary, restoring database availability without waiting for the automated health check to complete. Option D is correct because it directly initiates the failover process, leveraging the existing Multi-AZ setup.

Exam trap

The trap here is that candidates assume they can directly access or promote the standby instance (Option C), but RDS does not expose the standby as a connectable endpoint, and the only manual failover mechanism is the reboot with failover option.

How to eliminate wrong answers

Option A is wrong because restoring from the latest automated snapshot to a new DB instance is a time-consuming process (can take minutes to hours depending on size) and does not utilize the existing standby replica, which is already synchronized and ready to take over. Option B is wrong because modifying the Multi-AZ setting to 'enable automatic failover' is not a valid action; Multi-AZ is already enabled and the setting cannot be toggled to 'enable' failover—failover is inherent to Multi-AZ and the issue is that the automatic health check did not trigger it. Option C is wrong because you cannot directly connect to the standby instance in Amazon RDS Multi-AZ; the standby is not accessible as a standalone database endpoint and there is no 'promote' operation available to the user—RDS manages the standby entirely.

140
Matchingmedium

Match each AWS automation or configuration management tool to its description.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Operational hub for managing AWS resources at scale

Configuration management service using Chef and Puppet

PaaS for deploying and scaling web applications

Infrastructure as Code using templates

Create and manage approved IT service catalogs

Why these pairings

AWS CloudFormation is for IaC, OpsWorks for configuration management, Elastic Beanstalk for PaaS, and CodeDeploy for automated deployments. Distractors swap definitions.

141
MCQeasy

An application running on Amazon ECS experiences intermittent failures. The DevOps engineer wants to capture the application's standard output and error logs and send them to CloudWatch Logs. What is the simplest way to achieve this?

A.Install the CloudWatch Agent in each container.
B.Configure AWS CloudTrail to capture logs.
C.Use the awslogs log driver in the task definition.
D.Write logs to a file and use an S3 bucket with event notifications.
AnswerC

The `awslogs` log driver, configured under the `logConfiguration` element in the ECS task definition, makes Docker send the container's `stdout` and `stderr` directly to a specified CloudWatch Logs group and stream. The ECS agent automatically creates log streams, and you can set `awslogs-group`, `awslogs-region`, and `awslogs-stream-prefix`; the task execution role must have `logs:CreateLogStream` and `logs:PutLogEvents` permissions. This approach requires no changes to the application image, works on both Fargate and EC2 launch types, and provides near-real-time access to logs through the CloudWatch console, CLI, or APIs. It is the native, recommended method for centralizing ECS container logs.

Why this answer

The awslogs log driver is the simplest native integration between Amazon ECS and CloudWatch Logs. By specifying the log driver in the task definition, the ECS container agent automatically captures stdout and stderr from the container and streams them to CloudWatch Logs without any additional agents or custom code.

Exam trap

The trap here is that candidates may overcomplicate the solution by choosing the CloudWatch Agent (Option A) because they assume an agent is needed, but the awslogs log driver is the built-in, simpler mechanism for ECS tasks.

How to eliminate wrong answers

Option A is wrong because installing the CloudWatch Agent inside each container adds unnecessary complexity and overhead; the awslogs log driver handles log forwarding at the container runtime level, making an in-container agent redundant. Option B is wrong because AWS CloudTrail captures API activity and management events, not application stdout/stderr logs, so it cannot fulfill the requirement to capture application output. Option D is wrong because writing logs to a file and using S3 event notifications introduces latency, extra components (S3 bucket, notifications), and does not provide real-time log streaming to CloudWatch Logs; the awslogs driver is far simpler and more direct.

142
Multi-Selecthard

A company uses Amazon RDS for MySQL with Multi-AZ deployment. The database experiences a sudden spike in connections, causing the application to timeout. The DevOps engineer notices that the 'DatabaseConnections' metric is high, but the 'CPUUtilization' is low. Which THREE actions should the engineer take to diagnose the issue?

Select 3 answers
A.Check the 'max_connections' parameter in the DB parameter group and increase it if needed.
B.Add a read replica to offload read traffic.
C.Scale up the DB instance class to handle more connections.
D.Enable the 'general_log' and 'log_output' parameters to capture connection attempts.
E.Enable Performance Insights and review the top SQL queries and sessions.
AnswersA, D, E

The max_connections parameter in the RDS parameter group defines the hard ceiling of allowed simultaneous client connections; if this limit is being hit, legitimate application requests are rejected with 'too many connections'. Reviewing this value against your connection pool settings and raising it if it's legitimately too low can provide immediate relief, but doing so without understanding why the number of connections jumped can exhaust memory or threads because each connection consumes resources. Also, changing the parameter group may require a reboot depending on whether it's a static or dynamic parameter, so monitor after applying.

Why this answer

Option A is correct because in RDS for MySQL the maximum number of simultaneous client connections is governed by the max_connections parameter in the DB parameter group, so if DatabaseConnections is high and the app times out, verifying and raising max_connections (within instance memory limits) directly addresses connection exhaustion. Option D is correct because enabling general_log with log_output set to FILE or TABLE captures every connection attempt and query, letting the engineer confirm whether connections are being rejected or piling up at the server. Option E is correct because Performance Insights provides session-level and SQL-level visibility (db.load, top wait events, active sessions), which is the fastest way to see what those connections are doing when CPU is low.

Option B is not appropriate because a read replica offloads read queries but does not increase the connection limit on the primary endpoint and would not resolve connection timeouts. Option C is not appropriate because CPUUtilization is low, so scaling up the instance class is not justified and would not necessarily raise max_connections without also adjusting the parameter.

Exam trap

DOP-C02 often tests the misconception that high connections must mean high CPU, leading candidates to pick vertical scaling when the real fix is connection-limit and pooling diagnostics.

143
MCQeasy

A company runs a critical web application on AWS. The application is deployed on EC2 instances behind an Application Load Balancer (ALB). The instances are in an Auto Scaling group across multiple Availability Zones. The company uses Amazon Route 53 for DNS with a failover routing policy. Recently, the operations team noticed that during a regional outage, the failover did not trigger as expected, and users experienced downtime. The health checks in Route 53 are configured to check the ALB endpoint. The ALB's health checks are configured to check the instances. What is the MOST likely reason the failover did not work?

A.The ALB remained healthy during the regional outage, so Route 53 did not fail over.
B.The failover routing policy requires manual intervention to switch traffic.
C.Route 53 health checks were not configured for the instance IP addresses.
D.Route 53 cannot failover to a different region when the primary endpoint is still reachable.
AnswerA

The health check associated with the primary ALB endpoint continued to receive a valid HTTP 200 response, so Route 53 considered the resource healthy and did not trigger failover to the secondary region. Even though an AZ was experiencing an outage, the ALB is a regional service that can remain available if it is deployed in other AZs, or the outage may have been in a different part of the region. Since the health check threshold was never breached, Route 53 had no reason to update DNS records.

Why this answer

Route 53 health checks are configured to check the ALB endpoint. During a regional outage, if the ALB itself remains healthy (e.g., the outage only affected the instances but not the ALB), the health check passes, so Route 53 does not trigger failover. This leads to users experiencing downtime because the instances behind the ALB are unhealthy, but Route 53 still directs traffic to the primary region.

Option B is incorrect because failover with Route 53 is automatic when a health check fails; no manual intervention is required. Option C is incorrect because Route 53 health checks do not need to check instance IPs; checking the ALB is sufficient for failover as long as the ALB's health reflects instance health, but in this case it doesn't. Option D is incorrect because Route 53 can failover to a different region when the primary endpoint becomes unhealthy, but the condition for failover (health check failure) was not met.

144
MCQmedium

A company's production EC2 instance running a web application becomes unresponsive. The operations team checks CloudWatch metrics and sees a CPU Utilization spike to 100% for the last 10 minutes. What is the MOST efficient first step to restore service?

A.Check the instance's system logs in CloudWatch Logs to identify the root cause
B.Create an AMI of the instance and launch a new instance from that AMI
C.Reboot the EC2 instance from the AWS Management Console or CLI
D.Terminate the instance and launch a new one from the latest AMI
AnswerC

A graceful reboot via the AWS Management Console or CLI sends an ACPI reboot signal to the guest OS, which restarts the operating system and all application processes while preserving the instance ID, private IP, EBS volumes, and attached Elastic IP. This is the fastest and least disruptive recovery action for transient conditions like memory leaks or a stuck process, and it typically resolves the issue within minutes without requiring reconfiguration.

Why this answer

Rebooting the EC2 instance is the most efficient first step to restore service when the instance is unresponsive due to a CPU spike. A reboot can quickly resolve transient software issues or resource exhaustion without data loss. Option A (checking CloudWatch Logs) delays recovery—investigation should come after restoration.

Option B (creating an AMI and launching a new instance) is time-consuming and unnecessary for initial recovery. Option D (terminating and launching a new instance) risks data loss and is more drastic than rebooting.

145
MCQeasy

A company uses AWS Key Management Service (KMS) to encrypt data at rest. The security team needs to know who attempted to decrypt data using a specific KMS key and whether the attempt succeeded. Which AWS service should the team use?

A.AWS Config
B.KMS key policies
C.AWS CloudTrail
D.CloudWatch Logs
AnswerC

AWS CloudTrail is the correct service because it records all KMS API requests as events, including both management-plane actions like CreateKey and EnableKeyRotation and data-plane operations like Encrypt, Decrypt, and GenerateDataKey. Each event includes details such as the caller identity, source IP address, key ID, and timestamp, which allows you to audit key usage and detect unauthorized access after the fact.

Why this answer

AWS CloudTrail is the correct service because it records all KMS API calls, including Decrypt, Encrypt, and GenerateDataKey, as events in the CloudTrail logs. By examining CloudTrail events for the specific KMS key ID, the security team can see who called the Decrypt API and whether the call succeeded (HTTP 200) or failed (e.g., AccessDenied). This provides the exact audit trail needed for incident response.

Exam trap

The trap here is that candidates confuse AWS Config (which tracks resource configuration) with CloudTrail (which tracks API activity), or they assume KMS key policies themselves provide audit logs, when in fact policies only control permissions and do not generate event records.

How to eliminate wrong answers

Option A is wrong because AWS Config evaluates resource compliance against rules and records configuration changes, but it does not capture API-level actions like decryption attempts. Option B is wrong because KMS key policies define who can use the key and under what conditions, but they do not generate logs or provide historical audit records of decryption attempts. Option D is wrong because CloudWatch Logs can store log data from various sources, but it is not the native service for capturing KMS API calls; CloudTrail is the service that generates those logs, which can optionally be sent to CloudWatch Logs.

146
MCQhard

A DevOps engineer observes the CloudWatch alarm output shown in the exhibit. The alarm is in ALARM state for instance i-0abcd1234efgh5678. The engineer checks the EC2 console and sees that the instance's CPU utilization is currently 10%. What is the MOST likely explanation?

A.The alarm is misconfigured with wrong metric
B.The threshold was set too low
C.The alarm has not yet evaluated enough low datapoints to change state
D.The CPUUtilization metric is not being emitted
AnswerC

For a CloudWatch alarm to transition from ALARM to OK, it must evaluate a specified number of consecutive periods where the metric is below the threshold (e.g., 3 of 3 datapoints below 90%). After the CPU spike dropped back down, the alarm has only recently started seeing low datapoints; until enough consecutive okay datapoints are evaluated, the alarm state remains ALARM. This is CloudWatch's default state-transition behavior to avoid flapping, not a misconfiguration.

Why this answer

CloudWatch alarms evaluate metrics over a specified number of datapoints and periods. Even if the current CPU utilization is 10%, the alarm may still be in ALARM state because it has not yet evaluated enough consecutive low datapoints to transition to OK. The alarm state changes only after the configured number of evaluation periods and datapoints breach or recover from the threshold.

Exam trap

DOP-C02 often tests the misunderstanding that CloudWatch alarms change state immediately when the metric value changes, but they actually require multiple datapoints and evaluation periods to transition, leading candidates to incorrectly blame misconfiguration or missing metrics.

How to eliminate wrong answers

Option A is wrong because if the alarm were misconfigured with the wrong metric, it would likely never enter ALARM state or would behave erratically, but the exhibit shows a specific alarm for CPU utilization. Option B is wrong because a threshold set too low would cause the alarm to trigger more easily, but the current CPU is 10%, which is likely below any reasonable threshold, so the alarm should be OK if the threshold were the issue. Option D is wrong because if the CPUUtilization metric were not being emitted, the alarm would be in INSUFFICIENT_DATA state, not ALARM.

147
MCQmedium

A company uses AWS CodePipeline for CI/CD. A recent deployment to an Amazon ECS service failed because the new task definition referenced an ECR image that does not exist. The pipeline uses a source stage (CodeCommit), build stage (CodeBuild), and deploy stage (ECS). The engineer wants to catch such errors earlier. What should the engineer add to the pipeline?

A.Add an invoke action that calls a Lambda function to check the image.
B.Add a manual approval step before the deploy stage.
C.Add a test stage that runs a script to verify the image exists in ECR.
D.Add a second build stage that re-builds the image.
AnswerC

Adding a test stage immediately after the build stage with a script using the AWS CLI or SDK to call ecr:DescribeImages for the specific repository and image tag validates that the expected artifact exists before deployment. This automated gate runs in every pipeline execution, catches missing or mis-tagged images at the earliest practical point, and fails the pipeline with a clear error rather than letting the deploy stage fail late. It is the correct, lightweight fix that does not alter build behavior.

Why this answer

Adding a test stage that runs a script to verify the image exists in ECR will catch the error earlier in the pipeline, before the deploy stage. This test stage can use AWS CLI or SDK to check if the image tag exists in the ECR repository, and fail the pipeline if it does not. This prevents the deployment from proceeding with a non-existent image.

Exam trap

DOP-C02 often tests the concept of adding validation stages to catch errors early. Candidates may think that a manual approval or a Lambda invoke is the answer, but the most straightforward and integrated solution is a test stage using CodeBuild. The trap is overlooking the simplicity of a script-based test stage.

How to eliminate wrong answers

Option A is wrong because an invoke action calling a Lambda function could work, but it is not a standard pipeline stage type and would require custom Lambda code; a test stage with a script is simpler and more direct. Option B is wrong because a manual approval step would not catch the error; it only requires human approval, and the human might not check the image existence. Option D is wrong because adding a second build stage that re-builds the image does not verify that the image exists in ECR; it might rebuild and push, but if the build fails, it still doesn't catch the missing image issue before deploy.

148
MCQhard

A company uses AWS Lambda functions to process events from Amazon SQS. Recently, the Lambda function has been throttled, causing messages to accumulate in the dead-letter queue (DLQ). The function’s reserved concurrency is set to 100, and the account’s regional concurrency limit is 1000. What is the MOST likely cause of the throttling?

A.The function’s concurrency is fully utilized due to long-running invocations
B.The Lambda function has a cold start issue
C.The SQS queue is not configured as a FIFO queue
D.The reserved concurrency is set too high, exceeding the account limit
AnswerA

Lambda concurrency is the number of in-flight invocations across all resources. When a function's execution time is long, each invocation holds a concurrency slot for the entire duration, so the configured reserved concurrency of 100 can be reached with relatively few requests. Once all slots are occupied, Lambda throttles additional invocations with a 429 error, and the SQS event source mapping receives a failure, causing messages to remain in the queue. This is the classic cause of throttling with long-running workers, not the other options.

Why this answer

The most likely cause of throttling is that the function's reserved concurrency of 100 is fully utilized due to long-running invocations. When invocations take longer to complete, they occupy concurrency for an extended period, preventing new invocations from starting. This leads to messages accumulating in the DLQ.

Option D is incorrect because reserved concurrency of 100 is well below the account limit of 1000, so that is not the cause. Option B is incorrect because cold starts cause latency but not throttling; they do not consume concurrency. Option C is incorrect because the queue type (standard vs.

FIFO) does not directly cause throttling; Lambda can process from both.

149
MCQhard

An organization uses a multi-account AWS environment with AWS Organizations. During an incident, the security team needs to isolate a compromised account by preventing all API calls from that account's root user and IAM users. Which action should be taken?

A.Create a new IAM group with a deny-all policy and add all users to it.
B.Apply a service control policy (SCP) that denies all actions to the affected account's root user and all IAM users.
C.Attach an IAM policy denying all actions to all IAM users in that account.
D.Apply an SCP that denies all actions to the root user only.
AnswerB

An SCP attached to the affected account or its OU in AWS Organizations acts as a guardrail across every principal in the account, including the account root user and all IAM users and roles. By explicitly denying * , the SCP reduces each principal's effective permissions to an empty set, and because SCPs can only be changed by an administrator in the management account, neither the compromised root user nor any IAM user can detach or bypass the lockdown. This provides full account quarantine, which is exactly what the scenario requires.

Why this answer

A service control policy (SCP) is a feature of AWS Organizations that allows you to centrally control permissions for all accounts in your organization. An SCP can be applied to the root of the organization, an OU, or a specific account. When an SCP that denies all actions is applied to an affected account, it restricts permissions for all principals, including the root user and IAM users in that account.

This effectively isolates the compromised account by preventing any API calls. Option A is incorrect because creating a new IAM group with a deny-all policy and adding all users would affect only IAM users, not the root user. Option C is incorrect because an IAM policy attached to IAM users does not affect the root user.

Option D is incorrect because applying an SCP that denies all actions only to the root user would leave IAM users unrestricted, failing to fully isolate the account.

150
MCQeasy

A DevOps engineer notices that an EC2 instance running a critical application is unresponsive. The instance is part of an Auto Scaling group with a minimum size of 2. What is the quickest way to restore service with minimal data loss?

A.Stop and start the instance from the EC2 console.
B.Create a new AMI from the instance and launch a replacement manually.
C.Terminate the instance and let the Auto Scaling group launch a new one.
D.Restore the instance from the most recent EBS snapshot.
AnswerC

Terminating the instance is the correct action because it triggers the Auto Scaling group to detect the lost capacity through its health checks and immediately launch a fresh instance from the launch template, restoring the desired instance count automatically. This approach ensures the replacement is clean, conforms to the group's configuration, and avoids the downtime and statefulness of manual repair, making it the most efficient and reliable recovery mechanism.

Why this answer

Terminating the unresponsive instance triggers the Auto Scaling group to automatically launch a replacement instance, restoring service with minimal data loss. Since the Auto Scaling group has a minimum size of 2, it will immediately detect the terminated instance and launch a new one using the launch template or configuration, ensuring the desired capacity is maintained without manual intervention.

Exam trap

The trap here is that candidates may think stopping and starting the instance (Option A) is the quickest fix, but they overlook that the Auto Scaling group's automated self-healing is designed exactly for this scenario and is faster than any manual recovery method.

How to eliminate wrong answers

Option A is wrong because stopping and starting the instance does not resolve the unresponsive state if the underlying issue is a software or OS hang; it also requires manual action and does not leverage the Auto Scaling group's self-healing capabilities. Option B is wrong because creating a new AMI from the unresponsive instance and manually launching a replacement is time-consuming, may propagate the failure state, and bypasses the automated recovery provided by the Auto Scaling group. Option D is wrong because restoring from the most recent EBS snapshot would revert the instance to a previous state, potentially causing significant data loss, and requires manual steps to attach the volume and launch a new instance, which is slower than letting the Auto Scaling group handle the replacement.

← PreviousPage 2 of 3 · 183 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Incident and Event Response questions.