Courseiva

CCNA Incident and Event Response Questions

75 of 183 questions · Page 1/3 · Incident and Event Response · Answers revealed

1
MCQmedium

A company uses Amazon S3 to store critical data. An incident occurs where an S3 bucket is accidentally deleted. The DevOps engineer needs to recover the bucket and its objects. What should the engineer do?

A.Restore the bucket from AWS CloudTrail event history
B.Contact AWS Support to restore the bucket from a backup if versioning was enabled
C.Recreate the bucket with the same name and restore objects from a previous backup
D.Use the AWS S3 console to undo the deletion
AnswerC

The only way to recover from a deleted bucket is to recreate it and restore objects from a backup. (Note: bucket name uniqueness may require waiting or using a different name.)

Why this answer

S3 bucket names are globally unique and, once deleted, cannot be restored by AWS — the bucket and its configuration are gone. The correct recovery approach is to recreate the bucket with the same name (if the name is still available) and restore the objects from a backup or from versioned copies if versioning was enabled. This is the only viable path to recovery.

Exam trap

DOP-C02 often tests whether candidates understand that S3 bucket deletion is irreversible and that versioning protects objects, not the bucket itself, leading candidates to incorrectly believe AWS Support or CloudTrail can restore a deleted bucket.

How to eliminate wrong answers

Option A is wrong because CloudTrail event history records API activity (who deleted the bucket) but does not store bucket contents or enable restoration. Option B is wrong because AWS Support cannot restore a deleted S3 bucket — AWS does not retain deleted buckets, and Support has no mechanism to recover them. Option D is wrong because there is no 'undo deletion' feature in the S3 console; deletion is permanent.

2
MCQmedium

A DevOps engineer notices that an EC2 instance running a critical web application has been terminated unexpectedly. The instance was part of an Auto Scaling group. Which step should the engineer take FIRST to investigate the root cause?

A.Review AWS CloudTrail logs for TerminateInstances API calls.
B.Look at the EC2 console's 'Termination Protection' setting.
C.Check the application logs on the instance's attached EBS volume (detached and attached to another instance).
D.Verify the Auto Scaling group's scaling policies and scheduled actions.
AnswerA

AWS CloudTrail is the authoritative source for who, when, and how an EC2 instance was terminated: every StopInstances/TerminateInstances API call is recorded as an event with the IAM user or role, source IP, user agent, and request parameters. Console clicks, CLI commands, SDK calls, and automated actions by services such as Auto Scaling or AWS Lambda all generate a TerminateInstances event. Reviewing CloudTrail event history or querying the CloudTrail S3 bucket with Athena reveals the entity that issued the termination and any accompanying error or access-denied information.

Why this answer

CloudTrail records all AWS API activity, including TerminateInstances calls, capturing the identity, source IP, timestamp, and whether the termination was user-initiated, by an Auto Scaling policy, or by another service. Reviewing CloudTrail first establishes who or what triggered the termination before examining instance-level artifacts.

Exam trap

The trap is jumping to Auto Scaling policies or instance settings as the cause — the exam expects you to know CloudTrail is the authoritative source for API-level 'who did what' before examining downstream artifacts.

How to eliminate wrong answers

Option B is wrong because Termination Protection is a setting that, if enabled, would have prevented the termination — checking it after the fact does not explain why the instance was terminated and is not the first investigative step. Option C is wrong because application logs on the EBS volume may show application behavior but not the cause of termination, and detaching/attaching volumes is a later forensic step, not the first. Option D is wrong because scaling policies and scheduled actions are a possible cause, but CloudTrail will reveal whether they were the actual trigger — checking policies without the API audit trail is speculative.

3
Multi-Selecteasy

A DevOps engineer is troubleshooting an Amazon RDS for PostgreSQL instance that is running out of storage. The engineer wants to resolve the issue without downtime. Which TWO actions can achieve this? (Choose two.)

Select 2 answers
A.Create a read replica and promote it to primary.
B.Enable storage auto scaling on the DB instance.
C.Delete old automated snapshots to free up storage.
D.Scale up the DB instance to a larger instance class.
E.Modify the DB instance to increase the allocated storage size.
AnswersB, E

Enabling storage auto scaling on the DB instance automatically increases the allocated storage when free space falls below the configured threshold, up to your specified maximum. Amazon RDS detects low storage conditions and modifies the storage volume dynamically, requiring no manual intervention or downtime. This is the most direct and proactive solution to prevent the instance from reaching a storage-full state.

Why this answer

Enabling storage auto scaling on an Amazon RDS for PostgreSQL instance allows the database to automatically increase its allocated storage when it detects that available storage is running low, preventing out-of-storage errors without requiring manual intervention or downtime. Option E is correct because modifying the DB instance to increase the allocated storage size is a dynamic operation that can be performed without downtime, as Amazon RDS supports online storage scaling for PostgreSQL instances, allowing the change to take effect while the database remains available.

Exam trap

The trap here is that candidates often confuse instance class scaling (compute/memory) with storage scaling, or mistakenly think that deleting snapshots (which are stored separately in S3) can free up space on the DB instance's attached storage volume.

4
MCQeasy

A DevOps engineer receives an alert that an Amazon ECS service is failing to start tasks. The service uses the Fargate launch type. The task definition includes a container that requires port 8080. The security group associated with the service allows inbound traffic on port 8080. What should the engineer check NEXT?

A.Verify that the VPC subnets have a route to a NAT Gateway or Internet Gateway.
B.Confirm that the task definition's container image exists in ECR.
C.Check if the task definition has sufficient CPU and memory allocated.
D.Review the security group rules for outbound traffic.
AnswerA

Fargate tasks must download their container images from Amazon ECR before they enter the RUNNING state. If the service is configured to use private subnets that lack a route to a NAT Gateway (or a subnet with an Internet Gateway for public IPs), the task cannot establish outbound connectivity to the image registry, causing the deployment to hang in PROVISIONING and eventually fail. Verifying the subnet route tables for a 0.0.0.0/0 route to a NAT Gateway or IGW directly addresses the most likely cause of image pull failures.

Why this answer

Fargate tasks require network connectivity to pull container images from ECR (or Docker Hub) and to send logs to CloudWatch. Without a route to a NAT Gateway (for private subnets) or an Internet Gateway (for public subnets), the task cannot pull the image and fails to start. Option B is incorrect: while the image must exist, the immediate symptom of tasks failing to start when the image is missing would be an 'image not found' error, not a generic failure; the security group already allows inbound traffic on port 8080, but outbound connectivity is the issue.

Option C is incorrect: insufficient CPU/memory would cause tasks to enter a 'CPU exhausted' or 'memory exhausted' state, not prevent them from starting entirely. Option D is incorrect: the security group allows inbound traffic, but the issue is about egress connectivity for the task to reach the image registry.

5
MCQhard

A DevOps team is debugging a production incident where an Application Load Balancer (ALB) is returning 503 errors for some requests. The target group instances are healthy. What is the most likely cause?

A.The security group for the ALB does not allow inbound traffic on port 443
B.Health checks are misconfigured to use an incorrect path
C.The deregistration delay setting on the target group is too long
D.Cross-zone load balancing is disabled
AnswerC

An excessively long deregistration delay prolongs the draining period in which a target is excluded from new request routing but still waits for in-flight requests to complete. For example, if the delay is set to 3,600 seconds (the maximum) and a rolling deployment drains instances, replacement targets may not become healthy quickly enough, leaving zero targets in rotation. With no healthy targets, the ALB returns HTTP 503 to new requests even though the underlying instances might function correctly—the bottleneck is the delayed deregistration.

Why this answer

The deregistration delay setting controls how long the ALB continues to send requests to an instance that is being deregistered. If this delay is too long, the ALB may route traffic to an instance that has already stopped accepting connections, resulting in 503 errors even though the health checks pass. Option A is incorrect because a missing security group rule would prevent any traffic from reaching the ALB, causing connection timeouts rather than 503 errors.

Option B is incorrect because the instance health checks are passing (as stated), so the health check path must be correct. Option D is incorrect because disabling cross-zone load balancing affects traffic distribution but does not cause 503 errors.

6
MCQmedium

After deploying a new application version using AWS CodeDeploy, an EC2 instance fails the deployment. The deployment group is configured with an in-place deployment. The engineer sees the error 'ScriptMissing' in the CodeDeploy logs. What should the engineer check?

A.The deployment group's deployment configuration
B.The file path defined in the appspec.yml for the lifecycle hook
C.The security group attached to the instance
D.The AMI used for the EC2 instance
AnswerB

The appspec.yml file is the deployment specification that maps each lifecycle hook (e.g., BeforeInstall, AfterInstall, ApplicationStart) to a script path. The 'ScriptMissing' error occurs precisely when the CodeDeploy agent attempts to execute the script declared in a hook and finds no file at that path, either because the path is mistyped, uses an incorrect relative reference, or the script file was omitted from the deployment revision. Since the agent only knows script locations from appspec.yml, an incorrect path there is the direct and immediate cause of this error, making this option correct.

Why this answer

The 'ScriptMissing' error in CodeDeploy indicates that a lifecycle hook script referenced in appspec.yml cannot be found at the specified path on the instance. The engineer should verify the file path defined for that hook in appspec.yml.

Exam trap

DOP-C02 often tests whether candidates can distinguish configuration-level causes (deployment config, security groups, AMI) from the actual appspec-driven cause of a specific lifecycle error like ScriptMissing.

How to eliminate wrong answers

Option A is wrong because the deployment configuration controls how many instances are deployed to at a time (e.g., one-at-a-time, half-at-a-time), not script resolution. Option C is wrong because security groups affect network access, not the presence of a local script file. Option D is wrong because the AMI determines the base OS image; while a missing script could theoretically stem from a bad AMI, the direct cause of 'ScriptMissing' is the appspec path.

7
MCQmedium

An EC2 instance shows as 'running' in the AWS console, but the system status check is 'impaired'. What is the most likely cause?

A.The instance's security group rules are blocking traffic.
B.The EBS root volume is corrupted.
C.The instance's operating system is not responding.
D.The underlying physical host has experienced a failure.
AnswerD

System status checks specifically monitor the health of the EC2 physical host and the surrounding AWS infrastructure, including power, network connectivity, and host maintenance events. When the underlying physical host experiences a hardware failure, loss of power, or network degradation, the system status check will fail because the host can no longer provide the required virtualization and I/O capabilities. This type of issue is considered host-level and is not fixable by rebooting the instance or making guest OS changes; instead, AWS recommends stopping and starting the instance to move it to a new, healthy host. The symptom described—an instance that is running in the console but failing system status checks—is the classic indicator of a physical host impairment.

Why this answer

The system status check specifically monitors the underlying physical host and hypervisor-level networking. When it is impaired, it indicates a failure of the AWS infrastructure supporting the instance, such as a power loss, network connectivity issue, or hardware degradation on the physical server. Since the instance status check (which monitors the guest OS) is not mentioned, the most likely cause is a host-level failure.

Exam trap

DOP-C02 often tests the distinction between system status checks and instance status checks, and candidates frequently confuse which check corresponds to host-level versus OS-level failures.

How to eliminate wrong answers

Option A is wrong because security group rules only affect network traffic to and from the instance; they do not cause status check failures. Option B is wrong because a corrupted EBS root volume would typically cause the instance status check to fail (or the instance to not boot), not the system status check. Option C is wrong because an unresponsive operating system would cause the instance status check to fail, not the system status check.

8
MCQeasy

A DevOps engineer receives a CloudWatch alarm for high CPU utilization on an EC2 instance. The engineer needs to investigate the cause. Which AWS service can provide a detailed analysis of the running processes and their resource consumption?

A.AWS Config
B.AWS CloudTrail
C.AWS Systems Manager Run Command
D.Amazon Inspector
AnswerC

AWS Systems Manager Run Command lets you securely execute commands on managed EC2 instances via the SSM Agent, without requiring SSH, RDP, or open inbound ports. To diagnose a high-CPU alarm, you can run an OS-level command such as `top -b -n1` or `ps aux --sort=-%cpu` to get a snapshot of per-process CPU and memory usage, and retrieve the output from the console, CLI, or an S3 bucket. This gives you process-level detail directly from the guest OS, making it the correct tool to determine the root cause of the CPU spike.

Why this answer

AWS Systems Manager Run Command lets you remotely execute commands on EC2 instances (via the SSM agent) without SSH or RDP, so you can run tools like top, ps, htop, or perf to inspect running processes and CPU/memory consumption in real time. This is the correct service for interactive, instance-level process investigation triggered by a CloudWatch CPU alarm.

Exam trap

DOP-C02 often tests whether candidates reach for monitoring/audit services (CloudWatch, CloudTrail, Config, Inspector) when the question actually asks for remote command execution — Run Command is the only option that gives OS-level process visibility.

How to eliminate wrong answers

Option A is wrong because AWS Config records resource configuration changes and evaluates compliance rules — it does not provide runtime process or CPU analysis on instances. Option B is wrong because CloudTrail logs API calls and account activity (who did what in AWS), not OS-level process metrics. Option D is wrong because Amazon Inspector is a vulnerability management service that scans instances and containers for CVEs and network exposure, not a live process/resource profiler.

9
Matchingmedium

Match each AWS compute or container service with its description.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Container orchestration service supporting Docker

Managed Kubernetes service

Serverless compute engine for containers

Serverless, event-driven compute service

Automatically adjusts EC2 capacity based on demand

Why these pairings

Correct matches: EC2 for virtual servers, ECS for Docker containers, EKS for Kubernetes, Lambda for serverless computing. Common confusions include mistaking EC2 as serverless or mixing up ECS and EKS.

10
MCQhard

A company uses AWS Config to track resource changes. They want to automatically remediate non-compliant security group rules that allow public SSH access. What is the MOST effective approach?

A.Set up an AWS Config rule that triggers a Lambda function to remove the SSH rule.
B.Use Amazon CloudWatch Events to detect the change and invoke a Lambda function.
C.Use AWS Service Catalog to enforce security group templates.
D.Create an AWS Config rule with an automatic remediation action using AWS Systems Manager Automation.
AnswerD

This is the correct approach because AWS Config rules continually evaluate resources against a desired policy, and when they detect non-compliance they can trigger an automatic remediation action—a Systems Manager Automation document—to fix the resource. In this case the rule (such as the managed RESTRICTED_SSH rule) would flag any security group with port 22 open to 0.0.0.0/0, and the associated SSM Automation document (for example, AWS-RevokeSecurityGroupIngress) would revoke the offending rule automatically. AWS Config tracks the remediation status and retries until the resource becomes compliant, providing a closed-loop, auditable remediation process without manual involvement.

Why this answer

AWS Config can directly associate an AWS Systems Manager Automation document as a remediation action for a non-compliant rule. This approach provides a fully managed, idempotent, and auditable remediation workflow without requiring custom Lambda code or external event orchestration. The automation document can be configured to automatically remove the SSH ingress rule (port 22) from the security group when the Config rule detects non-compliance.

Exam trap

The trap here is that candidates often assume a custom Lambda function (Option A) is the most flexible or effective approach, but AWS Config's native remediation with Systems Manager Automation is the recommended, fully managed, and less error-prone solution for automatic compliance enforcement.

How to eliminate wrong answers

Option A is wrong because while a Lambda function can remove the SSH rule, this approach requires you to write, deploy, and maintain custom code, and it does not natively integrate with AWS Config's remediation lifecycle (e.g., automatic retries, resource exclusion, or rollback). Option B is wrong because Amazon CloudWatch Events (now Amazon EventBridge) can detect security group changes, but it only provides an event notification; it does not include built-in remediation orchestration, compliance evaluation, or the ability to automatically trigger a remediation action directly from a Config rule evaluation. Option C is wrong because AWS Service Catalog is used to provision and govern pre-defined product templates, not to automatically remediate existing non-compliant resources; it cannot react to a Config compliance change or modify an already deployed security group.

11
MCQhard

An application running on an EC2 instance in a private subnet needs to access an S3 bucket. The instance has an IAM role with S3 access. However, the application is failing with timeout errors. The security group allows all outbound traffic, and the NACL allows outbound ephemeral ports. What is the most likely cause?

A.No VPC endpoint for S3
B.Missing route in the route table to an Internet Gateway
C.IAM role does not have correct trust policy
D.Missing HTTP proxy configuration
AnswerA

An EC2 instance in a private subnet has no route to the public internet unless a NAT device or VPC endpoint is provisioned. Since no VPC endpoint for S3 is listed as existing, traffic from the instance to S3 cannot traverse the AWS backbone via the private subnet. A gateway endpoint or interface endpoint for S3 is required to establish private connectivity without leaving the AWS network, making the absent endpoint the root cause of the failure.

Why this answer

A VPC endpoint for S3 (Gateway or Interface) is needed for private subnet access to S3 without NAT. Without a VPC endpoint, the EC2 instance in a private subnet cannot reach S3, resulting in timeout errors. The security group and NACL settings are permissive, so they are not the issue.

Option B is incorrect because the instance is in a private subnet; routing to an Internet Gateway is not necessary and would require a NAT device. Option C is incorrect because the IAM role has the necessary S3 permissions. Option D is incorrect because no HTTP proxy is required for S3 access.

12
MCQmedium

Refer to the exhibit. The DevOps engineer runs the commands and sees the output. What is the most likely issue with the instance?

A.The underlying hardware is having issues (system status check failed).
B.The instance is healthy and no issues exist.
C.The instance is stopped.
D.The instance has a failed status check due to OS-level issues.
AnswerA

The correct interpretation is that the system status check has failed, which indicates an AWS infrastructure-level problem with the physical host running the instance. This includes issues such as a failing disk, network connectivity loss, or hardware component degradation. Even though the instance is running and its instance status check may pass, the impaired system status means the underlying hardware is compromised and AWS may need to repair or replace the host, often requiring a stop/start or recovery action.

Why this answer

The SystemStatus is 'impaired', indicating a problem with the underlying physical host (system status check failed). The InstanceStatus is 'ok', so the OS is functioning normally. Option B is incorrect because the system status check shows impairment, so there is an issue.

Option C is incorrect because the instance is running, not stopped. Option D is incorrect because the impairment is at the system level, not the OS level.

13
Multi-Selectmedium

A company runs a critical application on Amazon ECS with Fargate launch type. During an incident, the DevOps engineer notices that tasks are failing with 'CannotPullContainerError: API error (500)'. Which TWO steps should the engineer take to resolve this issue?

Select 2 answers
A.Attach an EBS volume to the Fargate task for caching.
B.Check that the ECS service role has the required permissions.
C.Ensure that the ECR repository policy allows the task execution role to pull images.
D.Increase the task memory to accommodate the image pull.
E.Verify that the task execution IAM role has the necessary permissions to pull from Amazon ECR.
AnswersC, E

Amazon ECR uses both identity-based policies (attached to the principal, typically the task execution role) and resource-based policies (attached to the repository) to control access. If the repository policy does not explicitly grant the task execution role permission to perform actions like ecr:BatchGetImage and ecr:GetDownloadUrlForLayer, the pull will be denied even if the role's IAM policy allows those actions. The default repository policy is restrictive, and an overly narrow or misconfigured repository policy will cause a 'CannotPullContainerError' during task startup. Therefore, verifying that the ECR repository policy includes an Allow statement for the task execution role is a critical step.

Why this answer

Options C and E are correct. When a Fargate task fails with 'CannotPullContainerError: API error (500)', it typically indicates an issue with pulling the container image from Amazon ECR. The task execution IAM role (E) must have the necessary permissions (ecr:GetDownloadUrlForLayer, ecr:BatchGetImage, ecr:BatchCheckLayerAvailability) to pull images from ECR.

Additionally, if the image resides in a private ECR repository, the repository policy (C) must allow the task execution role to perform those actions. Option A is wrong because Fargate does not support attaching EBS volumes; it stores image layers ephemerally. Option B is incorrect because the ECS service role is used for load balancer integration, not for pulling images.

Option D is incorrect because increasing task memory does not fix image pull errors; memory affects running tasks, not the pull process.

14
MCQmedium

An IAM policy attached to a user is shown in the exhibit. The user reports that they are unable to delete an object in the 'example-bucket' bucket. What is the reason for this?

A.The resource ARN does not match the bucket name
B.The explicit Deny statement overrides the Allow
C.The user does not have permissions to perform s3:DeleteObject
D.The policy has a syntax error
AnswerB

This is the correct explanation. AWS IAM policy evaluation is based on a strict rule: an explicit deny from any applicable policy always overrides any allow, regardless of statement order or the number of allows. Even though the policy contains an Allow that includes `s3:DeleteObject` on the specified bucket and objects, the explicit Deny for the same action takes precedence, resulting in the user being denied. This is fundamental to AWS's default-deny model and is non-negotiable.

Why this answer

An explicit Deny overrides any Allow. The Deny action s3:DeleteObject explicitly denies the delete, even though the Allow all s3 actions includes delete. Option A is wrong because the resource ARN matches.

Option C is wrong because the policy allows all s3 actions, but the Deny blocks delete. Option D is wrong because the policy is valid.

15
MCQmedium

A company is using Amazon RDS for MySQL with Multi-AZ deployment. The database experiences a failover due to an availability zone outage. After the failover, the application team reports that the database endpoint is not resolving to the new primary. What is the most likely reason?

A.The RDS CNAME record was not updated by AWS after the failover.
B.The application is using the read replica endpoint instead of the primary endpoint.
C.The application is using a Route 53 health check that failed and redirected traffic away from the endpoint.
D.The application is using a cached DNS resolution that points to the old primary.
AnswerD

This is correct. After an RDS Multi-AZ failover, the RDS-managed DNS CNAME is updated to point to the new primary instance's underlying IP address. However, if the application's DNS resolver (or the application itself) has cached the old IP address from before the failover, it will continue to try to connect to the old primary instance until the cache expires (based on the DNS TTL, which is typically 30–60 seconds for RDS). A long TTL or a resolver that ignores TTL can keep the stale IP in effect, causing errors (e.g., 'Communications link failure') even though the new primary is healthy.

Why this answer

After an RDS Multi-AZ failover, the DNS CNAME record for the DB instance is updated to point to the new primary in the standby AZ. However, if the application or its DNS resolver has cached the previous DNS resolution, it will continue to use the old IP address, which is no longer reachable. This is a common issue that can be resolved by reducing the TTL on the DNS record or implementing retry logic with DNS re-resolution in the application.

Exam trap

The trap here is that candidates may assume AWS automatically handles DNS propagation instantly or that the CNAME record is not updated, but the real issue is client-side DNS caching, which is a common operational oversight in failover scenarios.

How to eliminate wrong answers

Option A is wrong because AWS automatically updates the RDS CNAME record to point to the new primary after a failover; it is not a manual process. Option B is wrong because the read replica endpoint is a separate endpoint used for read-only traffic; using it would not cause the primary endpoint to fail to resolve, and the application team reported the database endpoint is not resolving, not that it is resolving to the wrong instance. Option C is wrong because Route 53 health checks are not used for RDS DNS resolution; RDS uses its own internal DNS system with CNAME records, and Route 53 health checks are typically used for custom domain names pointing to RDS, not for the default RDS endpoint.

16
MCQmedium

A DevOps team uses AWS CodePipeline to deploy a web application. The pipeline has a manual approval step. During an incident, the deployment is stuck at the approval step because the approver is on leave. The team needs to unblock the pipeline quickly. What is the BEST action to take?

A.Update the pipeline definition to remove the manual approval step temporarily.
B.Use the CodePipeline console to approve the action directly as a different user.
C.Disable the transition to the approval stage and manually run the remaining stages.
D.Retry the action in the approval stage from the CodePipeline console.
AnswerB

The CodePipeline console provides Approve and Reject buttons for any pending manual approval action to any IAM user or role with the codepipeline:ApproveStage permission. The approval is bound to the action and its configured IAM permissions, not to the specific person who originally triggered it, so another authorized user can approve the action and immediately unblock the current pipeline execution without any code or pipeline definition changes.

Why this answer

In AWS CodePipeline, a manual approval action can be approved by any IAM user with the appropriate permissions, not just the designated approver. Using the console to approve as a different user is the quickest way to unblock the pipeline without modifying the pipeline structure. This action is immediate and preserves the pipeline's integrity.

Exam trap

The trap is thinking that only the designated approver can approve, when in fact any authorized IAM user can do so, making option B the fastest resolution.

How to eliminate wrong answers

Option A is wrong because updating the pipeline definition to remove the approval step is a configuration change that requires a pipeline update, which can take time and may have unintended consequences; it's not the best immediate action. Option C is wrong because disabling the transition to the approval stage would skip the stage entirely, potentially bypassing necessary checks, and manually running remaining stages is not a standard feature. Option D is wrong because retrying the action will not bypass the approval; it will simply re-trigger the same approval requirement.

17
MCQhard

A company uses AWS CloudTrail to log API calls. An IAM user's credentials are compromised, and the attacker launches multiple EC2 instances in regions that are not typically used. The security team wants to receive near-real-time notifications of any API calls from this user. What is the MOST effective solution?

A.Create an AWS Config rule that checks for EC2 instances in unauthorized regions
B.Configure CloudTrail to deliver logs to an S3 bucket and enable S3 event notifications to SQS
C.Create a CloudTrail trail that delivers to CloudWatch Logs, then set up a CloudWatch Events rule to invoke a Lambda function that sends an SNS notification
D.Use CloudWatch Logs Insights to query CloudTrail logs every 5 minutes and send results via email
AnswerC

This is the correct near-real-time solution. By delivering CloudTrail events to CloudWatch Logs, events are streamed continuously rather than waiting for S3 log file delivery. A CloudWatch Events (EventBridge) rule can match the specific API call from CloudTrail—filtering on the IAM user, event name, or region—and invoke a Lambda function, which publishes an SNS notification. This end-to-end path operates in seconds, giving you immediate alerting on unauthorized API calls.

Why this answer

The most effective solution for near-real-time notifications of specific API calls is to deliver CloudTrail logs to CloudWatch Logs and create a CloudWatch Events (now EventBridge) rule that triggers a Lambda function to send an SNS notification. This provides immediate alerting based on specific API calls, such as RunInstances, from a particular user.

Exam trap

The trap is choosing AWS Config or S3 event notifications, which are not real-time for API calls. Candidates must remember that CloudWatch Events with CloudTrail integration is the standard for real-time API monitoring.

How to eliminate wrong answers

Option A is wrong because AWS Config rules evaluate resource configurations periodically and are not designed for near-real-time API call notifications. Option B is wrong because S3 event notifications are triggered when objects are created, not when specific API calls occur, and there is inherent latency in log delivery to S3. Option D is wrong because CloudWatch Logs Insights queries are run on demand or on a schedule, and a 5-minute interval is not near-real-time; also, sending results via email is not automated alerting.

18
MCQeasy

An application running on AWS Lambda is experiencing increased error rates. The DevOps engineer needs to quickly identify the root cause. Which AWS service should the engineer use to analyze the logs and errors?

A.AWS X-Ray
B.AWS Trusted Advisor
C.AWS Config
D.AWS CloudTrail
AnswerA

AWS X-Ray traces requests through Lambda and downstream services, exposing latency, errors and their causal path. This gives the engineer the distributed tracing needed to pinpoint the root cause of increased error rates, which CloudWatch logs alone would require manual correlation to achieve.

Why this answer

AWS X-Ray is designed for distributed tracing and root-cause analysis of application errors, including Lambda functions. It captures traces, errors, and latency data across services, making it the right tool to quickly identify where failures originate in a Lambda-based application.

Exam trap

DOP-C02 often tests the distinction between observability services, so candidates pick CloudTrail or Config for application error analysis when those services only cover API auditing and configuration compliance, not runtime tracing.

How to eliminate wrong answers

Option B is wrong because AWS Trusted Advisor provides best-practice checks on cost, security, fault tolerance, and service limits — it does not analyze application logs or errors. Option C is wrong because AWS Config tracks resource configuration changes and compliance, not runtime application errors. Option D is wrong because AWS CloudTrail records API activity and audit events, not application-level error traces or performance data.

19
MCQeasy

A DevOps engineer receives an alert that an Amazon S3 bucket has become publicly accessible. The engineer needs to identify who made the bucket public. Which AWS service should the engineer use to find the API call that changed the bucket policy?

A.Amazon CloudWatch Logs
B.AWS CloudTrail
C.AWS Config
D.Amazon GuardDuty
AnswerB

AWS CloudTrail is the correct answer because it provides a complete audit history of API activity across AWS, including actions performed on S3 buckets such as CreateBucket, PutBucketPolicy, and DeleteBucket. For every management event, CloudTrail captures the IAM user or role, the source IP address, the request parameters, and the response returned, enabling you to answer exactly who did what and when. CloudTrail can also be configured to log data events for object-level S3 operations (GetObject, PutObject) for deeper forensic analysis.

Why this answer

CloudTrail records all S3 API calls, including bucket policy changes. Option A is wrong because CloudWatch monitors metrics. Option C is wrong because Config records configuration changes but not the identity.

Option D is wrong because GuardDuty detects threats but doesn't log API calls.

20
MCQhard

An organization uses AWS CloudFormation to manage infrastructure. During an incident, a stack update fails with 'UPDATE_ROLLBACK_FAILED' status. The engineer needs to bring the stack to a consistent state without losing data. What is the BEST approach?

A.Use the 'ContinueUpdateRollback' API to skip the resource that caused the failure.
B.Create a new stack from the same template and migrate resources.
C.Manually correct the resource configuration that caused the failure, then perform a stack update.
D.Delete the stack and then recreate it from the same template.
AnswerA

The `ContinueUpdateRollback` API is the designed recovery action when a CloudFormation stack is stuck in the `UPDATE_ROLLBACK_FAILED` state. By invoking it with the `ResourcesToSkip` parameter, you explicitly instruct CloudFormation to skip the specific resource that caused the rollback failure, allowing the stack to return to a stable `UPDATE_COMPLETE` state. This bypasses the problematic resource without requiring manual intervention. It is the recommended and least disruptive method to recover from a failed stack update.

Why this answer

The 'ContinueUpdateRollback' API is the best approach because it allows the stack to resume the rollback process, skipping the resource that caused the failure, and bringing the stack to a consistent 'UPDATE_ROLLBACK_COMPLETE' state without manual intervention or data loss. This API is specifically designed for the 'UPDATE_ROLLBACK_FAILED' status, enabling you to skip resources that cannot be rolled back (e.g., due to a non-reversible change) while preserving the rest of the stack's state.

Exam trap

The trap here is that candidates often choose manual correction (Option C) thinking they can fix the resource and retry the update, but they overlook that the stack is in a failed rollback state that blocks further updates until the rollback is resolved, making 'ContinueUpdateRollback' the only viable path to a consistent state without data loss.

How to eliminate wrong answers

Option B is wrong because creating a new stack from the same template and migrating resources is time-consuming, risks data loss during migration, and does not address the immediate need to recover the existing stack to a consistent state. Option C is wrong because manually correcting the resource configuration and then performing a stack update assumes the failure is fixable via a new update, but the stack is stuck in 'UPDATE_ROLLBACK_FAILED' and cannot accept further updates until the rollback is completed or continued; this approach may also lead to configuration drift and potential data loss. Option D is wrong because deleting the stack would destroy all resources, including any data stored in them (e.g., databases, EBS volumes), which violates the requirement to avoid data loss.

21
MCQmedium

A DevOps engineer notices that an EC2 instance running a critical application is unresponsive. CloudWatch alarms for CPU utilization and memory usage did not trigger. The engineer checks the system logs and finds an 'Out of memory: Kill process' error. What is the MOST likely cause of the missed alarms?

A.The CloudWatch agent is not installed or configured to collect memory metrics.
B.The instance is using instance store volumes instead of EBS, which prevents metric collection.
C.The CloudWatch metrics retention period is set to 1 day, so old alarms were deleted.
D.The EC2 instance's root EBS volume is encrypted, blocking CloudWatch agent logs.
AnswerA

The CloudWatch agent is required to emit in-guest memory metrics. The default EC2 monitoring collects only hypervisor-level metrics such as CPU utilization, disk I/O, and network throughput; memory utilization is not visible from the hypervisor and must be reported by the agent inside the OS. If the agent is not installed or its configuration does not define memory as a collected metric, the MemoryUtilization metric will be absent, causing any alarm relying on it to remain in INSUFFICIENT_DATA and never trigger.

Why this answer

The 'Out of memory: Kill process' error indicates the OS OOM killer terminated a process due to memory exhaustion, but the CloudWatch memory alarm did not fire because the default CloudWatch metrics for EC2 do not include memory utilization. Memory metrics are only available if the CloudWatch agent is installed and configured to collect them via the mem_used_percent metric. Therefore, the most likely cause is that the agent was not installed or not configured for memory metrics.

Exam trap

The trap is assuming that all EC2 metrics are available by default — candidates forget that memory and disk space require the CloudWatch agent, and they may incorrectly blame storage type or encryption for the missing alarm.

How to eliminate wrong answers

Option B is wrong because instance store vs. EBS has no bearing on CloudWatch metric collection — the CloudWatch agent runs at the OS level and can collect metrics regardless of the underlying storage type. Option C is wrong because CloudWatch metric retention (e.g., 1-day for high-resolution) does not delete alarms; alarms persist and evaluate against available data points, and retention affects historical data, not alarm existence.

Option D is wrong because EBS encryption does not block CloudWatch agent logs or metrics — encryption at rest is transparent to the OS and the agent, and CloudWatch agent communicates over the network, not via the EBS volume directly.

22
MCQhard

A company runs a containerized application on Amazon ECS with Fargate launch type. The application experiences periodic spikes in response times. The CloudWatch metrics show high CPU and memory usage for the tasks during these spikes. What is the MOST effective approach to handle these spikes?

A.Use a larger Fargate task size to handle the spikes
B.Increase the CPU and memory limits for the ECS task definition
C.Set up a scheduled scaling action to add tasks during peak hours
D.Configure target tracking scaling policies for the ECS service using CPU or memory utilization
AnswerD

Configuring target tracking scaling policies for the ECS service lets AWS automatically scale the number of tasks in response to actual load, using CloudWatch metrics such as CPUUtilization or MemoryUtilization. The policy works by maintaining a specified target value (for example, 70% CPU) and triggers scale-out or scale-in based on aggregated utilization across the service. This is the recommended pattern for handling spikes in containerized workloads on ECS, as it provides reactive, horizontal scaling without manual intervention.

Why this answer

Target tracking scaling policies for Amazon ECS services using CPU or memory utilization are the most effective approach because they dynamically adjust the number of tasks in response to real-time demand, automatically adding capacity during spikes and removing it when load subsides. This aligns with the AWS Well-Architected Framework's principle of elasticity, ensuring the application scales out precisely when high CPU/memory usage is detected, without manual intervention or over-provisioning.

Exam trap

The trap here is that candidates often confuse increasing task-level resources (CPU/memory limits) with horizontal scaling, or assume scheduled scaling is sufficient, failing to recognize that unpredictable spikes require reactive, metric-based auto scaling.

How to eliminate wrong answers

Option A is wrong because using a larger Fargate task size (e.g., increasing vCPU and memory) addresses the spike by over-provisioning resources for each task, which is cost-inefficient and does not scale the number of tasks; it may still hit limits if the spike exceeds the larger size. Option B is wrong because increasing CPU and memory limits in the task definition only raises the maximum resources a single task can use, but does not add more tasks to handle increased load; it can also lead to throttling if the underlying Fargate platform cannot allocate the requested resources. Option C is wrong because scheduled scaling actions are predictive, not reactive, and cannot adapt to unpredictable spikes; they may add tasks at the wrong times, leading to either insufficient capacity during unexpected spikes or wasted resources during off-peak periods.

23
MCQmedium

A DevOps engineer supports a microservice running on Amazon EKS. During an incident, pods in one node group are repeatedly evicted and the engineer sees node memory pressure conditions. The team wants future incidents to trigger automatic replacement of unhealthy nodes and alerting without manual intervention. Which combination should the engineer implement?

A.Deploy the AWS Node Termination Handler and configure a CloudWatch alarm on node memory metrics to page the team.
B.Configure the Cluster Autoscaler to add nodes when pods are pending and rely on the pod eviction events for alerting.
C.Increase the pod memory requests and limits so the kubelet stops evicting pods, and add a PodDisruptionBudget for alerting.
D.Use a managed node group with health checks enabled and configure the cluster to replace nodes that fail health checks, plus CloudWatch alarms on node conditions for alerting.
AnswerD

Managed node groups perform health checks and can automatically replace nodes that fail them, which addresses unhealthy nodes without manual action. Pairing that with CloudWatch alarms on node condition metrics provides the alerting path, satisfying both automatic replacement and notification requirements for future incidents.

Why this answer

The requirement is twofold: automatically replace unhealthy nodes and alert on node conditions. A managed node group with health checks enabled replaces nodes that fail health checks, and CloudWatch alarms on node condition metrics deliver the alerting. The other approaches either only scale on pending pods, only handle planned interruptions, or tune pod resources without addressing node-level failure.

Exam trap

The trap here is confusing Cluster Autoscaler or Node Termination Handler behavior with node health remediation, when neither automatically replaces a node that is failing due to memory pressure.

24
MCQmedium

A company runs a containerized application on Amazon ECS with Fargate launch type. The application is behind an Application Load Balancer (ALB). The operations team notices that the ALB's 5xx error rate increases periodically. The ECS service is configured with a target tracking scaling policy based on CPU utilization. The CloudWatch logs from the application show no errors. The health check on the ALB is configured to hit the /health endpoint. What is the MOST likely cause of the 5xx errors?

A.The ECS tasks are running on an underlying host that is being patched.
B.The health check endpoint is returning a 503 status due to a dependency failure.
C.The target tracking scaling policy is not responding quickly enough to traffic spikes.
D.The application is throwing exceptions that are not logged.
AnswerB

This is the correct explanation because the ALB health check probes the configured endpoint and expects a successful status code, and a 503 returned by that endpoint indicates the application's dependency (for example, a database or downstream API) is unavailable. When the health check receives 503 for all tasks, the ALB marks the targets unhealthy and stops routing traffic, ultimately returning HTTP 503 to clients. The symptom therefore matches the health check failure, not an application exception or scaling issue.

Why this answer

The ALB periodically reports 5xx errors because the health check endpoint /health is returning a 503 status code, likely due to a transient dependency failure. When the health check fails, the ALB considers the target unhealthy and returns a 503 (or 502) to clients. The application code itself logs no errors because the failure occurs in a downstream dependency that the health check probes, not in the main application logic.

Option A is incorrect because ECS with Fargate does not expose underlying host patching; the ALB would not detect such patching as 5xx errors. Option C is incorrect because a target tracking CPU scaling policy, even if slow, would not cause 5xx errors—it would affect performance but not directly trigger health check failures. Option D is incorrect because the application logs show no errors, ruling out unlogged exceptions as the source.

25
MCQmedium

An organization uses AWS Systems Manager Incident Manager for incident response. They have created a response plan with an engagement plan that pages the on-call engineer via SMS. The engineer acknowledges the incident but then does not take any further action. What is the BEST way to automate escalation?

A.Manually re-page the on-call engineer with a higher urgency.
B.Use Amazon CloudWatch Events to trigger a second SMS if the incident is not resolved within a time frame.
C.Create an AWS Lambda function that checks the incident status and pages the next responder if no action is taken.
D.Configure an escalation plan in the response plan that pages a secondary contact after a specified timeout.
AnswerD

An escalation plan is a first-class component of an AWS Systems Manager Incident Manager response plan. You define a total duration (e.g., 10 minutes) and add engagement targets such as individual contacts, chat channels, or a full on-call schedule; if the primary responder does not acknowledge the incident within that window, Incident Manager automatically pages the secondary/next-level contact using the configured contact channels (SMS, voice, mobile push). This is the intended solution because it is fully managed, idempotent, and does not require any custom code or additional AWS services. Escalation plans also support multiple levels, and you can set the engagement duration per target to progressively move up the chain until someone acknowledges.

Why this answer

Systems Manager Incident Manager supports escalation plans with timeouts and multiple engagement levels. Option A is wrong because manual re-paging does not provide automated escalation and requires human intervention. Option B is wrong because CloudWatch Events can trigger actions but is not the built-in escalation mechanism within Incident Manager; that capability is provided by escalation plans.

Option C is wrong because while a Lambda function could be used, it is not the best or simplest approach; Incident Manager natively supports escalation plans.

26
Multi-Selectmedium

A DevOps engineer is investigating a security incident where an EC2 instance was compromised. The engineer needs to collect forensic data without losing volatile information. Which TWO actions should the engineer take? (Choose two.)

Select 2 answers
A.Detach the EBS volumes and attach them to a forensic instance.
B.Retrieve the instance metadata from the console.
C.Create a snapshot of the attached EBS volumes.
D.Collect a memory dump from the instance before stopping it.
E.Terminate the instance immediately to prevent further access.
AnswersC, D

Creating a snapshot of the attached EBS volumes is a core forensic preservation technique because it captures the full disk state—including deleted file remnants, user-space artifacts, logs, and malware binaries—without stopping the instance. Unlike a live filesystem copy, an EBS snapshot is crash-consistent (or application-consistent with pre-freeze), providing a point-in-time image that can be analyzed on a separate forensic instance without risking further alteration of the original evidence. This must be done before any stop/termination, since those actions can change or destroy disk data.

Why this answer

Option C is correct because creating an EBS snapshot captures a point-in-time, crash-consistent copy of the attached volumes, preserving disk-based forensic evidence (file system, logs, malware artifacts) without altering the running instance. Option D is correct because volatile data such as RAM contents, running processes, network connections, and encryption keys exist only in memory and are lost once the instance is stopped or terminated, so a memory dump must be collected first. Option A is wrong because detaching EBS volumes from a running instance is not supported and would disrupt the live system before volatile data is captured.

Option B is wrong because instance metadata contains only configuration data (instance ID, AMI, IAM role, user data) and provides no forensic value for the compromise. Option E is wrong because terminating the instance destroys both volatile memory and the instance store, and may also delete EBS volumes depending on the DeleteOnTermination setting, irreversibly destroying evidence.

Exam trap

The trap is choosing actions that seem to preserve evidence (detaching volumes, terminating) but actually destroy volatile data or alter the scene; the exam tests knowledge of order of volatility and proper forensic sequence.

27
MCQhard

An incident response team is analyzing an IAM policy attached to a role used by a forensic tool. The tool needs to create snapshots of EBS volumes during an incident. However, when the tool runs from an IP address in the 203.0.113.0/24 range, the CreateSnapshot API call fails with an access denied error. What is the MOST likely cause?

A.The policy does not grant ec2:CreateSnapshot on specific resource ARNs, only on all resources.
B.The aws:ViaAWSService condition is set to false, but the tool is invoked by an AWS service such as Systems Manager, making the condition evaluate to true and denying access.
C.The Deny statement explicitly denies ec2:DeleteSnapshot, but the error is for CreateSnapshot, so it is unrelated.
D.The source IP address 203.0.113.0/24 is not included in the Condition block, so access is implicitly denied.
AnswerB

The aws:ViaAWSService global condition key is true when an AWS service, such as Systems Manager, makes the API call on the principal's behalf rather than the principal making a direct call. The policy's condition requires this key to be false, so when the tool is invoked via Systems Manager the actual value is true and the Allow statement does not match. With no other matching Allow, the request is implicitly denied, which is exactly the error observed.

Why this answer

The aws:ViaAWSService condition key evaluates to true when an API call is made by an AWS service on behalf of a principal. If the policy sets this condition to false, it denies any call that originates from an AWS service (e.g., Systems Manager Automation). In this scenario, the forensic tool is likely invoked by Systems Manager, causing the condition to evaluate to true and triggering the deny, even though the source IP is allowed.

This explains why CreateSnapshot fails with access denied despite the IP being in the allowed range.

Exam trap

The trap here is that candidates focus on the IP address condition and assume the error is due to an IP mismatch, overlooking the subtle aws:ViaAWSService condition that denies calls made through AWS services even when the source IP is allowed.

How to eliminate wrong answers

Option A is wrong because granting ec2:CreateSnapshot on all resources ("*") would not cause an access denied error; the error is due to a condition key, not resource ARN specificity. Option C is wrong because a deny on ec2:DeleteSnapshot is unrelated to the CreateSnapshot failure; IAM evaluates deny statements independently per action. Option D is wrong because the source IP 203.0.113.0/24 is included in the Condition block (as stated in the question), so implicit denial does not apply; the error is caused by the aws:ViaAWSService condition, not the IP condition.

28
Multi-Selecthard

A DevOps team is investigating a performance issue where an application's response time spiked during a deployment. The deployment used AWS CodeDeploy to update an Auto Scaling group. Which THREE actions should the team take to identify the root cause? (Choose THREE.)

Select 3 answers
A.Review the CodeDeploy deployment logs for errors.
B.Examine application logs on the new EC2 instances launched during the deployment.
C.Review the CodeDeploy deployment group configuration.
D.Check AWS CloudTrail for any unauthorized API calls during the deployment.
E.Compare CloudWatch metrics for the Auto Scaling group before and after the deployment.
AnswersA, B, E

CodeDeploy deployment logs capture lifecycle event hook failures, invalid scripts, and resource timing issues during deployment (e.g., BeforeInstall/AfterInstall failures) that can leave instances in a degraded state causing performance hits. Any failed or aborted deployment step may cause new instances to be registered with incomplete configuration, leading to CPU/memory pressure or misrouted traffic. Scrutinizing these logs pinpoints whether the performance spike correlates with deployment execution timeouts, file overwrite errors, or instance registration failures.

Why this answer

CodeDeploy deployment logs contain detailed information about the deployment process, including any errors or failed steps that could impact performance. Option B is correct because application logs on the new EC2 instances can reveal errors, misconfigurations, or resource contention that may have caused the spike. Option E is correct because comparing CloudWatch metrics (e.g., CPU utilization, latency, request count) before and after the deployment helps pinpoint changes that correlate with the performance issue.

Option C is wrong because reviewing the deployment group configuration—which defines how deployments occur (e.g., traffic routing, instance selection)—is unlikely to directly identify the root cause of a performance spike; it is more relevant for deployment strategy issues. Option D is wrong because AWS CloudTrail records API calls for auditing and security, not application performance; unauthorized API calls are unlikely to cause a transient performance spike during deployment.

29
Multi-Selecthard

An e-commerce platform uses Amazon DynamoDB as its primary database. During a flash sale, the application experiences throttling errors. The operations team needs to implement a solution to handle sudden traffic spikes while keeping costs under control. Which TWO actions should the team take? (Choose two.)

Select 2 answers
A.Increase the read and write capacity units manually before the sale.
B.Switch from on-demand to provisioned capacity with auto scaling.
C.Implement DynamoDB Accelerator (DAX) to cache read-intensive data.
D.Use application-level retry logic with exponential backoff to handle throttling gracefully.
E.Enable DynamoDB Streams and replicate data to a read replica.
AnswersC, D

Implementing DynamoDB Accelerator (DAX) is correct because it puts a write-through, in-memory cache directly in front of your DynamoDB table. DAX intercepts repeated read requests—such as product details, pricing, or inventory views during a sale—and serves them in microseconds, dramatically reducing the read capacity units consumed by the table. This offloading lowers the chance of throttling your primary table while keeping latency low for read-heavy traffic.

Why this answer

DynamoDB Accelerator (DAX) is an in-memory cache that reduces read latency from milliseconds to microseconds, offloading read requests from the main DynamoDB table. During a flash sale, caching read-intensive data (e.g., product details) with DAX reduces the number of read capacity units consumed, helping to avoid throttling while keeping costs under control by not requiring a permanent increase in provisioned capacity.

Exam trap

The trap here is that candidates often confuse DynamoDB Streams with read replicas, or assume that provisioned capacity with auto scaling is always cost-effective for spikes, when in fact on-demand capacity is designed for unpredictable traffic and avoids the cold-start throttling risk of auto scaling.

30
Multi-Selecteasy

Which TWO actions should be taken to ensure a highly available and resilient architecture for a critical web application on AWS? (Choose two.)

Select 2 answers
A.Enable Amazon CloudFront with multiple origins.
B.Use an Auto Scaling group to maintain a desired number of instances.
C.Use a Multi-AZ RDS deployment with read replicas.
D.Store backups in a different AWS Region.
E.Deploy the application across multiple Availability Zones.
AnswersB, E

An Auto Scaling group with a desired capacity continuously monitors instance health via EC2 status checks and optionally Elastic Load Balancing health checks, automatically terminating and relaunching failed instances. By spreading the ASG across multiple Availability Zones and setting the desired count, you ensure that if an instance or an entire AZ fails, replacement capacity is launched to maintain the required number of instances, making this a core high-availability mechanism.

Why this answer

Correct: B and E. Option B ensures that the desired number of EC2 instances is maintained, providing automatic scaling and fault tolerance. Option E deploys the application across multiple Availability Zones, which protects against an AZ failure.

Option A (CloudFront) enhances content delivery but does not directly ensure high availability of the web application. Option C (Multi-AZ RDS with read replicas) improves read performance and provides disaster recovery, but write availability depends on the primary instance. Option D (backups in a different region) is for disaster recovery, not for immediate availability.

31
MCQhard

A DevOps engineer is configuring an AWS Lambda function that processes messages from an Amazon SQS queue. The function must handle transient failures gracefully and avoid reprocessing the same message multiple times. The engineer sets the maximum receives to 3 and configures a dead-letter queue (DLQ) for the source queue. After several days, the engineer notices that some messages are being processed more than once, even though they were successfully processed. What is the MOST likely cause of this issue?

A.The dead-letter queue is misconfigured, causing messages to be sent back to the source queue.
B.The SQS queue's visibility timeout is shorter than the Lambda function's execution time.
C.The Lambda function is not idempotent, so it processes the same message differently each time.
D.The Lambda function's timeout is set too low, causing it to fail and retry messages.
AnswerB

If the visibility timeout is shorter than the time it takes for the Lambda function to process the message, the message becomes visible again in the queue before the function completes. Another Lambda invocation can then receive and process the same message, leading to duplicate processing. This is a common misconfiguration. The visibility timeout should be set to at least the function's timeout plus a buffer to prevent this. Even if the function eventually succeeds, the duplicate processing has already occurred.

Why this answer

The most likely cause is that the SQS queue's visibility timeout is shorter than the Lambda function's execution time. When a Lambda function polls an SQS queue, it receives a batch of messages and each message becomes invisible for the duration of the visibility timeout. If the function takes longer than that timeout to process a message, the message becomes visible again and can be picked up by another Lambda invocation.

This results in duplicate processing. To resolve this, set the visibility timeout to at least the function's timeout plus a buffer, and ensure the function is idempotent.

Exam trap

The trap here is focusing on idempotency as the cause of duplicate processing, when idempotency is a mitigation for duplicates, not the reason they occur.

32
MCQhard

A company uses AWS Organizations with multiple accounts. The security team notices that an IAM user in the production account has been making changes to security group rules that are not compliant with the company's policy. The team wants to automatically revoke any non-compliant security group rules and notify the security team. What is the MOST efficient way to achieve this?

A.Create a CloudWatch alarm on the SecurityGroupEvent metric to notify the security team.
B.Apply a Service Control Policy (SCP) that denies changes to security groups in the production account.
C.Set up a CloudTrail trail that logs security group modifications and use Amazon Detective to analyze the changes.
D.Use an AWS Config managed rule to detect non-compliant security group rules, and configure an automatic remediation action with AWS Systems Manager Automation.
AnswerD

AWS Config's managed rules continuously evaluate security group configurations against policies such as 'restricted-ssh' or 'vpc-sg-open-only-to-a-specific-port'. When a rule detects a non-compliant inbound rule, a configured AWS Systems Manager Automation document, such as AWS-DisablePublicAccessForSecurityGroup, automatically removes or tightens the offending rule, providing the required detect-and-remediate control loop without manual intervention.

Why this answer

AWS Config managed rules can continuously evaluate security group rules against a desired policy (e.g., disallowing SSH from 0.0.0.0/0). When a non-compliant change is detected, AWS Config can trigger an automatic remediation action using an AWS Systems Manager Automation document that revokes the offending rule. This provides both detection and automated correction without manual intervention, making it the most efficient solution.

Exam trap

The trap here is that candidates often confuse detective controls (CloudTrail, CloudWatch alarms) with corrective controls (AWS Config remediation), and fail to recognize that SCPs are preventive and cannot selectively revoke existing non-compliant rules.

How to eliminate wrong answers

Option A is wrong because CloudWatch alarms on the SecurityGroupEvent metric can only notify on the occurrence of an event, not automatically revoke the non-compliant rule; it lacks remediation capability. Option B is wrong because Service Control Policies (SCPs) apply to all IAM users and roles in an account and cannot selectively revoke specific security group rules after they are created; SCPs are preventive, not detective or corrective, and would block all security group changes, which may be too restrictive. Option C is wrong because CloudTrail logs and Amazon Detective can analyze changes after the fact but cannot automatically revoke non-compliant rules; they provide visibility and investigation, not automated remediation.

33
Multi-Selectmedium

A company uses Amazon CloudWatch for monitoring. The operations team wants to receive an alert when an EC2 instance's status check fails for 2 consecutive minutes. Which THREE resources should the team configure? (Choose three.)

Select 3 answers
A.CloudWatch Events rule
B.CloudWatch Logs
C.CloudWatch alarm
D.EC2 StatusCheckFailed metric
E.Amazon SNS topic
AnswersC, D, E

A CloudWatch alarm is the correct monitoring construct to watch a metric such as StatusCheckFailed or CPUUtilization. It evaluates the metric against a threshold over a specified number of evaluation periods and transitions to ALARM, OK, or INSUFFICIENT_DATA, then triggers a configured SNS action. The alarm is the central component that converts raw metric data into an operational notification, making it the appropriate mechanism for alerting on the instance's status check result.

Why this answer

To alert when an EC2 instance's status check fails for 2 consecutive minutes, you need to create a CloudWatch alarm on the StatusCheckFailed metric (options C and D). The alarm needs to send notifications via an SNS topic (option E). Option A (CloudWatch Events rule) is not used for metric-based alerts; CloudWatch Events triggers on events or schedules, not metric thresholds.

Option B (CloudWatch Logs) is for log data, not metrics.

Exam trap

A common trap is confusing CloudWatch Events with CloudWatch Alarms. CloudWatch Events are for event-driven actions based on state changes or schedules, not for monitoring metric thresholds over time. Metric alarms require the CloudWatch Alarm resource.

34
MCQmedium

A DevOps team observes that an Amazon CloudFront distribution is returning HTTP 504 errors for a small percentage of requests. The origin is an Application Load Balancer (ALB) that distributes traffic to EC2 instances. The team has already checked the ALB's access logs and found that the ALB returns 200 OK for all requests. What should the team investigate NEXT?

A.Check the ALB target group health check settings and ensure instances are healthy.
B.Examine the request headers in CloudFront logs to identify unusual patterns.
C.Review the CloudFront cache hit ratio and optimize caching strategies.
D.Check the ALB's idle timeout settings and compare with CloudFront origin timeout.
AnswerD

The correct root cause is a timeout mismatch: CloudFront waits for an origin response within its configured origin timeout (default 30 seconds), while the ALB has an idle timeout (default 60 seconds) that can close the connection to the backend if no bytes flow. If the ALB idle timeout is shorter than CloudFront’s timeout, the ALB can terminate a long-running request before the backend finishes, causing CloudFront to receive no response and return 504. Since the ALB logs show 200 for completed requests, the occasional slow requests are being cut off by this idle setting; aligning the ALB idle timeout to be greater than CloudFront’s origin timeout (or tuning backend latency) is the fix.

Why this answer

The ALB returns 200 OK for all requests, so the origin itself is not failing. However, HTTP 504 errors from CloudFront typically indicate that the origin (ALB) is not responding within CloudFront's timeout window. The ALB's idle timeout (default 60 seconds) can cause the ALB to close idle connections, while CloudFront's origin timeout (default 30 seconds) is separate.

If the ALB's idle timeout is shorter than the time CloudFront waits for a response, the ALB may close the connection before CloudFront receives the full response, leading to a 504. Option D directly addresses this mismatch.

Exam trap

The trap here is that candidates assume 504 errors always indicate an unhealthy origin, but the ALB logs show 200 OK, so they incorrectly focus on health checks or caching instead of the timeout mismatch between CloudFront and the ALB.

How to eliminate wrong answers

Option A is wrong because the ALB access logs show 200 OK for all requests, meaning the target group health check settings and instance health are not the issue—healthy instances would still produce 200 responses. Option B is wrong because examining request headers in CloudFront logs for unusual patterns would not explain a consistent 504 error when the ALB itself is responding successfully; the issue is at the transport layer, not the application layer. Option C is wrong because a low cache hit ratio would cause more origin requests but not 504 errors; optimizing caching strategies would reduce origin load but not fix a timeout mismatch between CloudFront and the ALB.

35
Multi-Selectmedium

A company uses AWS Lambda with an Amazon DynamoDB trigger. Recently, the Lambda function started failing with 'ProvisionedThroughputExceededException' errors. The DevOps team needs to mitigate the issue. Which TWO actions should the team take? (Choose TWO.)

Select 2 answers
A.Increase the Lambda function's reserved concurrency
B.Disable DynamoDB Streams on the table
C.Enable DynamoDB Accelerator (DAX) for the table
D.Increase the DynamoDB table's write capacity
E.Reduce the batch size for the DynamoDB stream event source mapping
AnswersD, E

DynamoDB throttling occurs when write requests exceed the provisioned write capacity (WCUs) of the table. If the Lambda function writes processed items back to the same table, insufficient WCUs will cause ProvisionedThroughputExceededException, leading to retries and stream processing failures. Increasing the write capacity reduces throttling, allowing the stream-triggered writes to succeed and the function to make progress.

Why this answer

To mitigate 'ProvisionedThroughputExceededException' errors when a Lambda function is triggered by DynamoDB Streams, two actions are effective. Option D: Increase the DynamoDB table's write capacity to handle the write demand from the stream processing. Option E: Reduce the batch size for the DynamoDB stream event source mapping to lower the number of writes per invocation, reducing the chance of exceeding throughput.

Option A is wrong because Lambda reserved concurrency controls how many concurrent executions Lambda can run, but the issue is DynamoDB throttling, not Lambda capacity. Option B is wrong because disabling DynamoDB Streams would stop the trigger entirely, which is not a mitigation. Option C is wrong because DynamoDB Accelerator (DAX) is an in-memory cache for reads, not writes, and does not affect write throughput.

36
Drag & Dropmedium

Drag and drop the steps to implement a blue/green deployment using AWS CodeDeploy.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

First create the application and deployment group, then configure blue/green settings, then deploy, then validate, then reroute traffic.

37
MCQeasy

A company uses Amazon CloudFront to serve static content from an S3 bucket. Users report that they see outdated content even after the engineer has updated the files in the S3 bucket. What should the engineer do to ensure users see the latest content?

A.Create an invalidation for the updated file paths.
B.Change the S3 bucket policy to allow public access.
C.Reduce the TTL for the CloudFront distribution.
D.Delete and recreate the CloudFront distribution.
AnswerA

CloudFront caches objects at edge locations until their TTL expires, so after you update files in Amazon S3 the edge locations continue serving the old copies. An invalidation request explicitly removes the specified file paths from every CloudFront edge cache, forcing subsequent requests to fetch the latest version from the S3 origin. You can target an individual file with /path/file.js or use a wildcard like /images/* to clear a whole directory. This is the intended, immediate mechanism for propagating content updates without waiting for natural cache expiry.

Why this answer

Creating a CloudFront invalidation for the specific file paths forces the edge locations to fetch the updated content from the S3 origin immediately, ensuring users see the latest files. Option B is incorrect because changing the bucket policy to allow public access does not affect CloudFront's cache; it only controls direct access to the bucket. Option C is incorrect because reducing the TTL affects how long new content is cached but does not clear already-cached outdated content.

Option D is incorrect because deleting and recreating the distribution is an overly disruptive solution; a simple invalidation suffices.

38
MCQeasy

An application running on Amazon EC2 instances behind an Application Load Balancer (ALB) is experiencing intermittent 503 errors. The target group health checks are failing. The DevOps engineer checks the instance logs and finds that the application is running but taking longer than 30 seconds to respond. What is the MOST likely cause?

A.The Auto Scaling group's scaling policy is too aggressive, causing frequent instance replacements.
B.The security group for the ALB does not allow inbound traffic from the internet.
C.The health check timeout is set too low, causing the ALB to mark instances unhealthy.
D.The EC2 instances are running out of memory and the application is crashing.
AnswerC

A health check timeout set too low can cause the ALB to mark otherwise functional instances as unhealthy when the application's response time occasionally exceeds the timeout. The ALB health check settings include an interval, timeout, and unhealthy threshold; if a slow application misses the timeout a few consecutive times, the target is deregistered and the ALB returns 503 Service Unavailable when no healthy targets remain. This matches the symptom of intermittent errors under load, as the application may respond normally at times but exceed the timeout during traffic spikes.

Why this answer

The most likely cause of intermittent 503 errors and failing health checks is that the health check timeout is set too low. The application takes longer than 30 seconds to respond, but if the health check timeout is set to a value less than the application response time, the ALB will mark the instance as unhealthy, leading to 503 errors. Adjusting the health check timeout to accommodate the application's response time would resolve the issue.

Exam trap

DOP-C02 often tests the misconception that 503 errors are always due to security groups or instance failures, overlooking health check timeout misconfigurations.

How to eliminate wrong answers

Option A is wrong because an aggressive scaling policy would cause instances to be replaced frequently, but that would not directly cause health checks to fail due to response time; it might cause other issues. Option B is wrong because if the security group did not allow inbound traffic, the ALB would not be able to reach the instances at all, resulting in consistent failures, not intermittent ones. Option D is wrong because if instances were running out of memory and crashing, the application would not be running, but the logs show it is running and taking longer than 30 seconds.

39
MCQmedium

A company uses AWS CloudTrail to monitor API activity. The security team notices that an IAM user 'dev-user' deleted an S3 bucket. They need to quickly identify the source IP address of the delete request. Which CloudTrail feature should they use to find this information?

A.Use CloudTrail Lake to query the event and extract the IP address from the userIdentity field.
B.Check S3 server access logs for the bucket deletion event.
C.Enable CloudTrail Insights to analyze unusual activity.
D.Search the CloudTrail event history for the delete event and review the sourceIPAddress field.
AnswerD

CloudTrail event history retains management events for the last 90 days, and each event record includes the sourceIPAddress field, which identifies the IP address from which the API call was made. Searching event history for the DeleteBucket event, or a similar deletion event, and expanding the event details will show the sourceIPAddress in the raw event record. This is the direct and correct way to retrieve the requester's IP for a specific CloudTrail event.

Why this answer

CloudTrail Event History retains 90 days of management events and each event record includes a sourceIPAddress field that captures the IP address from which the API call was made. Searching Event History for the DeleteBucket event and inspecting sourceIPAddress is the fastest way to identify the source IP without additional setup.

Exam trap

DOP-C02 often tests whether candidates know that S3 server access logs do not capture control-plane API calls like DeleteBucket — only CloudTrail does.

How to eliminate wrong answers

Option A is wrong because CloudTrail Lake is a paid, query-based service for long-term retention and SQL analysis; while it can return the IP, it is overkill and not the 'quick' method for a recent event. Option B is wrong because S3 server access logs record bucket-level access requests but do not capture the DeleteBucket API call itself (which is a control-plane action logged by CloudTrail, not S3 access logs). Option C is wrong because CloudTrail Insights detects anomalous API call rates and error rates — it does not provide per-event source IP details.

40
MCQeasy

A company uses AWS CloudFormation to deploy infrastructure. A stack update fails with the error 'UPDATE_ROLLBACK_FAILED'. What should the engineer do to resolve this?

A.Retry the stack update with the same parameters.
B.Delete the stack and recreate it.
C.Ignore the error and continue using the stack.
D.Use the 'ContinueUpdateRollback' operation to fix the resource that caused the failure.
AnswerD

The ContinueUpdateRollback operation is the correct recovery mechanism specifically designed for stacks that fail during an update and subsequently fail to roll back automatically, landing in UPDATE_ROLLBACK_FAILED. It instructs CloudFormation to resume the rollback to the last known good state, optionally skipping resources that are causing repeated failures via the ResourcesToSkip parameter after manually fixing them. This restores the stack to a stable, updatable condition without destroying the underlying infrastructure.

Why this answer

When a CloudFormation stack update fails and rollback also fails, the stack enters UPDATE_ROLLBACK_FAILED, a terminal state where CloudFormation cannot automatically recover. The ContinueUpdateRollback API operation tells CloudFormation to resume rolling back the stack, optionally skipping specific resources that are stuck, so the engineer can manually fix the problematic resource and then complete the rollback. This is the documented recovery path for this state.

Exam trap

The trap is thinking a failed rollback can be fixed by simply retrying or deleting the stack, when the correct action is the ContinueUpdateRollback recovery operation.

How to eliminate wrong answers

Option A is wrong because retrying the same update with identical parameters will fail again for the same underlying reason and does not address the stuck rollback. Option B is wrong because deleting and recreating the stack destroys all resources and state, causing data loss and downtime, and is a last resort rather than the correct recovery procedure. Option C is wrong because ignoring the error leaves the stack in a non-operational, inconsistent state where further updates are blocked and resources may be partially configured.

41
MCQmedium

An application running on Amazon ECS Fargate is experiencing intermittent 'CannotPullContainerError' errors. The task definition references a Docker image in a private Amazon ECR repository. The task execution role has the 'AmazonECSTaskExecutionRolePolicy' policy attached. What is the most likely cause?

A.The Fargate task is in a private subnet without a NAT gateway or VPC endpoint
B.The task execution role does not have sufficient permissions
C.The ECS service is not configured with Auto Scaling
D.The ECR repository is not in the same region as the ECS cluster
AnswerA

Fargate tasks provisioned in a private subnet have no route to the internet unless a NAT gateway is configured in a public subnet. Because ECR's API and Docker Hub need outbound HTTPS access, and image layers are fetched from Amazon S3, the task's image pull fails without a NAT gateway or VPC endpoints for ECR (both API and DKR) and S3. This manifests as a 'CannotPullContainerError' or 'ResourceInitializationError' in the task's stopped reason. Adding a NAT gateway or the appropriate VPC endpoints resolves the issue.

Why this answer

The 'CannotPullContainerError' occurs when the ECS task cannot retrieve the container image from ECR. Since the task execution role has the 'AmazonECSTaskExecutionRolePolicy' attached, which includes the necessary permissions (ecr:GetAuthorizationToken, ecr:BatchCheckLayerAvailability, ecr:BatchGetImage, ecr:GetDownloadUrlForLayer), the issue is not permissions. The most likely cause is that the Fargate task is running in a private subnet that lacks a route to the internet (via NAT gateway) or a VPC endpoint for ECR.

Without either, the task cannot reach the ECR API to pull the image. Option A is correct. Option B is wrong because the policy provides sufficient permissions.

Option C is irrelevant; Auto Scaling does not affect image pulling. Option D is less likely because ECR repositories are typically in the same region, and cross-region pulls would still be possible with proper permissions and networking.

42
MCQhard

A company runs a multi-tier web application on AWS. The application consists of an Application Load Balancer (ALB), an EC2 Auto Scaling group (ASG) for web servers, and an Amazon RDS Multi-AZ DB instance. The ASG uses a launch template with Amazon Linux 2 and a user data script that installs the web application and connects to the RDS database using a static password stored in the user data. Recently, the security team discovered that the user data script is exposed in the EC2 console and could be viewed by anyone with EC2 describe-instances permissions. The team wants to remediate this immediately without causing downtime. The ASG is configured with a min size of 2, max size of 6, and desired capacity of 4. The application is currently under load. Which option describes the best course of action?

A.Create a new launch template version that retrieves the password from AWS Secrets Manager. Update the ASG to use the new template version and perform an instance refresh with a minimum healthy percentage of 100%.
B.Immediately modify the user data on each running EC2 instance to remove the password, then update the launch template to reference AWS Secrets Manager.
C.Update the existing launch template to use AWS Secrets Manager for the database password. The ASG will automatically apply the change to existing instances.
D.Delete the existing launch template and create a new one with secrets from AWS Secrets Manager. Then terminate all running instances and let the ASG launch new ones.
AnswerA

This action creates a new launch template version that retrieves the password from AWS Secrets Manager, then performs an instance refresh with a minimum healthy percentage of 100%. This replaces instances one by one without downtime, remediating the security issue on all instances.

Why this answer

It uses an instance refresh with a minimum healthy percentage of 100% to replace instances without downtime, while the new launch template version retrieves the password from AWS Secrets Manager, eliminating the static password exposure. This approach ensures that the security vulnerability is remediated immediately without disrupting the running application under load.

Exam trap

The trap here is that candidates assume updating the launch template automatically propagates to existing instances, but in reality, the ASG only applies the launch template to new instances, so an instance refresh or manual replacement is required to remediate existing instances.

How to eliminate wrong answers

Option B is wrong because modifying user data on running instances does not change the launch template, so any new instances launched by the ASG will still use the exposed static password; also, manually editing instances is not scalable and risks configuration drift. Option C is wrong because updating the launch template does not automatically apply changes to existing instances; the ASG only uses the launch template for new instances, so existing instances remain vulnerable until replaced. Option D is wrong because terminating all running instances at once would cause downtime, violating the requirement to avoid disruption, and the ASG would launch replacements based on the new template, but the immediate termination is not safe under load.

43
MCQhard

A critical application is deployed on Amazon EKS. The DevOps team notices that pods are failing with 'CrashLoopBackOff' status. The team needs to capture the application logs before the pod restarts to debug the issue. Which approach should the team use?

A.Use 'kubectl logs' command immediately after the crash
B.Configure a sidecar container to stream logs to Amazon CloudWatch Logs
C.Store logs in a ConfigMap
D.Use 'kubectl exec' to access the container and check logs
AnswerB

A sidecar container, such as aws-for-fluent-bit, runs alongside the application in the same pod and streams log events to Amazon CloudWatch Logs in near real time. Even if the main application container crashes and immediately restarts, the sidecar remains operational, and the log events already shipped are safely retained in CloudWatch, enabling immediate debugging and automated alarms. This decouples log shipping from the application's lifetime and provides durable, searchable history that survives pod restarts and rescheduling.

Why this answer

Configuring a sidecar container to stream logs to Amazon CloudWatch Logs ensures that logs are persisted and available for debugging even if the pod crashes and restarts. This approach decouples log collection from the pod's lifecycle, allowing the DevOps team to analyze logs from the crash without needing to capture them in real-time. It aligns with the incident response best practice of centralized logging for ephemeral environments like EKS.

Exam trap

The trap here is that candidates assume 'kubectl logs' can always capture logs from a crashed pod, but they overlook that CrashLoopBackOff causes the container to restart, overwriting previous logs in the default Kubernetes logging setup (which only retains logs for the current container instance).

How to eliminate wrong answers

Option A is wrong because 'kubectl logs' retrieves logs from the current container instance, and after a crash and restart, the logs from the previous instance are lost unless the container has a logging driver that persists them; in a CrashLoopBackOff scenario, the pod may restart before logs can be captured. Option C is wrong because ConfigMaps are designed for storing configuration data (e.g., environment variables, configuration files), not for dynamic application logs, and they have a size limit of 1 MiB, making them impractical for log storage. Option D is wrong because 'kubectl exec' requires a running container to execute commands, and in a CrashLoopBackOff state, the container may be in a crash loop or not running, making exec inaccessible.

44
Multi-Selecthard

A security team is investigating a potential data exfiltration from an S3 bucket. They need to identify which IAM user accessed a specific object and whether the access was from a known IP address. Which THREE AWS services or features should they use together to conduct this investigation?

Select 3 answers
A.AWS Config
B.VPC Flow Logs
C.AWS CloudTrail
D.S3 server access logs
E.Amazon Athena
AnswersC, D, E

AWS CloudTrail is correct for this investigation because it records S3 data-plane API calls such as GetObject, PutObject, and ListObjects, along with the requesting IAM identity, source IP, and timestamp. Object-level logging must be enabled on the trail or bucket, but once active, it provides a comprehensive audit trail of exactly which objects were retrieved. This makes CloudTrail a primary evidence source for S3 data exfiltration.

Why this answer

AWS CloudTrail (C) is correct because it records S3 data-plane API calls such as GetObject, including the IAM identity (user or role) that made the request, the source IP address, and the timestamp, which directly answers who accessed the object and from where. S3 server access logs (D) are correct because they provide detailed, object-level records for every request against the bucket, including the requester, source IP, request URI, and HTTP status, giving a second authoritative source for the specific object access. Amazon Athena (E) is correct because it lets the team query CloudTrail logs and S3 access logs stored in S3 using standard SQL, so they can efficiently correlate the IAM user, object key, and source IP across large log datasets.

AWS Config (A) is not appropriate because it tracks resource configuration changes and compliance, not individual object access events or source IPs. VPC Flow Logs (B) capture IP traffic metadata at the ENI/subnet level and do not identify IAM users or S3 object-level requests, so they cannot answer who accessed the object.

Exam trap

The trap is selecting VPC Flow Logs or AWS Config because they sound like network/audit tools — but Flow Logs lack IAM identity and object names, and Config tracks configuration not access events, so neither can answer 'which IAM user accessed this object from which IP.'

45
MCQmedium

A company stores sensitive data in Amazon S3. A security audit reveals that several S3 buckets are publicly accessible. The DevOps engineer needs to implement a solution that automatically detects and alerts on any S3 bucket that becomes public. Which AWS service should the engineer use?

A.Amazon Macie
B.S3 Block Public Access
C.AWS Trusted Advisor
D.AWS Config
AnswerA

Amazon Macie is a fully managed data security service that uses machine learning and pattern matching to automatically discover sensitive data, such as personally identifiable information (PII), stored in Amazon S3. It continuously monitors bucket policies and access control lists (ACLs) to detect publicly accessible buckets and generates security findings that can be sent in real time via Amazon EventBridge. This makes it a detective control perfectly suited for alerting auditors on public exposure.

Why this answer

Amazon Macie uses machine learning to discover, classify, and protect sensitive data, and it can automatically detect and alert on S3 buckets that become publicly accessible. Option B is incorrect because S3 Block Public Access is a preventive control, not detective. Option C is incorrect because AWS Trusted Advisor can check for public buckets but does not provide real-time alerts.

Option D is incorrect because AWS Config can track bucket policies but does not automatically alert on public access.

46
MCQhard

A company runs a multi-tier web application on EC2 instances behind an Application Load Balancer. The application experiences intermittent 503 errors during peak traffic. The Auto Scaling group is configured with a step scaling policy based on CPU utilization. CloudWatch metrics show that CPU utilization never exceeds 70%, but the ALB target group reports that some targets are unhealthy. What is the MOST likely cause?

A.The application health check endpoint is returning HTTP 5xx or timing out.
B.The ALB is misconfigured with an incorrect security group blocking traffic to the targets.
C.The ALB connection draining settings are too short, causing in-flight requests to fail.
D.The step scaling policy is too aggressive and is terminating instances prematurely.
AnswerA

The ALB health check is performing HTTP requests against the configured health check path on each target. When the application returns any 5xx status code (or the request times out because the app hangs under load), the ALB marks that target as unhealthy. With an intermittent application bug (e.g., a memory leak or a connection pool exhaustion), the health check will fail sporadically, causing the ALB to periodically stop routing traffic to that instance. If all targets become unhealthy at the same time, the ALB returns HTTP 503 Service Unavailable to clients, which exactly matches the intermittent nature of the reported errors.

Why this answer

ALB target groups mark targets unhealthy when the configured health check fails — typically because the application endpoint returns 5xx or times out. Unhealthy targets are removed from rotation, and if too few healthy targets remain to serve peak traffic, the ALB returns 503 Service Unavailable. Since CPU never exceeds 70%, the ASG is not scaling out, confirming the bottleneck is health-check failures rather than capacity.

Exam trap

DOP-C02 often tests the assumption that 503 errors always mean capacity problems, leading candidates to blame Auto Scaling policies when the real cause is failing health checks removing targets from rotation.

How to eliminate wrong answers

Option B is wrong because a security group blocking ALB-to-target traffic would cause all targets to be unhealthy consistently, not intermittently during peak traffic, and would also prevent the health checks from ever succeeding. Option C is wrong because connection draining (deregistration delay) affects in-flight requests during scale-in or deployment, producing client-side errors on specific requests, not ALB-level 503s from unhealthy targets. Option D is wrong because step scaling terminating instances prematurely would show as scale-in events and reduced capacity, but CPU at 70% indicates the policy is not even triggering scale-out, and premature termination would not explain health-check failures.

47
MCQmedium

An application runs on EC2 instances behind an ALB. Users report intermittent 503 errors. The engineer checks ALB metrics and sees 'SurgeQueueLength' increasing periodically. What is the most likely cause?

A.The application instances are not able to process requests quickly enough, causing the request queue to back up.
B.The target group health checks are failing, causing the ALB to route traffic to unhealthy instances.
C.The SSL certificate on the ALB has expired.
D.The ALB security group is blocking traffic from the clients.
AnswerA

When the application instances cannot process requests fast enough, their internal request queue gradually fills up. The ALB forwards new connections to these instances, but if the backend fails to accept or complete the connection within the configured timeout, the load balancer returns HTTP 503 Service Unavailable. This is a performance/capacity bottleneck, not a network or configuration fault.

Why this answer

A high SurgeQueueLength indicates that the ALB is receiving more requests than the target instances can process, causing the request queue to back up. When the queue exceeds the limit (1024 requests), the ALB returns 503 errors. Option B is incorrect because health check failures typically cause the ALB to stop routing traffic to unhealthy instances, reducing the queue rather than increasing it.

Option C is incorrect because an expired SSL certificate would cause TLS handshake failures, not 503 errors. Option D is incorrect because a misconfigured security group would block traffic entirely, resulting in connection timeouts or 503s, but it would not cause the SurgeQueueLength to increase.

48
Drag & Dropmedium

Drag and drop the steps to troubleshoot an AWS CloudTrail that is not logging API calls.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

First verify CloudTrail is enabled, then check bucket policy, then check integrity, then check IAM role, then test.

49
MCQeasy

A company uses Amazon RDS for MySQL and has enabled automated backups. The database administrator accidentally deleted a critical row from a table. The deletion occurred 15 minutes ago. What is the fastest way to recover the lost data?

A.Perform a Point-in-Time Restore to a time just before the deletion.
B.Use the RDS console to undo the last transaction.
C.Use the binary log to replay transactions before the deletion.
D.Create a manual snapshot of the current instance and restore it.
AnswerA

Point-in-Time Restore is the fastest AWS-native recovery method because RDS never leaves you without a timeline: it automatically captures daily snapshots and continuously records MySQL binary logs, so you can pick a timestamp just before the offending DDL/DML statement. RDS then creates a new DB instance by restoring the last automated snapshot before that moment and replaying transaction logs up to the exact second you specify, avoiding any manual log analysis. The restored instance reflects the database before the change, making this both the quickest and most reliable choice.

Why this answer

Amazon RDS automated backups enable Point-in-Time Restore (PITR), which uses continuous transaction log backups to restore the instance to any second within the retention window (up to 35 days). Since the deletion happened 15 minutes ago, performing a PITR to a time just before the deletion is the fastest and most precise recovery method. This restores to a new instance, from which the lost row can be extracted.

Exam trap

DOP-C02 often tests whether candidates know that RDS has no transaction-undo feature and that PITR is the canonical recovery path, while snapshots only capture point-in-time states that may already include the damage.

How to eliminate wrong answers

Option B is wrong because the RDS console provides no 'undo last transaction' feature — relational databases do not expose a rollback button for committed transactions after the fact. Option C is wrong because while MySQL binary logs do record transactions, RDS does not expose them for direct replay by customers, and manually parsing binlogs is far slower and more error-prone than PITR. Option D is wrong because a manual snapshot captures the current state, which already includes the deletion, so restoring it would not recover the lost row and would also lose any changes made after the snapshot.

50
MCQmedium

A company uses AWS Lambda functions to process messages from an Amazon SQS queue. The Lambda function is configured with a reserved concurrency of 5. The SQS queue has a large backlog of messages, and the Lambda function is processing them slowly. The DevOps team wants to increase throughput without making changes to the Lambda code. The team decides to increase the reserved concurrency to 10. However, after the change, the Lambda function starts to experience throttling errors (RateExceeded). The team also notices that other Lambda functions in the same account are also being throttled. What is the MOST likely cause?

A.The SQS queue's polling interval is too high, causing Lambda to poll infrequently.
B.The account's Lambda concurrency limit has been reached due to the increased reserved concurrency.
C.The Lambda function's execution role does not have permission to invoke the function.
D.The SQS queue's visibility timeout is too short, causing messages to be processed multiple times.
AnswerB

When you increase reserved concurrency for a function, you allocate a specific slice of the account's total concurrency limit (e.g., 1000). The remaining functions share the leftover pool, and if the total request rate exceeds that remaining capacity, new invocations for those functions are throttled with a 429 RateExceeded error. This is especially common when a burst of SQS messages triggers many concurrent invocations across multiple functions, saturating the account-wide limit and causing messages to remain in queue.

Why this answer

Lambda concurrency is governed by an account-level limit (default 1,000) shared across all functions in a region. Reserved concurrency carves out a guaranteed slice for a function and subtracts from the unreserved pool available to other functions. Raising the reserved concurrency from 5 to 10 consumes more of the account pool, and if the account is already near its limit, other functions get throttled with RateExceeded errors.

Exam trap

The trap is forgetting that reserved concurrency is deducted from the shared account-level concurrency pool — candidates assume raising reserved concurrency only affects the target function, when it actually reduces capacity available to every other function in the account.

How to eliminate wrong answers

Option A is wrong because SQS polling interval is not a configurable Lambda setting in the way described; Lambda uses event source mapping with batch size and polling behavior, and a slow poll would not cause RateExceeded throttling errors. Option C is wrong because a missing IAM permission would produce AccessDenied errors, not RateExceeded throttling, and the function was already running successfully. Option D is wrong because a short visibility timeout causes duplicate message processing and possibly reprocessing, not RateExceeded throttling errors on the function or other functions in the account.

51
MCQmedium

A company uses AWS CloudTrail to audit API activity. During an incident investigation, they find that a user with the IAM policy 'AdministratorAccess' deleted an S3 bucket. The security team wants to know the source IP address and user agent used for the delete operation. Which action should the team take to obtain this information?

A.View the CloudTrail event history for the delete-bucket event.
B.Check the S3 server access logs for the deleted bucket.
C.Use CloudWatch Logs to search for the event in the CloudTrail log group.
D.Query AWS Config to find the configuration item for the bucket deletion.
AnswerA

Viewing the CloudTrail event history for the delete-bucket event provides the source IP address and user agent because CloudTrail records management API calls, including DeleteBucket. This is the direct and correct method to obtain the required information.

Why this answer

CloudTrail event history captures all management events, including DeleteBucket, and records the source IP address and user agent for each API call. By viewing the event history for the specific delete-bucket event, the security team can directly retrieve the required metadata without needing additional log sources or configurations. Option B is incorrect because S3 server access logs log object-level operations, not management events like bucket deletion.

Option C is not the most direct method; CloudWatch Logs can be used if CloudTrail is configured to send events to a log group, but the simplest way is from CloudTrail event history directly. Option D is incorrect because AWS Config tracks resource configuration changes, not API call details like source IP.

Exam trap

The trap here is that candidates confuse S3 server access logs (which log object-level operations) with CloudTrail management events, leading them to incorrectly choose option B for a bucket deletion that is a management API call.

How to eliminate wrong answers

Option B is wrong because S3 server access logs record object-level requests (e.g., GET, PUT, DELETE on objects), not management-level API calls like DeleteBucket, and they do not capture the user agent or IAM user identity. Option C is wrong because CloudTrail does not automatically deliver events to a CloudWatch Logs log group unless a specific trail is configured with CloudWatch Logs integration; the default event history is not searchable via CloudWatch Logs. Option D is wrong because AWS Config records configuration changes to resources (e.g., bucket existence), but it does not capture the source IP address or user agent of the API call that triggered the change.

52
MCQhard

A company runs a critical application on a fleet of EC2 instances managed by an Auto Scaling group. The application generates logs that are sent to CloudWatch Logs using the CloudWatch agent. Recently, the operations team noticed that some instances are missing logs for certain periods. The CloudWatch agent is configured to batch log events and send them every 5 seconds. The instances have high CPU utilization (90%+) during the missing periods. The DevOps engineer suspects that the agent is being throttled or failing. Which of the following is the MOST likely cause and the BEST course of action?

A.The network bandwidth is saturated, causing log delivery to fail. Increase instance network performance.
B.The CloudWatch Logs retention policy is set to 1 day, so older logs are deleted. Increase retention.
C.The CloudWatch agent is being starved of CPU resources, causing it to drop logs. Increase the CPU credits or instance size.
D.The instances are running out of disk space, preventing log buffering. Add more EBS volume space.
AnswerC

The CloudWatch agent runs as a separate user-space daemon that periodically reads log files and sends them to the CloudWatch Logs API. When the host's CPU is saturated — especially on T-series instances with exhausted CPU credits — the agent's log collection and flush loop can be delayed or preempted for long enough that it begins dropping buffered events to avoid creating an ever-growing backlog. Increasing instance size or CPU credits gives the agent the scheduling time it needs to reliably process and upload log batches, directly resolving the observed missing periods.

Why this answer

When CPU utilization is sustained at 90%+, the CloudWatch agent competes for CPU and can be starved, causing it to drop or delay log batches. The agent buffers events in memory and on disk; under CPU starvation, the buffer may overflow or the agent may fail to flush within the 5-second interval. Increasing CPU credits (for burstable instances) or moving to a larger instance size gives the agent the resources it needs.

Exam trap

DOP-C02 often tests resource contention — candidates blame network or disk because logs are I/O, but the question explicitly states high CPU, pointing to agent starvation.

How to eliminate wrong answers

Option A is wrong because network saturation would affect all traffic, not just logs, and the symptom is CPU-related (90%+ utilization). Option B is wrong because a 1-day retention policy deletes logs after one day, not during the missing periods, and retention does not cause gaps in delivery. Option D is wrong because disk space issues would produce agent errors about buffer overflow, and the question points to CPU as the stressor.

53
Multi-Selecthard

A DevOps team is troubleshooting an application that occasionally throws 'Connection reset by peer' errors when connecting to an RDS MySQL instance. The errors are intermittent and seem to correlate with high traffic. Which TWO steps should the team take to diagnose the issue?

Select 2 answers
A.Check the RDS error log for messages about connection timeouts or aborted connections.
B.Increase the max_connections parameter in the DB parameter group.
C.Enable Multi-AZ deployment to provide a standby instance.
D.Enable Performance Insights to analyze database load and find bottlenecks.
E.Review the security group rules to ensure the application can connect.
AnswersA, D

The RDS error log is the first place to look because it records 'Aborted connection' entries with explicit reasons such as 'Got timeout reading communication packets' or 'Client has exceeded the max_user_connections' when a reset occurs. These entries include the source host, user, and exact error code, which lets you determine whether the reset originates from the database server (e.g., idle timeout) or from the client side. This evidence directly confirms or rules out connection-level failures before you change any configuration or architecture.

Why this answer

Option A is correct because the RDS error log records server-side events such as 'Aborted connection' and connection timeout messages that directly explain why the server is resetting client connections under load. Option D is correct because Performance Insights captures database load (DBLoad) broken down by wait events and top SQL, letting the team pinpoint the bottleneck causing intermittent resets during high traffic. Option B is not a diagnostic step and blindly raising max_connections may not address the root cause.

Option C is a high-availability configuration change, not a troubleshooting action, and Multi-AZ does not resolve connection resets. Option E is unlikely to be the cause since security group misconfiguration would produce consistent connection failures, not intermittent resets correlated with traffic.

Exam trap

DOP-C02 often tests the tendency to jump to configuration changes (like increasing max_connections) instead of first diagnosing via logs and performance tools.

54
MCQeasy

A company uses Amazon RDS for MySQL with Multi-AZ deployment. The application experiences increased latency during peak hours. The DevOps engineer investigates and notices that the Read Replicas are not being utilized effectively. The application is configured to use the primary database endpoint. The engineer wants to offload read traffic to the Read Replicas without changing the application code. What is the BEST solution?

A.Increase the instance size of the primary database to handle the load.
B.Modify the application to use separate endpoints for read and write operations.
C.Create a new Multi-AZ cluster with a read-only endpoint.
D.Configure Amazon RDS Proxy in front of the database and enable read/write splitting.
AnswerD

Amazon RDS Proxy sits between the application and the database, pooling and reusing connections to reduce connection overhead, and when read/write splitting is enabled it can automatically route read queries to one or more Read Replicas while sending write transactions to the primary. This transparently offloads read traffic from the primary without requiring application code changes or manual endpoint management. The proxy also maintains session state and preserves transaction semantics, making it the correct solution for reducing load on a Multi-AZ RDS MySQL instance.

Why this answer

Amazon RDS Proxy with read/write splitting allows the application to offload read traffic to Read Replicas without code changes. It automatically directs read queries to Read Replicas and write queries to the primary, using a single endpoint. Option A (increasing instance size) does not leverage Read Replicas.

Option B (modifying application) requires code changes. Option C (Multi-AZ cluster with read-only endpoint) is invalid because Multi-AZ clusters use a writer and reader endpoint, but the question requires no code changes, and the existing configuration uses a primary endpoint; also, creating a new cluster may not be the simplest solution compared to RDS Proxy.

55
MCQeasy

A company's production environment uses an Amazon ElastiCache Redis cluster for session caching. The operations team reports that the cache hit ratio has dropped significantly, causing increased load on the backend database. What is the MOST likely cause?

A.The cache is under memory pressure and evicting keys to make room.
B.The cluster was resized from a single node to a cluster mode.
C.The encryption in transit was enabled, adding latency.
D.There is a network partition between the application and the cache.
AnswerA

When an ElastiCache cluster reaches its maxmemory limit, the eviction policy (e.g., allkeys-lru or volatile-lru) kicks in to free space by deleting keys. Evicted keys disappear before they can serve valid reads, so subsequent lookups for those keys become cache misses, directly reducing the CacheHitRate metric. This is the most consistent explanation for a sustained drop in hit ratio: the cache is actively discarding data to accommodate new writes, and the eviction is a continuous symptom of memory pressure, not a one-time event.

Why this answer

A drop in cache hit ratio with increased backend load is most commonly caused by memory pressure leading to eviction of keys. When ElastiCache Redis reaches maxmemory, it evicts keys according to the eviction policy (e.g., LRU), causing more cache misses and higher database load.

Exam trap

DOP-C02 often tests the difference between symptoms of memory pressure (evictions, low hit ratio) and other issues like network partitions or encryption, tricking candidates into selecting configuration changes that do not address the root cause.

How to eliminate wrong answers

Option B is wrong because resizing to cluster mode generally increases capacity and does not inherently reduce hit ratio; it might require reconfiguration but not eviction. Option C is wrong because enabling encryption in transit adds minor latency but does not cause cache misses or reduce hit ratio. Option D is wrong because a network partition would cause complete connectivity loss, not a gradual drop in hit ratio; it would result in errors, not just misses.

56
MCQmedium

A company uses AWS Organizations with multiple accounts. The security team wants to ensure that all accounts automatically forward their CloudWatch Logs to a central logging account. Which solution should the team implement?

A.Enable AWS Config aggregator in the central account
B.Use AWS Service Catalog to create a product for log forwarding
C.Use AWS CloudFormation StackSets to deploy a subscription filter and Lambda function in each account
D.Configure AWS Organizations to automatically forward logs
AnswerC

AWS CloudFormation StackSets extends CloudFormation to deploy stacks across multiple accounts and regions within AWS Organizations, automatically and in a single operation. You can define a template containing a CloudWatch Logs subscription filter and a Lambda function, and StackSets will create those resources in every member account. The subscription filter streams selected log events from each account's log groups to the Lambda function, which then processes and forwards them to a centralized destination such as an S3 bucket, Kinesis, or a central CloudWatch account. This approach is fully automated, consistent, and the recommended pattern for centralized log forwarding.

Why this answer

AWS CloudFormation StackSets allows you to deploy infrastructure components across multiple accounts and regions in an AWS Organization. In this case, you can create a StackSet that includes a CloudWatch Logs subscription filter and a Lambda function to forward logs from each account to a central logging account. The subscription filter triggers the Lambda function to forward logs to a destination in the central account.

This ensures all accounts automatically forward their CloudWatch Logs. Option A is incorrect because AWS Config aggregator aggregates configuration and compliance data, not logs. Option B is incorrect because AWS Service Catalog is used to create and manage IT service catalogs for approved products, not for log forwarding.

Option D is incorrect because AWS Organizations does not have a native capability to forward logs; it manages policies and account structure.

57
MCQeasy

A DevOps engineer is investigating an incident where an EC2 instance became unreachable. The engineer checks the AWS Management Console and finds the instance is running, but the status check shows '2/2 checks passed' and the system log shows no errors. What should the engineer do NEXT to diagnose the connectivity issue?

A.Review the CloudWatch metrics for CPU utilization and network throughput.
B.Reboot the instance to reset the network interface.
C.Stop and start the instance to move it to new underlying hardware.
D.Check the security group and network ACL rules to ensure inbound traffic is allowed.
AnswerD

Security groups act as a stateful firewall at the instance level, while network ACLs are stateless at the subnet level, so both must allow the relevant inbound traffic and the NACL must also allow the corresponding outbound return traffic. An incorrect deny rule, an overly restrictive CIDR, or a missing allow for the source IP/port pairs will cause exactly this kind of unreachability even when the instance is running and healthy, making this the correct first troubleshooting step.

Why this answer

Since the instance is running, status checks pass, and the system log shows no errors, the issue is not with the operating system or underlying hardware. The most likely cause is a network-layer restriction, such as security group or network ACL rules blocking inbound traffic. Checking these rules is the correct next step because they control traffic at the instance and subnet levels, respectively, and misconfigurations here are a common cause of unreachability despite healthy instance status.

Exam trap

The trap here is that candidates assume a 'running' instance with passing status checks guarantees network reachability, overlooking that security groups and NACLs can silently drop traffic without any error in system logs or status checks.

How to eliminate wrong answers

Option A is wrong because CloudWatch metrics for CPU utilization and network throughput measure performance, not connectivity; they would not reveal whether inbound traffic is being blocked by security groups or NACLs. Option B is wrong because rebooting the instance resets the OS but does not change network configurations or underlying hardware; if the instance is running and status checks pass, a reboot is unlikely to resolve a network-level block. Option C is wrong because stopping and starting the instance moves it to new underlying hardware, which could help if the issue were hardware-related, but the status checks passing indicates the hardware is healthy; this action is more disruptive and unnecessary for a likely network configuration problem.

58
MCQmedium

A DevOps engineer is troubleshooting an issue where an EC2 instance running a web application becomes unresponsive every few hours. CloudWatch logs show no application errors, but the instance's status checks are passing. The engineer suspects a memory leak. Which AWS service can be used to capture memory utilization metrics at a granular level to confirm the leak?

A.EC2 Status Checks
B.AWS Config
C.AWS CloudTrail
D.CloudWatch Agent
AnswerD

The unified CloudWatch Agent is the correct solution because it runs as a service on your EC2 instance and reads guest-OS metrics directly from /proc and the operating system, allowing it to report memory utilization as a custom metric to CloudWatch. You can install and configure it via the command line, SSM Run Command, or an AWS-supplied recipe, and attach an IAM role with `cloudwatch:PutMetricData` permissions. It also supports collecting disk usage, CPU per-core statistics, and logs, all of which are invisible to hypervisor-level default EC2 metrics.

Why this answer

The CloudWatch Agent can be installed on EC2 instances to collect custom metrics such as memory usage and send them to CloudWatch, allowing granular monitoring. Option A is incorrect because EC2 Status Checks only verify system status (e.g., network reachability). Option B is incorrect because AWS Config tracks configuration changes, not performance metrics.

Option C is incorrect because CloudTrail records API activity, not system metrics.

59
MCQeasy

A company uses CloudWatch Synthetics canaries to monitor a critical API endpoint. Recently, a canary started failing with a '403 Forbidden' error. The DevOps engineer verifies that the canary's IAM role has the necessary permissions to invoke the API and that the API endpoint is publicly accessible. What should the engineer check NEXT?

A.Review the canary's CloudWatch Logs for any runtime errors.
B.Increase the canary's memory to 512 MB to prevent timeout-related issues.
C.Check if the API requires an API key or other authentication that the canary is not providing.
D.Verify that the canary is attached to the correct VPC and subnet.
AnswerC

A 403 Forbidden response from an API means the server understood the request but refused to authorize it; public APIs commonly require an API key, an Authorization header, or a signed payload, and the canary's HTTP request may be missing those credentials. Verify the canary script attaches the API key (e.g., x-api-key header) or uses the same authentication mechanism as your test client; also ensure the key is valid and not expired.

Why this answer

A 403 Forbidden error indicates that the request reached the API but was denied due to authentication or authorization issues. Since the canary's IAM role has the necessary permissions and the endpoint is publicly accessible, the most likely cause is that the API requires an API key or other authentication mechanism (e.g., a token) that the canary is not providing. Checking for missing authentication is the logical next step.

Exam trap

DOP-C02 often tests the distinction between IAM permissions and application-level authentication; candidates may assume that IAM roles are sufficient for API access, ignoring API keys or other auth mechanisms.

How to eliminate wrong answers

Option A is wrong because runtime errors would typically produce different error codes (e.g., 500 or timeout) and the 403 specifically points to an authentication/authorization issue, not a code error. Option B is wrong because increasing memory addresses timeout issues, which would manifest as timeouts, not 403 errors. Option D is wrong because VPC/subnet misconfiguration would result in network connectivity failures (e.g., timeouts or connection refused), not a 403 response from the API.

60
Multi-Selectmedium

A DevOps team is investigating a production incident where an Amazon RDS for MySQL database experienced a sudden spike in connections and CPU utilization. The team suspects a SQL injection attack. Which TWO actions should the team take to investigate and mitigate the incident?

Select 2 answers
A.Delete the error logs to free up storage space.
B.Enable automated backups and ensure point-in-time recovery is configured.
C.Enable RDS Enhanced Monitoring and audit logs to capture SQL queries.
D.Create a read replica to offload traffic from the primary instance.
E.Increase the DB instance size to handle the increased load.
AnswersB, C

Enabling automated backups with point-in-time recovery (PITR) is the correct immediate response because it lets you restore the database to any second within the retention window, such as a timestamp just before the compromise occurred. Automated backups take daily snapshots, and RDS continuously records transaction logs to enable PITR, minimizing data loss if tables were dropped or data was tampered with. This is the foundational recovery mechanism to return the system to a known-good state.

Why this answer

Enabling automated backups and point-in-time recovery ensures that the database can be restored to a state before the suspected SQL injection attack, preserving data integrity and enabling forensic analysis. Option C is correct because RDS Enhanced Monitoring provides OS-level metrics (CPU, memory, disk I/O) to correlate with the spike, while audit logs capture actual SQL queries, which are essential for identifying malicious patterns and confirming the attack vector.

Exam trap

The trap here is that candidates confuse reactive scaling (Option E) or read replicas (Option D) with proper incident response, failing to recognize that investigation and mitigation require enabling logging and backup capabilities, not just increasing capacity.

61
Multi-Selecteasy

A company runs a critical application on EC2 instances in an Auto Scaling group. The application must be highly available across multiple Availability Zones. Which TWO configurations are necessary to achieve this? (Choose TWO.)

Select 2 answers
A.Use a single Availability Zone to reduce latency.
B.Use Spot Instances to reduce costs.
C.Use a Classic Load Balancer to distribute traffic.
D.Place an Application Load Balancer in front of the Auto Scaling group.
E.Configure the Auto Scaling group to launch instances in multiple Availability Zones.
AnswersD, E

An Application Load Balancer operates at Layer 7 and can distribute HTTP/HTTPS traffic across instances in multiple Availability Zones, providing health checks and automatic registration of instances in the Auto Scaling group. This integration allows the ALB to route traffic only to healthy instances and to scale with the ASG's changes, enhancing availability and fault tolerance. It is the recommended pattern for web applications requiring high availability and advanced routing.

Why this answer

To achieve high availability across multiple Availability Zones (AZs), two key configurations are required. First, the Auto Scaling group must be configured to launch instances in multiple AZs (Option E). This ensures that if one AZ fails, the application can continue serving traffic from instances in other AZs.

Second, an Application Load Balancer (ALB) should be placed in front of the Auto Scaling group (Option D). The ALB distributes incoming traffic across instances in all AZs where the Auto Scaling group has launched instances. It also performs health checks and automatically routes traffic away from unhealthy instances.

Option A is incorrect because using a single AZ would create a single point of failure, undermining high availability. Option B is incorrect because Spot Instances can be terminated with little notice, which is unsuitable for a critical application requiring consistent availability. Option C is incorrect because a Classic Load Balancer lacks advanced features like path-based routing and native cross-zone load balancing, making the ALB a better choice for modern applications.

62
MCQeasy

A DevOps engineer is troubleshooting an Auto Scaling group (ASG) that is not launching instances as expected. The ASG is configured with a launch template that uses an Amazon Linux 2 AMI. The engineer checks the EC2 Auto Scaling console and sees that the group's desired capacity is set to 2, but only 1 instance is running. The last scaling activity shows 'Failed to launch instance. Error: Your quota allows for 0 more running instance(s).' What is the most likely cause?

A.The launch template has insufficient IAM permissions to create instances.
B.The account has reached the EC2 instance limit for the selected instance type in the region.
C.The VPC subnet does not have enough available IP addresses.
D.There is an instance that is not passing health checks, preventing new instances.
AnswerB

The account has reached the EC2 instance limit for the selected instance type in the region, which is the only option that matches the error 'You have requested more instances than your current EC2 instance limit allows'. Auto Scaling attempts to call RunInstances, but EC2 rejects it with an InstanceLimitExceeded error, causing the scaling activity to fail. This is a service quota that applies per instance type per region, and the fix is to request a quota increase via the Service Quotas console.

Why this answer

The error message 'Your quota allows for 0 more running instance(s)' directly indicates that the AWS account has reached its EC2 instance limit for the specific instance type in the region. Auto Scaling groups cannot launch instances beyond the service quota, regardless of the desired capacity. This is a common issue when the default limit (e.g., 5 or 20 instances per instance family) has been exhausted.

Exam trap

The trap here is that candidates confuse IAM permissions with service quotas, or assume a VPC subnet IP shortage is the cause, when the explicit quota error message is the definitive clue.

How to eliminate wrong answers

Option A is wrong because IAM permissions affect the ability to call EC2 APIs (e.g., RunInstances), but the error message explicitly cites a quota issue, not an authorization failure (which would return 'UnauthorizedOperation' or 'AccessDenied'). Option C is wrong because insufficient IP addresses in the subnet would produce an error like 'InsufficientFreeAddressesInSubnet' or 'NoFreeAddressesInSubnet', not a quota limit message. Option D is wrong because an instance failing health checks does not prevent new instances from launching; the ASG would still attempt to launch replacements, and the error would relate to health check failures, not a quota limit.

63
MCQmedium

A company uses Amazon CloudFront to distribute content globally. Users in some regions report slow load times. The DevOps team wants to identify the geographic regions where performance is worst. Which tool should they use?

A.Amazon CloudWatch Metrics for CloudFront
B.CloudFront access logs in S3
C.Amazon Route 53 latency records
D.CloudFront reports in the AWS Management Console
AnswerD

CloudFront reports in the AWS Management Console include a Geo Distribution report and viewer reports that break down requests, bytes served, and other metrics by country and by edge location. These built-in reports are pre-aggregated by AWS and displayed directly in the console, giving a quick geographic view of traffic patterns and performance without needing additional infrastructure or custom queries. They are the intended way to answer global distribution questions out of the box.

Why this answer

CloudFront reports in the AWS Management Console provide performance metrics (e.g., total requests, error rates, latency) broken down by geographic region, enabling the team to identify regions with worst performance. Option A is wrong: CloudWatch metrics for CloudFront are aggregated per distribution, not per region. Option B is wrong: CloudFront access logs in S3 record individual requests but do not aggregate performance by region.

Option C is wrong: Amazon Route 53 latency records are used for DNS-based routing decisions, not for analyzing CloudFront performance.

64
Multi-Selectmedium

A company uses AWS CloudTrail to log API calls in a multi-account environment. The security team wants to be alerted immediately when an IAM user or role performs a specific sensitive action (e.g., DeleteTrail, DeleteDBInstance). Which TWO services can be used together to achieve near real-time alerting? (Choose TWO.)

Select 2 answers
A.CloudWatch Logs metric filters and alarms
B.CloudTrail with CloudWatch Logs integration
C.CloudTrail with Amazon S3 event notifications
D.Amazon Athena and CloudWatch dashboards
E.AWS Config and AWS Lambda
AnswersA, B

CloudWatch Logs metric filters are the correct mechanism for real-time alerting on CloudTrail API activity. A metric filter defines a pattern that matches specific CloudTrail log events, such as `errorCode = "UnauthorizedOperation"`, and continuously increments a custom CloudWatch metric. A CloudWatch alarm can then evaluate that metric over a fixed period (e.g., 5 minutes) and trigger an SNS notification when the threshold is breached, enabling fast, automated responses to suspicious API calls.

Why this answer

CloudTrail logs can be delivered to CloudWatch Logs, where metric filters can be created to match specific API actions (e.g., DeleteTrail, DeleteDBInstance) and trigger CloudWatch alarms for near real-time notification. Option B is correct because CloudTrail integration with CloudWatch Logs is the prerequisite step that enables the log delivery required for metric filters and alarms. Option C is incorrect: CloudTrail with Amazon S3 event notifications is not near real-time; S3 event notifications can have delays and are not designed for immediate alerting on specific API calls.

Option D is incorrect: Amazon Athena is an interactive query service for ad-hoc analysis, not for real-time alerting. Option E is incorrect: AWS Config is used for resource configuration compliance and change tracking, not for real-time alerting on API calls.

65
MCQmedium

A DevOps engineer manages a DynamoDB table that serves a high-traffic ordering API. During a flash sale, ProvisionedThroughputExceededException errors spike and the on-call engineer must reduce customer impact quickly. The table currently uses provisioned capacity and the workload is expected to surge unpredictably for the next several hours. Which action should the engineer take FIRST to mitigate the incident?

A.Enable DynamoDB Accelerator (DAX) in front of the table to absorb the surge.
B.Enable DynamoDB on-demand capacity mode for the table to absorb the unpredictable traffic without throughput errors.
C.Increase the table's provisioned read and write capacity units manually to a value matching the observed peak.
D.Create a global secondary index with higher capacity and redirect all API reads to it.
AnswerB

Switching the table to on-demand capacity mode removes the provisioned throughput ceiling and instantly accommodates unpredictable spikes, which is exactly the flash-sale pattern here. It is the fastest mitigation because the table continues serving reads and writes during the switch, so the ProvisionedThroughputExceededException errors stop without capacity math or client changes.

Why this answer

The scenario describes unpredictable burst traffic against a provisioned DynamoDB table producing throttling errors. Switching to on-demand capacity mode immediately removes the provisioned throughput limit and lets the table scale with the surge while still serving requests, making it the fastest and most reliable mitigation. Approaches that tune provisioned values or cache reads do not remove the write bottleneck during an unpredictable spike.

Exam trap

The trap here is assuming that caching with DAX or adding an index solves throughput throttling, when throttling in a write-heavy burst is caused by the base table's provisioned capacity ceiling.

66
MCQhard

A company runs a three-tier application on Amazon EC2 behind an Application Load Balancer. During an incident, users report intermittent HTTP 502 errors, and the engineer finds that some targets are failing health checks and being removed and re-added repeatedly. The application writes large log files to the instance store and the engineer suspects the health check configuration. Application startup takes about 90 seconds. Which configuration change should the engineer make to resolve the flapping targets?

A.Enable cross-zone load balancing on the Application Load Balancer to spread traffic more evenly.
B.Set the health check interval to 5 seconds and the unhealthy threshold to 1 so failures are detected faster.
C.Increase the health check interval and unhealthy threshold, and use a longer health check grace period or slow start so targets are not removed during startup.
D.Change the target group to use TCP health checks instead of HTTP health checks.
AnswerC

Flapping occurs when targets are marked unhealthy during normal startup or heavy I/O. Lengthening the interval and unhealthy threshold tolerates transient failures, while a health check grace period or slow start keeps new or busy targets in service until they are ready, which directly stops the repeated removal and re-addition causing 502s.

Why this answer

The targets flap because health checks fail during startup or heavy disk I/O, so the load balancer repeatedly removes and re-adds them, producing intermittent 502 responses. Extending the health check interval and unhealthy threshold and applying a grace period or slow start keeps targets in service until they are genuinely ready, which stops the flapping and the resulting errors.

Exam trap

The trap here is treating flapping targets as a load-balancing distribution problem and enabling cross-zone balancing, when the real cause is health check timing relative to application startup.

67
MCQeasy

A company stores application logs in Amazon CloudWatch Logs. During an incident, an engineer needs to search across multiple log groups for a specific request ID from the last hour and then preserve the findings for a post-incident review. Which approach meets both needs with the least operational effort?

A.Create a CloudWatch Logs subscription filter that streams matching events to AWS Lambda for processing.
B.Enable CloudTrail logging for the log groups and search the trail for the request ID.
C.Use CloudWatch Logs Insights to query the log groups for the request ID and save the query results for the post-incident review.
D.Export all log groups to Amazon S3 and use Amazon Athena to search for the request ID.
AnswerC

CloudWatch Logs Insights queries multiple log groups at once with a time range and supports saving or exporting results, directly matching the search and preservation requirements. It requires no infrastructure, so it is the lowest-effort option for finding the request ID and retaining the findings for review.

Why this answer

CloudWatch Logs Insights is designed to run interactive queries across one or more log groups over a specified time range, which fits searching the last hour of logs for a request ID. Because it also lets you save or export query results, it satisfies the need to preserve findings for a post-incident review with minimal setup and no additional infrastructure.

Exam trap

The trap here is reaching for S3 export and Athena for a quick, time-bounded log search, when CloudWatch Logs Insights already queries multiple log groups directly and can retain the results.

68
MCQeasy

A company hosts a static website on Amazon S3 with CloudFront as the CDN. Users report that they see an old version of the website even after the DevOps team updated the S3 objects. The team verified that the new objects are in the S3 bucket and are publicly accessible. The CloudFront distribution has a default TTL of 24 hours. To immediately serve the new content to users, the team needs to invalidate the CloudFront cache. Which of the following is the CORRECT approach to achieve this with minimal impact?

A.Create a CloudFront invalidation request for the path '/*'.
B.Change the CloudFront origin path to point to a new S3 bucket.
C.Update the CloudFront distribution's default TTL to 0 and wait for the changes to propagate.
D.Delete the S3 objects and re-upload them with different names.
AnswerA

An invalidation request for the '/*' path removes all objects from CloudFront's edge caches across every region, which forces the distribution to return to the S3 origin on the next request and fetch the updated content. This is the standard, immediate method for clearing cached content when you need to publish new website changes, and it does not require changing URLs or reconfiguring any origin settings.

Why this answer

A CloudFront invalidation for the path '/*' tells all edge locations to stop serving cached objects matching that pattern and fetch fresh copies from the S3 origin on the next request. This is the standard, immediate way to purge stale content without changing the distribution configuration or object keys. It has minimal impact because it only affects cached objects and does not require re-uploading or renaming anything.

Exam trap

The trap is confusing TTL changes with cache invalidation — candidates pick 'set TTL to 0' thinking it purges existing cached objects, but TTL changes only affect future caching, not objects already stored at edge locations.

How to eliminate wrong answers

Option B is wrong because changing the origin path to a new S3 bucket would require the new content to actually exist in that bucket and would break existing URLs; it is a disruptive configuration change, not a cache purge. Option C is wrong because setting the default TTL to 0 only affects future cache behavior — objects already cached at edge locations will still be served until their existing TTL expires, so it does not immediately serve new content. Option D is wrong because deleting and re-uploading objects with different names changes the object keys, breaking existing links and requiring HTML updates; it also does not invalidate the old cached keys.

69
MCQmedium

An application log excerpt shows repeated HTTP 500 errors for the /api/orders endpoint, with occasional successful health checks. The application runs on EC2 instances behind an ALB. What is the MOST likely cause of this pattern?

A.The backend service that the /api/orders endpoint depends on is unavailable or failing.
B.The EC2 instances are running out of memory and the application is crashing.
C.The ALB is misconfigured and routing requests to the wrong target group.
D.The EC2 instances are not passing health checks and are being deregistered from the target group.
AnswerA

The HTTP 500 on /api/orders is generated by the application after the ALB successfully forwards the request, meaning the instance is reachable and the web server process is handling traffic. However, the /api/orders endpoint depends on a downstream service (e.g., database, internal microservice, cache) whose failure causes the application to throw an unhandled exception and return a 500. Health checks succeed because the configured health check path (typically /health or /) exercises simple connectivity and doesn't call the failing dependency, so the instance remains in service. This is a classic partial-failure scenario where the dependency, not the instance, is unhealthy.

Why this answer

The pattern of repeated HTTP 500 errors for /api/orders with occasional successful health checks strongly indicates that the backend service dependency (e.g., a database, cache, or another microservice) is intermittently failing or unavailable. HTTP 500 errors are server-side errors, meaning the application code is running but cannot complete the request due to a downstream failure. Successful health checks confirm the EC2 instances themselves are healthy and in-service, ruling out instance-level or ALB misconfiguration issues.

Exam trap

The trap here is that candidates confuse HTTP 500 errors with instance-level failures (like OOM or health check failures), but the key differentiator is that successful health checks prove the instances are operational, shifting the root cause to a failing backend dependency rather than the compute layer.

How to eliminate wrong answers

Option B is wrong because running out of memory typically causes the application process to crash (e.g., OOM killer), leading to connection timeouts or immediate 503 errors, not repeated HTTP 500 errors with successful health checks. Option C is wrong because a misconfigured ALB routing to the wrong target group would cause requests to reach instances that don't serve the /api/orders endpoint, resulting in 404 or 503 errors, not 500 errors from the application itself. Option D is wrong because if instances were failing health checks and being deregistered, they would be removed from the target group and stop receiving traffic entirely, which contradicts the observed pattern of occasional successful health checks and persistent 500 errors on /api/orders.

70
MCQhard

A company runs a critical application on an Amazon RDS for MySQL DB instance. The application experiences intermittent connection timeouts. The DevOps team notices that the DB instance's CPU and memory metrics are normal. What should the team check NEXT to diagnose the issue?

A.Enable Enhanced Monitoring to check OS-level metrics
B.Examine the slow query log to identify long-running queries
C.Verify that the DB instance's storage is not full
D.Check the 'DatabaseConnections' CloudWatch metric to see if the connection count is near the max_connections limit
AnswerD

The DatabaseConnections CloudWatch metric directly tracks the number of client connections currently established to the RDS instance, and when this value approaches the max_connections parameter, the server begins rejecting or timing out new connection attempts. This pattern perfectly matches the symptom of intermittent connection timeouts while CPU and memory remain normal, because the connection ceiling is a hard limit unaffected by resource utilization. Comparing this metric against the configured max_connections (from the parameter group) confirms whether connection exhaustion is the cause.

Why this answer

The most likely cause of intermittent connection timeouts when CPU and memory are normal is that the number of database connections has reached the max_connections limit. The 'DatabaseConnections' CloudWatch metric shows the current number of connections; if it's near the limit, new connections will be refused or time out. Checking this metric is the next logical step to diagnose the issue.

Exam trap

DOP-C02 often tests the misconception that connection timeouts are always due to slow queries or resource exhaustion, but they can also be caused by connection limits, which are not reflected in CPU/memory metrics.

How to eliminate wrong answers

Option A is wrong because Enhanced Monitoring provides OS-level metrics, but CPU and memory are already normal, so OS-level metrics are less likely to reveal the cause. Option B is wrong because slow queries would typically increase CPU or memory usage, and they cause slow responses, not necessarily connection timeouts. Option C is wrong because a full storage would cause write failures, not intermittent connection timeouts, and it would likely be accompanied by other errors.

71
MCQmedium

A company's application running on EC2 instances behind an Application Load Balancer (ALB) is returning intermittent 504 errors. The instances are in an Auto Scaling group with a health check grace period of 300 seconds. What should the DevOps engineer check first to troubleshoot the issue?

A.Review the Auto Scaling group scaling policies.
B.Verify the target group health checks are passing.
C.Check security group rules for the ALB.
D.Check ALB access logs for target response times.
AnswerD

ALB access logs are the authoritative source for diagnosing 504 timeouts because they record the HTTP status code and three timing dimensions: request_processing_time, target_processing_time, and response_processing_time. The target_processing_time field directly measures how long the ALB waited for the target to begin sending a response; values at or near the idle timeout threshold precisely identify the request and target instance responsible for the 504. This option provides concrete, per-request evidence of backend latency, unlike health checks or security group rules, which only indicate availability and reachability.

Why this answer

A 504 error indicates the load balancer did not receive a response from the target within the idle timeout period. Checking ALB access logs for target response times is the first step to determine if the backend is slow or unresponsive. Option A is wrong because scaling policies affect the number of instances, not response times.

Option B is wrong because health checks verify instance availability, but intermittent slow responses may not cause health check failures. Option C is wrong because security group rules would cause different errors (e.g., connection timeouts) rather than 504s.

72
MCQeasy

A DevOps engineer is troubleshooting a Lambda function that processes S3 events. The function has been running successfully for months, but today it started timing out. The engineer checks CloudWatch Logs and sees 'Task timed out after 3.01 seconds' errors. The function is configured with a 3-second timeout. What should the engineer do to resolve the issue?

A.Increase the Lambda function reserved concurrency.
B.Increase the memory allocation for the Lambda function.
C.Increase the Lambda function timeout to 10 seconds.
D.Configure a dead-letter queue (DLQ) for the Lambda function.
AnswerC

The Lambda function is failing because its actual execution time is longer than the configured timeout. Raising the timeout to 10 seconds directly expands the allowed execution window, giving the code enough time to complete its work — as long as it stays below the 15-minute maximum. This is the only option that addresses the root cause of a timeout error rather than changing resources or failure handling.

Why this answer

The error 'Task timed out after 3.01 seconds' indicates the function is exceeding its configured 3-second timeout. Since the function has been running successfully for months and only recently started timing out, the workload has likely grown (larger S3 objects, downstream latency, or cold starts). Increasing the timeout to 10 seconds gives the function enough time to complete its processing, directly addressing the timeout error.

Exam trap

The trap here is confusing timeout with concurrency or memory: candidates often assume more memory or concurrency will fix a timeout, but the timeout is a separate, explicitly configured limit that must be raised.

How to eliminate wrong answers

Option A is wrong because reserved concurrency controls how many concurrent invocations can run, not how long a single invocation may run; it would not prevent a timeout. Option B is wrong because increasing memory can improve CPU and slightly speed execution, but it does not extend the timeout window and is not the direct fix for a timeout error. Option D is wrong because a DLQ captures failed asynchronous invocations after they fail; it does not prevent the timeout and would only record the failure.

73
MCQmedium

A company uses EC2 instances in an Auto Scaling group behind an ALB. The DevOps team receives alerts that the CPU utilization on the instances is consistently above 90% during peak hours. The Auto Scaling group is configured with a simple scaling policy that adds one instance when CPU exceeds 80% and removes one when below 30%. However, during sudden traffic spikes, the scaling policy reacts too slowly, causing performance degradation. The team wants to improve the scaling responsiveness without over-provisioning. What should the team do?

A.Increase the cooldown period for the simple scaling policy to allow more time for metrics to stabilize.
B.Replace the simple scaling policy with a step scaling policy that adds multiple instances when CPU exceeds 80%.
C.Create a scheduled scaling action to add instances before peak hours based on historical data.
D.Replace the simple scaling policy with a target tracking scaling policy based on average CPU utilization with a target value of 70%.
AnswerD

A target tracking scaling policy with a target value of 70% average CPU utilization proactively adds instances before utilization reaches 80% and continuously adjusts to maintain the target, providing fast response to spikes without over-provisioning.

Why this answer

A target tracking scaling policy automatically adjusts the size of the Auto Scaling group to keep the average CPU utilization close to the target value (70%). This provides a proactive and responsive scaling mechanism for sudden traffic spikes without manual intervention. Option A is incorrect because increasing the cooldown period would delay scaling actions, worsening the response time.

Option B is incorrect because although a step scaling policy can add multiple instances at once, it requires manual configuration of thresholds and step adjustments, and it may not adapt as smoothly to varying spikes as target tracking. Option C is incorrect because scheduled scaling only addresses predictable traffic patterns, not sudden, unpredictable spikes.

74
MCQmedium

A company uses AWS CloudTrail to log all API calls. During an incident investigation, a security engineer needs to identify who deleted an S3 bucket named 'critical-data' two days ago. Which approach will provide the necessary information?

A.Use AWS CloudTrail LookupEvents API to search for DeleteBucket events.
B.Check the AWS Management Console activity history.
C.Review the S3 access logs for the bucket.
D.Search CloudWatch Logs for 'DeleteBucket' events.
AnswerA

The AWS CloudTrail LookupEvents API is the correct method because it directly queries CloudTrail's event history, which records all control plane API calls such as DeleteBucket. You can use the LookupAttributes parameter with AttributeKey=EventName and AttributeValue=DeleteBucket, filtering by time range if needed. This returns the full event record, including the IAM user or role, source IP, and timestamp, for any deletion within the last 90 days. It is the native, low-latency mechanism for exactly this forensic search.

Why this answer

CloudTrail records management events such as DeleteBucket, and the LookupEvents API allows programmatic search of the last 90 days of event history by event name, so it can identify who deleted the bucket two days ago. This is the intended mechanism for API-level auditing. The other options either lack the event detail, cover only data-plane access, or require prior configuration that is not guaranteed.

Exam trap

DOP-C02 often tests the boundary between CloudTrail management events and S3 access logs, so candidates pick S3 access logs because the resource is an S3 bucket, missing that bucket deletion is a management event captured by CloudTrail.

How to eliminate wrong answers

Option B is wrong because the AWS Management Console activity history only shows actions taken by the current user in the console and does not provide a full account-wide audit trail with identity details. Option C is wrong because S3 access logs record data-plane requests to objects and buckets, not the DeleteBucket management API call or the caller identity in the same way CloudTrail does. Option D is wrong because CloudWatch Logs only contains CloudTrail events if a trail was configured to deliver them there, and even then the LookupEvents API is the direct, reliable way to query event history.

75
MCQeasy

A DevOps engineer receives a CloudWatch alarm that the 'StatusCheckFailed' metric for an EC2 instance is in ALARM state. The instance is part of an Auto Scaling group. What should the engineer do first to restore service?

A.Update the Auto Scaling group's launch configuration
B.Wait for Auto Scaling to replace the instance
C.Manually terminate the instance
D.Reboot the instance
AnswerB

Auto Scaling is configured to use EC2 status checks for health, so it will detect the instance's failed status check and automatically terminate it, launching a new instance with the same launch configuration to maintain desired capacity. Waiting is the correct action because the replacement process is fully automated and requires no manual intervention. The new instance will be created once the unhealthy instance is marked as unhealthy, typically within a few minutes.

Why this answer

The Auto Scaling group automatically replaces unhealthy instances based on EC2 status checks. Options A, C, and D are incorrect: Updating the launch configuration does not fix existing instances; manually terminating is unnecessary as Auto Scaling handles it; and rebooting may not resolve underlying issues.

Page 1 of 3 · 183 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Incident and Event Response questions.