Courseiva

AWS Certified DevOps Engineer Professional DOP-C02 (DOP-C02) — Questions 226–300

1298 questions total · 18pages · All types, answers revealed

Page 3

Page 4 of 18

Page 5
226
MCQeasy

A startup is using AWS CodeBuild to build and test their application. The build process takes about 10 minutes. Recently, they noticed that some builds are failing randomly with the error 'Could not download dependencies'. The build environment uses a custom Docker image stored in Amazon ECR. The team suspects that the issue is due to network connectivity problems when pulling the Docker image or dependencies from the internet. They want to ensure reliable and faster builds. Which solution should they implement?

A.Switch to using a public Docker image from Docker Hub
B.Increase the build timeout in CodeBuild project settings
C.Use a larger compute type for the CodeBuild project
D.Configure CodeBuild to use a VPC with a NAT gateway
AnswerD

Configuring CodeBuild to use a VPC with a NAT gateway is the correct solution because it provides a deterministic, controlled egress path for outbound internet traffic. By running the build in private subnets behind a NAT gateway that routes via an Internet Gateway, the build environment can reliably pull layers from Docker Hub, install packages, and access private VPC resources. This overrides the default AWS-managed network's unpredictable connectivity and gives you proper security group and routing control, making the build network behavior explicit and dependable.

Why this answer

To improve reliability and speed, configure CodeBuild to use a VPC with a NAT gateway. This provides consistent internet access for pulling dependencies and Docker images, and allows using VPC endpoints for Amazon ECR, reducing network failures. Option D is correct.

Option A (using a public Docker Hub) does not address the underlying network issues and may introduce additional points of failure. Option B (increasing build timeout) does not fix the root cause of connectivity problems. Option C (using a larger compute type) does not resolve network connectivity issues.

227
MCQmedium

A DevOps engineer is designing a CI/CD pipeline for a serverless application using AWS Lambda. They want to automatically deploy the latest version of the Lambda function to production after running integration tests. The source code is in AWS CodeCommit. Which pipeline configuration should they use?

A.CodeCommit -> CodeBuild (test) -> CodeDeploy (Lambda deployment) -> Lambda.
B.CodeCommit -> CodeBuild (test) -> Lambda (deploy via update-function-code).
C.CodeCommit -> Lambda (deploy via S3 trigger) -> CodeBuild (test) -> production.
D.CodeCommit -> CodeBuild (test and deploy) -> Lambda via AWS CLI in buildspec.
AnswerA

This is correct because CodeDeploy natively supports Lambda deployment with canary, linear, and all-at-once traffic-shifting strategies, letting you validate a new version before promoting it. CodeBuild runs automated tests first, then CodeDeploy updates the Lambda alias gradually, monitoring CloudWatch alarms for automatic rollback. This is the AWS-recommended managed deployment path for serverless applications.

Why this answer

It uses CodeDeploy's built-in Lambda deployment support, which enables safe, gradual traffic shifting (e.g., canary or linear deployments) and automatic rollback on CloudWatch alarm failures. This pipeline integrates CodeCommit for source, CodeBuild for integration tests, and CodeDeploy to orchestrate the Lambda update with minimal risk, aligning with AWS best practices for serverless CI/CD.

Exam trap

The trap here is that candidates assume any pipeline that runs tests before deploying to Lambda is sufficient, but the exam specifically tests the need for managed deployment strategies (CodeDeploy) over direct API calls or CLI commands to ensure production safety and rollback capabilities.

How to eliminate wrong answers

Option B is wrong because calling Lambda's update-function-code API directly from CodeBuild bypasses deployment safety features like traffic shifting, rollback, and pre/post-traffic hooks, which are critical for production deployments. Option C is wrong because triggering a Lambda deployment via an S3 trigger before running tests violates the CI/CD principle of testing before deployment, and S3 triggers are asynchronous and lack deployment orchestration. Option D is wrong because using the AWS CLI in a buildspec to deploy Lambda functions lacks the managed deployment strategies (e.g., canary, linear) and automatic rollback capabilities that CodeDeploy provides, making it error-prone for production.

228
MCQeasy

A DevOps engineer is setting up a CI/CD pipeline using AWS CodePipeline and AWS CodeBuild. The build environment requires specific software packages that are not available in the default CodeBuild environment. What is the MOST efficient way to customize the build environment?

A.Create a separate pipeline to pre-build the environment.
B.Add install commands in the buildspec file to install the packages during each build.
C.Modify the buildspec file to set environment variables that include the software packages.
D.Create a custom Docker image with the required software and push it to Amazon ECR.
AnswerD

Creating a custom Docker image that includes all required software and pushing it to Amazon ECR lets you reference that image in the CodeBuild project's environment, so CodeBuild launches build containers from a pre-baked image. This avoids repeated package installation on every run, significantly reducing build time and removing dependence on internet-accessible package repositories during the build. The image provides a deterministic and immutable toolchain, which improves reproducibility, and you can version the image (using tags or image digests) and update it via its own build pipeline for rolling upgrades.

Why this answer

Creating a custom Docker image with the required software and pushing it to Amazon ECR allows you to pre-configure the build environment with all necessary packages. This approach avoids the overhead of installing packages during each build, reduces build time, and ensures consistency across builds. AWS CodeBuild supports custom images from Amazon ECR, making this the most efficient method for customizing the environment.

Exam trap

The trap here is that candidates often confuse environment variables with actual software installation, thinking that setting variables in the buildspec file can bring in packages, when in fact environment variables only affect runtime behavior, not the software available in the build environment.

How to eliminate wrong answers

Option A is wrong because creating a separate pipeline to pre-build the environment introduces unnecessary complexity and overhead; it does not directly customize the CodeBuild environment for the main pipeline. Option B is wrong because adding install commands in the buildspec file to install packages during each build is inefficient, as it repeats the installation process for every build, increasing build time and potential for failure. Option C is wrong because modifying the buildspec file to set environment variables does not install software packages; environment variables only configure runtime settings, not the software dependencies themselves.

229
MCQeasy

Refer to the exhibit. The IAM policy above is attached to a Lambda function's execution role. The Lambda function is supposed to publish custom metrics to CloudWatch using PutMetricData. However, the metrics are not appearing. What is the most likely reason?

A.The policy does not include the 'cloudwatch:PutMetricData' action.
B.The policy includes unnecessary actions that conflict with each other.
C.The function needs to specify a metric name and value when calling PutMetricData.
D.The policy uses a wildcard resource, which is not allowed for the PutMetricData action.
AnswerC

Granting cloudwatch:PutMetricData only authorizes the API call; it does not construct or complete it. The Lambda function must send a MetricDatum object that includes both a MetricName and a numeric Value (or StatisticValues/Values), along with other optional fields like Dimensions and Timestamp. If these required fields are missing or empty, the API call cannot create a metric data point, regardless of how permissive the attached IAM policy is.

Why this answer

The most likely reason the metrics are not appearing is that the Lambda function is not providing the required parameters—specifically a metric name and value—when calling PutMetricData. The IAM policy correctly grants the cloudwatch:PutMetricData action, but the API call itself must include at least a MetricName and a Value (or StatisticValues) in the MetricDatum array; otherwise, CloudWatch silently drops the request without publishing any metric. This is a common coding error where the execution role permissions are sufficient but the function logic is incomplete.

Exam trap

The trap here is that candidates assume the issue is always a missing IAM permission (Option A) when the real problem is often a missing required parameter in the API call, especially since PutMetricData returns success even with invalid data.

How to eliminate wrong answers

Option A is wrong because the policy explicitly includes 'cloudwatch:PutMetricData' as an action, so the Lambda function has the necessary permission to publish metrics. Option B is wrong because the presence of multiple actions (like PutMetricData, GetMetricData, ListMetrics) does not cause conflicts; IAM policies allow multiple actions and they do not interfere with each other unless there are explicit deny statements. Option D is wrong because the PutMetricData action supports a wildcard resource ('*') in the policy; CloudWatch metrics are global and do not require a specific resource ARN, so using '*' is both allowed and standard practice.

230
Multi-Selecteasy

A company is designing a disaster recovery strategy for its application. The application runs on EC2 instances and uses an RDS MySQL database. The RTO is 1 hour, and the RPO is 15 minutes. Which TWO approaches meet these requirements?

Select 2 answers
A.Use a warm standby strategy: run a scaled-down version of the application in the DR region with RDS Multi-AZ across regions.
B.Use a pilot light strategy: replicate data using RDS cross-region automated backups and have a small environment running in the DR region.
C.Use a read replica in the DR region and promote it on failover.
D.Use a Multi-Zone deployment with RDS in the same region.
E.Use a backup and restore strategy: take snapshots every hour and restore in the DR region on failover.
AnswersA, B

The warm standby pattern keeps a fully functional, scaled-down copy of the application running in the DR region, so the entire stack is ready for production traffic almost immediately after failover. RDS data is continuously replicated to the DR region (via cross-region read replicas or Aurora global replication), keeping the RPO near zero—well within the 15-minute target. Because compute, storage, and networking are already provisioned, RTO is also met by simply scaling up and shifting traffic rather than building infrastructure from scratch.

Why this answer

Options A and B are correct. A warm standby with RDS Multi-AZ across regions ensures a standby database is ready and can be promoted quickly, meeting the 1-hour RTO. A pilot light with RDS cross-region automated backups provides replication with a 15-minute RPO; a small environment is running, allowing faster failover than a full pilot light.

Option C is wrong because RDS read replicas do not support automatic failover; manual promotion can take longer than 1 hour. Option D is wrong because Multi-AZ in the same region does not protect against region failure. Option E is wrong because hourly snapshots meet RPO but restoring from snapshots typically exceeds the 1-hour RTO.

231
MCQhard

A company runs a microservices architecture on Amazon ECS with Fargate. Each service is deployed in its own ECS service. The company wants to ensure that if one Availability Zone (AZ) fails, the services can continue to operate with minimal impact. What is the MOST resilient task placement strategy?

A.Use a task placement constraint to run tasks on distinct instances.
B.Use a task placement strategy that uses the random algorithm.
C.Use a task placement strategy that uses the binpack algorithm to maximize resource utilization.
D.Use a task placement strategy that spreads tasks across Availability Zones.
AnswerD

A spread placement strategy with the field attribute:ecs.availability-zone explicitly distributes tasks evenly across the Availability Zones used by the cluster or service. For Fargate, this strategy works in conjunction with your VPC subnets, ensuring tasks are placed in each configured AZ before any are duplicated. This directly satisfies the requirement that a single AZ failure does not take down all instances of the microservice.

Why this answer

The 'spread across Availability Zones' strategy explicitly distributes ECS tasks across multiple AZs, ensuring that if one AZ fails, the remaining AZs continue to run the service. This is the most resilient approach for Fargate tasks, as it leverages the AZ isolation provided by AWS to minimize the blast radius of a single-AZ failure.

Exam trap

The trap here is that candidates often confuse 'high availability' with 'resource efficiency' and choose binpack (Option C) because it reduces cost, but the question explicitly asks for resilience, not cost optimization.

How to eliminate wrong answers

Option A is wrong because 'distinct instances' is a constraint for EC2 launch type, not Fargate; Fargate tasks run on AWS-managed infrastructure, so this constraint is irrelevant and does not provide AZ-level resilience. Option B is wrong because the 'random' algorithm distributes tasks without any awareness of AZ boundaries, potentially placing all tasks in a single AZ and leaving the service vulnerable to that AZ's failure. Option C is wrong because 'binpack' maximizes resource utilization by packing tasks onto the fewest underlying resources, which often concentrates tasks in one AZ, reducing resilience and increasing the impact of an AZ failure.

232
MCQhard

A company runs a microservices application on Amazon EKS. The DevOps team wants to collect and visualize metrics such as pod CPU and memory usage, and set up alerts. Which combination of AWS services should be used?

A.Prometheus and Grafana on EC2
B.AWS X-Ray and Amazon CloudWatch ServiceLens
C.AWS CloudTrail and Amazon CloudWatch Logs
D.Amazon CloudWatch Container Insights and CloudWatch Alarms
AnswerD

Amazon CloudWatch Container Insights automatically discovers and collects pod-level metrics from Amazon EKS, including CPU, memory, network, and disk utilization, and presents them through pre-built dashboards. By defining CloudWatch Alarms on these metrics, operators receive proactive notifications when thresholds are breached, making this a fully managed, integrated solution for monitoring and alerting on container resource usage.

Why this answer

Amazon CloudWatch Container Insights and CloudWatch Alarms. Container Insights collects, aggregates, and summarizes metrics and logs from containerized applications and microservices running on Amazon EKS. It provides pre-built dashboards for pod CPU, memory, network, and disk metrics.

CloudWatch Alarms can be set on these metrics to trigger notifications or automated actions. Option A (Prometheus and Grafana on EC2) is possible but not an AWS service combination; it requires manual setup and maintenance. Option B (AWS X-Ray and Amazon CloudWatch ServiceLens) is for distributed tracing and service maps, not for collecting CPU/memory metrics.

Option C (AWS CloudTrail and Amazon CloudWatch Logs) is for auditing API calls and storing log data, not for metrics visualization and alerting.

233
MCQeasy

A company runs a critical application on Amazon EC2 instances in an Auto Scaling group. To ensure high availability, the instances are deployed across three Availability Zones. Which additional step should the company take to protect against a regional failure?

A.Place all instances in a single Availability Zone to simplify management.
B.Use EC2 Dedicated Hosts to ensure capacity.
C.Increase the minimum size of the Auto Scaling group to 10 instances.
D.Deploy the application in a second AWS Region and use Route 53 with failover routing.
AnswerD

Deploying the application in a second AWS Region and using Route 53 failover routing gives you active-passive or active-active DNS-level failover: Route 53 health checks continuously monitor the primary endpoint, and when it is unhealthy, DNS queries are answered with the secondary Region's IP addresses. This directly addresses an entire Region becoming unavailable, as long as the secondary Region has the resources and the data needed to serve traffic. For a critical application, combine this with RDS Cross-Region Read Replicas or Aurora Global Database for proper data durability.

Why this answer

Deploying the application in a second AWS Region and using Route 53 failover routing provides protection against a regional failure, because the application remains available even if an entire AWS Region becomes unavailable. Route 53 health checks detect the failure and automatically redirect traffic to the standby Region. This is the only option that addresses regional-level disasters.

Exam trap

DOP-C02 often tests the misconception that multi-AZ deployment alone protects against regional failure, when true regional resilience requires a multi-Region architecture with Route 53 failover.

How to eliminate wrong answers

Option A is wrong because placing all instances in a single Availability Zone reduces availability and does not protect against regional failure; it actually increases risk. Option B is wrong because EC2 Dedicated Hosts provide physical server isolation for licensing or compliance, not regional redundancy. Option C is wrong because increasing the Auto Scaling group size within the same Region does not protect against a regional outage; all instances would still be in the affected Region.

234
MCQeasy

A company wants to encrypt data at rest in Amazon S3 using server-side encryption. Which AWS service can automatically manage the encryption keys with minimal configuration?

A.SSE-C
B.SSE-KMS
C.SSE-S3
D.Client-side encryption
AnswerC

SSE-S3 is correct because it is Amazon S3's built-in server-side encryption that requires no additional configuration that the customer must perform. With SSE-S3, Amazon owns and manages the encryption keys on your behalf, using AES-256 with 256-bit keys, and applies encryption automatically to new objects written to S3. This provides a simple, fully managed, and effortless way to encrypt data at rest, aligning exactly with the stated requirement.

Why this answer

SSE-S3 (Server-Side Encryption with S3-Managed Keys) is the correct answer because it requires minimal configuration: you simply enable it on the bucket or object, and AWS fully manages the encryption keys, including rotation and protection, without any additional setup or key management overhead.

Exam trap

The trap here is that candidates often confuse SSE-S3 with SSE-KMS, assuming that any key management service (KMS) is required for automated encryption, but SSE-S3 provides fully automated key management with even less configuration than SSE-KMS.

How to eliminate wrong answers

Option A (SSE-C) is wrong because it requires you to provide and manage your own encryption keys, which adds configuration complexity and does not automate key management. Option B (SSE-KMS) is wrong because while it automates key management, it requires you to create and configure a KMS key, set IAM policies, and optionally manage key rotation, which is more configuration than SSE-S3. Option D (Client-side encryption) is wrong because it requires you to encrypt data before uploading to S3, meaning you must manage keys and encryption logic entirely on the client side, which is the opposite of minimal configuration.

235
Multi-Selecteasy

Which TWO actions should be taken to ensure a highly available and resilient architecture for a critical web application on AWS? (Choose two.)

Select 2 answers
A.Enable Amazon CloudFront with multiple origins.
B.Use an Auto Scaling group to maintain a desired number of instances.
C.Use a Multi-AZ RDS deployment with read replicas.
D.Store backups in a different AWS Region.
E.Deploy the application across multiple Availability Zones.
AnswersB, E

An Auto Scaling group with a desired capacity continuously monitors instance health via EC2 status checks and optionally Elastic Load Balancing health checks, automatically terminating and relaunching failed instances. By spreading the ASG across multiple Availability Zones and setting the desired count, you ensure that if an instance or an entire AZ fails, replacement capacity is launched to maintain the required number of instances, making this a core high-availability mechanism.

Why this answer

Correct: B and E. Option B ensures that the desired number of EC2 instances is maintained, providing automatic scaling and fault tolerance. Option E deploys the application across multiple Availability Zones, which protects against an AZ failure.

Option A (CloudFront) enhances content delivery but does not directly ensure high availability of the web application. Option C (Multi-AZ RDS with read replicas) improves read performance and provides disaster recovery, but write availability depends on the primary instance. Option D (backups in a different region) is for disaster recovery, not for immediate availability.

236
Multi-Selecteasy

A company wants to protect its AWS account credentials. Which TWO practices are recommended by AWS? (Choose TWO.)

Select 2 answers
A.Generate and share access keys for all users.
B.Store IAM user passwords in a shared document.
C.Enable multi-factor authentication (MFA) for privileged users.
D.Use the root user for daily administrative tasks.
E.Use IAM roles for applications that require AWS access.
AnswersC, E

Enabling MFA for privileged users requires a second authentication factor, so a stolen or phished password alone cannot produce a console login or an authenticated API call. This control directly blocks account takeover from reused passwords, keylogging, or credential stuffing, because the attacker must also possess the virtual or hardware token. AWS recommends enforcing MFA with an IAM policy that explicitly denies actions unless aws:MultiFactorAuthPresent is true, which makes MFA mandatory rather than optional.

Why this answer

Option C is correct because AWS recommends enabling multi-factor authentication (MFA) for privileged users, adding a second authentication factor (such as a virtual MFA device, hardware TOTP token, or FIDO2 security key) so that a stolen password alone cannot compromise the account. Option E is correct because IAM roles provide temporary credentials via AWS STS (AssumeRole), so applications, EC2 instances, and Lambda functions can access AWS services without embedding long-lived access keys in code or configuration. Option A is wrong because AWS advises against generating and sharing access keys; keys are long-lived credentials that should be rotated and never shared, and permissions should be granted per identity.

Option B is wrong because storing IAM user passwords in a shared document exposes credentials; passwords should be managed with strong policies and never stored in plaintext shared locations. Option D is wrong because the root user has unrestricted access and AWS strongly recommends locking away root credentials, enabling MFA on root, and using IAM users or roles with least privilege for daily administrative tasks.

Exam trap

DOP-C02 often tests the misconception that access keys are the standard way to grant AWS access to applications, when in fact IAM roles with temporary credentials are the recommended approach.

237
Multi-Selectmedium

A company is designing a highly available architecture for a web application using AWS services. The application must be resilient to the failure of an entire AWS Region. Which TWO strategies should the company implement? (Choose TWO.)

Select 2 answers
A.Deploy the application in multiple AWS Regions and use Route 53 with failover routing policy.
B.Use Amazon CloudFront with multiple origins in the same region.
C.Enable S3 cross-Region replication for static assets.
D.Configure Amazon RDS for Multi-AZ and enable cross-Region read replicas.
E.Use Auto Scaling groups in a single region with multiple Availability Zones.
AnswersA, D

This configuration provides global DNS-level failover by associating Route 53 health checks with endpoints in each region. If the primary region's health check fails, Route 53 automatically rewrites DNS responses to direct traffic to the standby region, with failover typically occurring within a few minutes depending on TTL and health-check intervals. However, this requires the application to be designed for multi-region operation, with data replication between regions and compute capacity pre-provisioned in the secondary region to actually serve traffic. This is the core pattern for regional disaster recovery and meets the requirement for a highly available architecture across regions.

Why this answer

Deploying to multiple regions with Route 53 failover provides cross-region disaster recovery. Option D is correct because using Amazon RDS Multi-AZ with cross-Region read replicas or Aurora Global Database ensures database resilience across regions. Option B is wrong because CloudFront alone does not provide compute failover.

Option C is wrong because S3 cross-Region replication is for data, not compute. Option E is wrong because single-region Auto Scaling does not protect against region failure.

238
MCQeasy

A company wants to receive notifications when an EC2 instance's CPU utilization exceeds 90% for 10 consecutive minutes. Which AWS service should be used?

A.Amazon CloudWatch alarm
B.AWS Config rule
C.AWS CloudTrail event
D.Amazon Inspector
AnswerA

A CloudWatch alarm is the correct mechanism because it continuously evaluates an EC2 instance's time-series metrics, such as CPUUtilization or StatusCheckFailed, against a defined threshold. When the metric crosses the threshold, the alarm state transitions to ALARM and automatically publishes a message to an SNS topic, which can deliver email, SMS, or invoke a Lambda function. Alarms can also be configured to perform EC2 actions like stop or reboot, making them the native AWS service for threshold-based metric notification.

Why this answer

Amazon CloudWatch alarms monitor specified metrics (like CPU utilization) and trigger actions (e.g., SNS notification) when a threshold is breached for a given period. Option A is correct. Option B is incorrect because AWS Config rules evaluate configuration compliance, not metric thresholds.

Option C is incorrect because AWS CloudTrail records API activity, not metric monitoring. Option D is incorrect because Amazon Inspector assesses security vulnerabilities, not performance metrics.

239
MCQhard

A DevOps engineer is configuring an AWS Lambda function that processes messages from an Amazon SQS queue. The function must handle transient failures gracefully and avoid reprocessing the same message multiple times. The engineer sets the maximum receives to 3 and configures a dead-letter queue (DLQ) for the source queue. After several days, the engineer notices that some messages are being processed more than once, even though they were successfully processed. What is the MOST likely cause of this issue?

A.The dead-letter queue is misconfigured, causing messages to be sent back to the source queue.
B.The SQS queue's visibility timeout is shorter than the Lambda function's execution time.
C.The Lambda function is not idempotent, so it processes the same message differently each time.
D.The Lambda function's timeout is set too low, causing it to fail and retry messages.
AnswerB

If the visibility timeout is shorter than the time it takes for the Lambda function to process the message, the message becomes visible again in the queue before the function completes. Another Lambda invocation can then receive and process the same message, leading to duplicate processing. This is a common misconfiguration. The visibility timeout should be set to at least the function's timeout plus a buffer to prevent this. Even if the function eventually succeeds, the duplicate processing has already occurred.

Why this answer

The most likely cause is that the SQS queue's visibility timeout is shorter than the Lambda function's execution time. When a Lambda function polls an SQS queue, it receives a batch of messages and each message becomes invisible for the duration of the visibility timeout. If the function takes longer than that timeout to process a message, the message becomes visible again and can be picked up by another Lambda invocation.

This results in duplicate processing. To resolve this, set the visibility timeout to at least the function's timeout plus a buffer, and ensure the function is idempotent.

Exam trap

The trap here is focusing on idempotency as the cause of duplicate processing, when idempotency is a mitigation for duplicates, not the reason they occur.

240
MCQhard

A DevOps team is designing a configuration management solution for a microservices architecture running on Amazon ECS. The team wants to ensure that container configurations are automatically updated when a new version of a parameter is stored in AWS Systems Manager Parameter Store. Which approach best meets this requirement with minimal operational overhead?

A.Use AWS AppConfig to create a configuration profile that references the parameter. Configure a Lambda function as a validator and deploy strategy. When the parameter changes, AppConfig triggers a deployment that updates the ECS service.
B.Use an AWS CloudFormation custom resource that updates the ECS service when the parameter changes.
C.Use a CI/CD pipeline that monitors the parameter store and triggers a new build and deploy of the container image with the updated parameter.
D.Use Amazon EventBridge to detect changes to the parameter and invoke a Lambda function that updates the ECS task definition and forces a new deployment.
AnswerA

AWS AppConfig is purpose-built for this exact scenario because it separates configuration from code and provides a managed deployment lifecycle. By creating a configuration profile that references the SSM parameter, AppConfig treats each change as a new configuration version, runs your Lambda validator to catch malformed or unsafe values, and then rolls out the update using the specified deploy strategy (e.g., linear or canary with bake time). The ECS service is updated through AppConfig's integration with ECS—either via the AppConfig agent that supplies the new configuration or by triggering a task definition update—so the application receives the new value without a container rebuild. This gives you controlled rollout, automatic rollback on alarms, and auditability, which is exactly what a configuration management solution should provide.

Why this answer

AWS AppConfig is purpose-built for managing application configuration and supports automatic deployment of configuration changes to ECS services when a parameter in Systems Manager Parameter Store is updated. By creating a configuration profile that references the parameter, AppConfig can trigger a deployment that updates the ECS service without requiring custom code or manual intervention, minimizing operational overhead.

Exam trap

The trap here is that candidates may assume EventBridge with Lambda (Option D) is the simplest solution, but they overlook that AppConfig provides a managed, lower-overhead alternative specifically designed for configuration management and automatic deployment to ECS.

How to eliminate wrong answers

Option B is wrong because AWS CloudFormation custom resources require manual invocation or a separate trigger to execute the update logic; they do not automatically detect parameter changes and would add operational overhead for polling or event handling. Option C is wrong because using a CI/CD pipeline to rebuild and redeploy container images for every parameter change is inefficient and introduces unnecessary overhead, as the container image itself does not change—only the runtime configuration does. Option D is wrong because while EventBridge can detect parameter changes and invoke a Lambda function, this approach requires custom code to update the ECS task definition and force a new deployment, increasing operational complexity compared to AppConfig's managed deployment strategy.

241
MCQeasy

A company wants to automate the rotation of IAM user access keys every 90 days. Which AWS service can be used to achieve this?

A.Store the access keys in AWS Secrets Manager and enable automatic rotation.
B.Use AWS CloudTrail to detect old keys and send notifications to administrators.
C.Use IAM's built-in access key rotation feature.
D.Use AWS Config with a custom Lambda function to rotate keys when they are older than 90 days.
AnswerD

AWS Config enables this by hosting a custom rule that invokes a Lambda function on a schedule using the rule's 'maximum execution frequency' setting, evaluating all IAM users' access keys for age. The Lambda function can use IAM API calls such as ListAccessKeys and GetAccessKeyLastUsed to identify keys older than 90 days, then rotate them by creating a new access key, deactivating the old key, and deleting the old key after a grace period. This is a well-known pattern because AWS Config provides the periodic orchestration and compliance evaluation while Lambda handles the actual IAM operations.

Why this answer

AWS IAM does not have a built-in automatic rotation feature for access keys. AWS Secrets Manager can store secrets and rotate them, but it does not natively rotate IAM access keys; it supports rotation for RDS, Redshift, and DocumentDB, but not IAM. AWS CloudTrail is a logging service and cannot rotate keys.

However, AWS Config can be used with a custom AWS Lambda function to create a rule that triggers rotation when access keys are older than 90 days. Therefore, option D is the correct answer.

242
MCQhard

A company uses AWS CodeDeploy to deploy applications to an Auto Scaling group. During a deployment, the new instances fail the health check and are terminated. The deployment fails. The team wants to automatically roll back to the previous working version. What should they do?

A.Set up an Auto Scaling lifecycle hook to terminate instances and trigger a rollback.
B.Configure the deployment group to automatically roll back when a deployment fails.
C.Manually redeploy the last successful deployment revision after investigating the failure.
D.Configure the deployment group to automatically redeploy the same revision on failure.
AnswerB

Automatic rollback in the deployment group triggers CodeDeploy to redeploy the last known-good revision once the health check failures terminate instances and the deployment fails, satisfying the requirement to restore the previous working version without manual intervention.

Why this answer

AWS CodeDeploy provides a built-in rollback configuration that can be triggered automatically when a deployment fails. By enabling automatic rollback in the deployment group settings, CodeDeploy will redeploy the last successful revision when the current deployment fails health checks, without requiring manual intervention or additional infrastructure.

Exam trap

The trap here is that candidates may confuse Auto Scaling lifecycle hooks with CodeDeploy rollback mechanisms, or think that redeploying the same revision (option D) would fix the issue, when in fact it would just repeat the failure.

How to eliminate wrong answers

Option A is wrong because Auto Scaling lifecycle hooks are used to perform custom actions during instance launch or termination (e.g., draining connections or running scripts), but they do not trigger CodeDeploy rollbacks; rollback logic must be configured within CodeDeploy itself. Option C is wrong because manually redeploying the last successful revision is a valid recovery method but does not meet the requirement for automatic rollback; the team wants an automated solution, not manual steps. Option D is wrong because redeploying the same revision on failure would repeat the same failing deployment, not restore the previous working version; automatic rollback specifically redeploys the last known good revision, not the failed one.

243
MCQmedium

A company manages its AWS infrastructure using AWS CloudFormation templates stored in an Amazon S3 bucket. The DevOps team needs to enforce that all new CloudFormation stacks are created only from templates that have been validated by AWS CloudFormation Guard. The team wants to integrate this validation into their existing CI/CD pipeline built with AWS CodePipeline. Which approach will meet these requirements with the LEAST operational overhead?

A.Add a CodeBuild action to the pipeline that runs `cfn-guard validate` against the template and fails the build if validation fails.
B.Use AWS CloudFormation change sets to preview changes and require manual approval before execution.
C.Enable AWS CloudFormation drift detection on all stacks and automatically remediate any drift using AWS Systems Manager Automation.
D.Configure an AWS Lambda function that downloads the template and runs `cfn-guard validate`, then triggers the pipeline only if validation succeeds.
AnswerA

AWS CloudFormation Guard is a policy-as-code tool that can be run in CodeBuild. Integrating `cfn-guard validate` as a build step ensures templates are validated before deployment, and the pipeline stops on failure. This requires minimal setup: install Guard in the build environment, run the command, and check the exit code. It is a native, serverless approach with no infrastructure to manage.

Why this answer

Running AWS CloudFormation Guard in CodeBuild is the most efficient way to enforce template validation. Guard integrates seamlessly with CodePipeline, and CodeBuild provides a managed environment where you can install and execute `cfn-guard validate`. This ensures that only validated templates proceed to deployment, with minimal operational effort compared to custom Lambda functions or post-deployment checks.

Exam trap

The trap here is assuming that CloudFormation change sets or drift detection can enforce template policy validation, when they actually only show differences or detect drift after deployment.

244
MCQmedium

A company is running a critical web application on Amazon EC2 instances behind an Application Load Balancer (ALB). The DevOps team wants to monitor HTTP 5xx errors and receive alerts when the error rate exceeds 5% over a 5-minute period. Which combination of services and configurations should be used to meet these requirements?

A.Enable CloudWatch Logs for the ALB and use CloudWatch Logs Insights to query 5xx logs, then create a metric filter and alarm.
B.Configure AWS Config rules to check ALB 5xx error counts and trigger alarms.
C.Use CloudWatch ALB metrics (HTTPCode_ELB_5XX_Count) and create a CloudWatch Alarm on the Sum statistic with a threshold based on total request count.
D.Use AWS X-Ray to trace requests and create a CloudWatch alarm based on X-Ray error rate.
AnswerC

Correct: The Application Load Balancer natively emits the HTTPCode_ELB_5XX_Count metric to CloudWatch, representing the number of 5xx responses returned by the load balancer itself. Create a CloudWatch Alarm on this metric using the Sum statistic over a period (e.g., 5 minutes) and set a threshold, optionally using a math expression to divide by RequestCount to track the error ratio. This is the simplest and most direct method because it uses existing metrics with no additional setup, latency, or cost.

Why this answer

ALB automatically publishes the `HTTPCode_ELB_5XX_Count` metric to CloudWatch, and you can create a CloudWatch alarm using the `Sum` statistic over a 5-minute period. To detect when the error rate exceeds 5%, you need to combine this metric with the `RequestCount` metric in a math expression (e.g., `m1/m2*100 > 5`) or use a composite alarm, as the alarm threshold must be based on the ratio of 5xx errors to total requests, not just the raw count.

Exam trap

The trap here is that candidates often assume they need to parse logs (Option A) or use a separate tracing service (Option D) for error rate monitoring, when in fact the ALB's built-in CloudWatch metrics and metric math provide a simpler, real-time, and cost-effective solution without additional log ingestion or query overhead.

How to eliminate wrong answers

Option A is wrong because CloudWatch Logs Insights is a query tool for analyzing log data, not a real-time alerting mechanism; while you can create a metric filter from ALB logs to count 5xx errors, this approach introduces latency and additional cost, and it is not the simplest or most direct method when ALB metrics are already available. Option B is wrong because AWS Config rules are designed for compliance and resource configuration auditing (e.g., checking if ALB is configured with a specific security policy), not for monitoring real-time error rates or triggering alarms on metric thresholds. Option D is wrong because AWS X-Ray traces individual requests to identify latency and errors, but it does not aggregate HTTP 5xx error rates over a time window or natively publish a metric that can be used directly in a CloudWatch alarm for this specific requirement.

245
MCQmedium

A company uses AWS CodePipeline with a manual approval stage before deploying to production. The approval notification is sent via Amazon SNS. The approvers report that they are not receiving the email notifications. What should the DevOps engineer check first?

A.Confirm that the email subscriptions to the SNS topic have been confirmed by clicking the link in the initial confirmation email.
B.Ensure that the IAM role for CodePipeline has permission to publish to the SNS topic.
C.Verify that the SNS topic's subscription has a filter policy that matches the approval event.
D.Check the email recipients' mailbox quota to see if it is full.
AnswerA

SNS email subscriptions start in the 'Pending Confirmation' state; SNS sends an initial confirmation email to the endpoint and no notifications are delivered until the recipient clicks the link in that message. If the recipient never confirms, the subscription stays inactive, so even though CodePipeline publishes approval messages to the topic, SNS discards them for that unconfirmed endpoint. Therefore, confirming the subscription is the correct first troubleshooting step.

Why this answer

Amazon SNS requires that email subscribers confirm their subscription by clicking the link in the initial confirmation email before they can receive notifications. If the approvers never confirmed the subscription, the SNS topic will not deliver any messages to them, even though the pipeline and IAM permissions are correctly configured. This is the most common cause of missing SNS email notifications and should be the first check.

Exam trap

The trap here is that candidates often jump to IAM permissions or SNS configuration details, overlooking the fundamental requirement that email subscriptions must be explicitly confirmed before any notifications can be delivered.

How to eliminate wrong answers

Option B is wrong because CodePipeline does not need an IAM role to publish to SNS; the pipeline uses an SNS topic ARN and the publish action is performed by the pipeline service itself, which already has the necessary permissions via the service-linked role. Option C is wrong because SNS filter policies are optional and used to filter messages based on attributes; the approval notification does not require a filter policy to be delivered, and a missing or mismatched filter policy would not prevent the initial subscription confirmation email. Option D is wrong because mailbox quota issues would cause bounce or rejection after delivery, but the core problem is that the subscription itself is not confirmed, so no emails are ever sent.

246
Multi-Selectmedium

A company uses AWS CloudFormation to provision infrastructure. They have a stack that creates an Amazon RDS DB instance. They want to update the stack to change the DB instance class from db.t2.micro to db.t3.medium. Which THREE of the following must be true for the update to succeed? (Choose three.)

Select 3 answers
A.The stack must not be in a state that prevents updates, such as ROLLBACK_COMPLETE.
B.Deletion protection must be disabled on the DB instance.
C.The new DB instance class must be available in the same VPC and subnet group as the existing DB instance.
D.The IAM role used by CloudFormation must have permissions to modify the RDS instance.
E.A change set must be created and executed for the update.
AnswersA, C, D

CloudFormation refuses to update any stack in a failed terminal state such as ROLLBACK_COMPLETE, because the stack's resources are in an inconsistent or partially provisioned state. The stack must be deleted and recreated (or, if applicable, brought back to a workable state via ContinueUpdateRollback for other rollback states) before an update operation can run.

Why this answer

CloudFormation stacks in a terminal failure state like ROLLBACK_COMPLETE cannot be updated; they must be deleted and recreated. This is a fundamental constraint of the CloudFormation service, as the stack is considered non-functional and unable to process further operations.

Exam trap

The trap here is that candidates often confuse deletion protection with modification protection, assuming it blocks all changes, when in fact it only blocks deletion operations.

247
MCQhard

A company uses AWS Organizations with multiple accounts. The security team notices that an IAM user in the production account has been making changes to security group rules that are not compliant with the company's policy. The team wants to automatically revoke any non-compliant security group rules and notify the security team. What is the MOST efficient way to achieve this?

A.Create a CloudWatch alarm on the SecurityGroupEvent metric to notify the security team.
B.Apply a Service Control Policy (SCP) that denies changes to security groups in the production account.
C.Set up a CloudTrail trail that logs security group modifications and use Amazon Detective to analyze the changes.
D.Use an AWS Config managed rule to detect non-compliant security group rules, and configure an automatic remediation action with AWS Systems Manager Automation.
AnswerD

AWS Config's managed rules continuously evaluate security group configurations against policies such as 'restricted-ssh' or 'vpc-sg-open-only-to-a-specific-port'. When a rule detects a non-compliant inbound rule, a configured AWS Systems Manager Automation document, such as AWS-DisablePublicAccessForSecurityGroup, automatically removes or tightens the offending rule, providing the required detect-and-remediate control loop without manual intervention.

Why this answer

AWS Config managed rules can continuously evaluate security group rules against a desired policy (e.g., disallowing SSH from 0.0.0.0/0). When a non-compliant change is detected, AWS Config can trigger an automatic remediation action using an AWS Systems Manager Automation document that revokes the offending rule. This provides both detection and automated correction without manual intervention, making it the most efficient solution.

Exam trap

The trap here is that candidates often confuse detective controls (CloudTrail, CloudWatch alarms) with corrective controls (AWS Config remediation), and fail to recognize that SCPs are preventive and cannot selectively revoke existing non-compliant rules.

How to eliminate wrong answers

Option A is wrong because CloudWatch alarms on the SecurityGroupEvent metric can only notify on the occurrence of an event, not automatically revoke the non-compliant rule; it lacks remediation capability. Option B is wrong because Service Control Policies (SCPs) apply to all IAM users and roles in an account and cannot selectively revoke specific security group rules after they are created; SCPs are preventive, not detective or corrective, and would block all security group changes, which may be too restrictive. Option C is wrong because CloudTrail logs and Amazon Detective can analyze changes after the fact but cannot automatically revoke non-compliant rules; they provide visibility and investigation, not automated remediation.

248
MCQhard

Refer to the exhibit. The CloudWatch alarm is set on CPUUtilization. The instance's CPU at 10:00 is 75%, at 10:05 is 82%, and at 10:10 is 85%. Will the alarm trigger?

A.Yes, because the CPU exceeded 80% at 10:05.
B.Yes, but only after 10:10 when two consecutive periods are above the threshold.
C.No, because the alarm uses Average statistic and the average over the entire time is below 80%.
D.No, because the first data point at 10:00 is below the threshold.
AnswerB

Correct. When the alarm evaluates at 10:10, it considers the two most recent consecutive periods—10:05 and 10:10—both of which have CPU utilization above 80%. Since the alarm is configured with two evaluation periods and the threshold of 80%, both consecutive periods satisfy the condition, causing the alarm to enter ALARM state. At 10:05, only one period had breached, so the alarm was still OK.

Why this answer

The alarm evaluates 2 consecutive periods (each 5 minutes). At 10:05, the average for that period is 82% (above 80), but the previous period (10:00) average is 75% (below 80). So only one period is breached.

At 10:10, the average for that period is 85% (above 80), and the previous period (10:05) average is 82% (above 80). So two consecutive periods are breached, triggering the alarm.

249
Multi-Selecteasy

A company is adopting Infrastructure as Code (IaC) using AWS CloudFormation. They want to ensure that stack updates are safe and minimize the risk of resource replacement. Which TWO of the following strategies should they use?

Select 2 answers
A.Always create a change set before executing a stack update.
B.Use a stack policy to prevent updates to critical resources.
C.Perform updates directly from the AWS Management Console to see immediate results.
D.Delete the stack and create a new one with the updated template.
E.Use the --disable-rollback flag to avoid unnecessary rollbacks.
AnswersA, B

A change set is a read-only preview of the exact modifications CloudFormation will make to a running stack, including whether each resource is created, updated, replaced, or deleted. By generating a change set before executing, you can inspect the action on every resource and confirm that the update does not introduce unintended destructive changes or interruption. Execution is a separate, explicit step, which gives engineers a checkpoint in the IaC workflow.

Why this answer

Creating a change set before executing a stack update allows you to review the proposed changes, including whether any resources will be replaced or interrupted. This provides a safety net by letting you validate the impact of the update before committing, reducing the risk of unintended resource replacement.

Exam trap

The trap here is that candidates often confuse the --disable-rollback flag with a safety mechanism, but it actually prevents recovery from failures and does nothing to mitigate resource replacement risks.

250
MCQhard

A company uses AWS CloudTrail to monitor API activity. The DevOps team needs to ensure that any deletion of an S3 bucket is detected in real time and triggers an automated response. Which combination of AWS services should be used to meet these requirements?

A.Use CloudWatch Logs to monitor the logs, and create a metric filter to trigger an alarm when the DeleteBucket event appears.
B.Configure S3 event notifications to send to an SQS queue, and poll the queue with a Lambda function.
C.Configure CloudTrail to deliver logs to an S3 bucket, and use S3 event notifications to invoke a Lambda function.
D.Send CloudTrail logs to CloudWatch Logs, create a CloudWatch Events rule matching the DeleteBucket event, and target a Lambda function.
AnswerD

Configuring CloudTrail to stream events to CloudWatch Logs provides near-real-time delivery of management events, including DeleteBucket. A CloudWatch Events (EventBridge) rule can use an event pattern to match the specific DeleteBucket API call and directly target a Lambda function as its target. This creates a low-latency, event-driven pipeline that automatically invokes the Lambda function without polling or human intervention, enabling immediate remediation or alerting.

Why this answer

CloudTrail logs can be sent to CloudWatch Logs, and a CloudWatch Events rule (now Amazon EventBridge) can be created to match the DeleteBucket event and trigger a Lambda function for automated response in real time. Option A is incorrect because CloudWatch Logs monitors log data, but a metric filter on DeleteBucket events in CloudTrail logs can trigger an alarm, but that is not as direct as EventBridge. More importantly, CloudWatch Logs metric filters have latency and are not the best for real-time response.

Option B is incorrect because S3 event notifications are for operations on objects, not for bucket-level operations like deletion. Option C is incorrect because while CloudTrail can deliver logs to S3, S3 event notifications are not triggered by CloudTrail log file delivery; they are for object-level events. Using EventBridge directly with CloudTrail is the correct real-time approach.

251
MCQmedium

A company is using AWS CodeBuild to run integration tests. The tests require access to an Amazon RDS instance in a private subnet. The CodeBuild project is configured with a VPC ID, subnet IDs, and security group IDs. However, the tests fail with a connection timeout. What is the MOST likely cause?

A.The security group attached to the RDS instance does not allow inbound traffic from the CodeBuild security group.
B.The CodeBuild project does not have internet access to download packages.
C.The CodeBuild project is not associated with a VPC.
D.The RDS instance is not publicly accessible and requires a NAT gateway.
AnswerA

CodeBuild placed in the VPC uses its own security group as the traffic source. If the RDS security group lacks an inbound rule permitting that CodeBuild security group on the database port, connections time out, matching the stem's VPC-configured constraint.

Why this answer

The most likely cause is that the security group attached to the RDS instance does not allow inbound traffic from the CodeBuild security group. CodeBuild runs inside the VPC using the specified security group, so it sends traffic to the RDS instance on port 3306 (or the appropriate database port). If the RDS security group's inbound rules do not explicitly permit traffic from the CodeBuild security group (or its CIDR), the connection is dropped, resulting in a timeout.

Exam trap

The trap here is that candidates often assume a NAT gateway or internet access is required for VPC-based resources, but the core issue is security group ingress rules, not network connectivity to the internet.

How to eliminate wrong answers

Option B is wrong because CodeBuild projects configured with a VPC can access the internet via a NAT gateway or VPC endpoints if needed, but the failure here is a connection timeout to RDS, not a package download issue. Option C is wrong because the question states the CodeBuild project is configured with a VPC ID, subnet IDs, and security group IDs, so it is associated with a VPC. Option D is wrong because RDS instances in private subnets do not need to be publicly accessible; CodeBuild can reach them directly via the VPC without a NAT gateway, as long as security group rules and network ACLs permit the traffic.

252
Multi-Selecteasy

A company is deploying a web application on Amazon ECS with Fargate. The application consists of a frontend service and a backend service. The DevOps team needs to ensure that the frontend service can communicate with the backend service securely without exposing the backend to the internet. Which THREE steps should the team take? (Choose THREE.)

Select 3 answers
A.Deploy the backend service in a private subnet with no internet access.
B.Use AWS Cloud Map service discovery for the backend service.
C.Configure a security group for the backend service that allows inbound traffic only from the frontend service's security group.
D.Deploy the backend service in a public subnet with an internet-facing Application Load Balancer.
E.Use an internet-facing Network Load Balancer for the backend service.
AnswersA, B, C

Placing the backend ECS service in a private subnet with no route to an internet gateway ensures it has no public IP and cannot be reached from the internet. This is correct for internal-only workloads because the backend does not need outbound internet access, and any required image pulls can be handled via VPC endpoints or pre-pulled images. This isolation reduces the attack surface and aligns with the principle of least privilege.

Why this answer

Deploying the backend service in a private subnet with no internet access ensures that the backend is not reachable from the internet, which is a fundamental security requirement. In Amazon ECS with Fargate, tasks in a private subnet use an elastic network interface (ENI) with no public IP address, and outbound traffic can be routed through a NAT gateway if needed, but inbound traffic from the internet is blocked. This isolates the backend from direct external exposure while still allowing communication from the frontend service within the same VPC.

Exam trap

The trap here is that candidates might think a load balancer is required for service-to-service communication in ECS, but AWS Cloud Map service discovery combined with security group rules can achieve secure, direct communication without exposing the backend to the internet.

253
MCQhard

A company runs a critical application on EC2 instances in an Auto Scaling group behind an ALB. They want to ensure that if an instance fails, the application remains available with minimal disruption. Which combination of services provides the best resilience?

A.Auto Scaling group with minimum 2 in a single AZ.
B.EC2 instance recovery with CloudWatch alarms.
C.Auto Scaling group with desired capacity of 2 and a lifecycle hook.
D.Auto Scaling group with ELB health checks and multiple AZs.
AnswerD

An Auto Scaling group spanning multiple Availability Zones and using Elastic Load Balancing health checks meets the resilience requirement: the ELB sends HTTP/HTTPS health checks to each instance, and when an instance is marked unhealthy, the ASG terminates it and launches a new one to maintain desired capacity. Multi-AZ distribution ensures that even if one AZ fails, the remaining instances continue serving traffic, and the ASG can launch replacements in other AZs. This combination provides both automated fault replacement and cross-AZ high availability.

Why this answer

Deploying an Auto Scaling group across multiple Availability Zones (AZs) with Elastic Load Balancer (ALB) health checks ensures that if an EC2 instance fails in one AZ, the ALB automatically routes traffic to healthy instances in other AZs, and Auto Scaling replaces the failed instance. This combination provides both fault isolation and automated recovery, minimizing disruption to the application.

Exam trap

The trap here is that candidates often think a single AZ with multiple instances (Option A) or instance recovery (Option B) provides sufficient resilience, but they overlook the need for AZ-level fault isolation and integrated health-check-driven replacement that only multi-AZ Auto Scaling with ELB health checks provides.

How to eliminate wrong answers

Option A is wrong because a single AZ is a single point of failure; if that AZ experiences an outage, all instances become unavailable regardless of the minimum count. Option B is wrong because EC2 instance recovery with CloudWatch alarms only recovers the same instance (e.g., after a hardware failure) but does not handle AZ-level failures or provide load balancing; it also does not automatically replace instances that fail health checks. Option C is wrong because a lifecycle hook is used for custom actions during instance launch or termination (e.g., draining connections), not for resilience; it does not distribute instances across AZs or provide health-check-based replacement.

254
Multi-Selecteasy

Which TWO AWS services can be used to manage secrets and database credentials securely? (Choose TWO.)

Select 2 answers
A.AWS CloudFormation
B.AWS Secrets Manager
C.Amazon S3
D.AWS Identity and Access Management (IAM)
E.AWS Systems Manager Parameter Store
AnswersB, E

AWS Secrets Manager is a purpose-built service for centrally managing the entire secret lifecycle, including storing, retrieving, and automatically rotating secrets such as database credentials, API keys, and OAuth tokens. It enforces fine-grained IAM access policies, integrates with AWS KMS for encryption, and supports automatic rotation via AWS Lambda with built-in integrations for RDS, Redshift, and DocumentDB. This makes it the most comprehensive choice for managing secrets.

Why this answer

AWS Secrets Manager is purpose-built for securely storing, rotating, and managing secrets such as database credentials, API keys, and other sensitive data. It provides built-in integration with Amazon RDS, Redshift, and DocumentDB to automatically rotate credentials on a schedule, eliminating the need for manual updates. This makes it a correct choice for the question's requirement to manage secrets and database credentials securely.

Exam trap

The trap here is that candidates often confuse AWS Systems Manager Parameter Store (which can store secure strings) with a full secrets management solution, but Parameter Store lacks native automatic rotation and is better suited for configuration data rather than database credentials that require scheduled rotation.

255
Multi-Selectmedium

A company uses Amazon CloudWatch for monitoring. The operations team wants to receive an alert when an EC2 instance's status check fails for 2 consecutive minutes. Which THREE resources should the team configure? (Choose three.)

Select 3 answers
A.CloudWatch Events rule
B.CloudWatch Logs
C.CloudWatch alarm
D.EC2 StatusCheckFailed metric
E.Amazon SNS topic
AnswersC, D, E

A CloudWatch alarm is the correct monitoring construct to watch a metric such as StatusCheckFailed or CPUUtilization. It evaluates the metric against a threshold over a specified number of evaluation periods and transitions to ALARM, OK, or INSUFFICIENT_DATA, then triggers a configured SNS action. The alarm is the central component that converts raw metric data into an operational notification, making it the appropriate mechanism for alerting on the instance's status check result.

Why this answer

To alert when an EC2 instance's status check fails for 2 consecutive minutes, you need to create a CloudWatch alarm on the StatusCheckFailed metric (options C and D). The alarm needs to send notifications via an SNS topic (option E). Option A (CloudWatch Events rule) is not used for metric-based alerts; CloudWatch Events triggers on events or schedules, not metric thresholds.

Option B (CloudWatch Logs) is for log data, not metrics.

Exam trap

A common trap is confusing CloudWatch Events with CloudWatch Alarms. CloudWatch Events are for event-driven actions based on state changes or schedules, not for monitoring metric thresholds over time. Metric alarms require the CloudWatch Alarm resource.

256
MCQmedium

A DevOps team observes that an Amazon CloudFront distribution is returning HTTP 504 errors for a small percentage of requests. The origin is an Application Load Balancer (ALB) that distributes traffic to EC2 instances. The team has already checked the ALB's access logs and found that the ALB returns 200 OK for all requests. What should the team investigate NEXT?

A.Check the ALB target group health check settings and ensure instances are healthy.
B.Examine the request headers in CloudFront logs to identify unusual patterns.
C.Review the CloudFront cache hit ratio and optimize caching strategies.
D.Check the ALB's idle timeout settings and compare with CloudFront origin timeout.
AnswerD

The correct root cause is a timeout mismatch: CloudFront waits for an origin response within its configured origin timeout (default 30 seconds), while the ALB has an idle timeout (default 60 seconds) that can close the connection to the backend if no bytes flow. If the ALB idle timeout is shorter than CloudFront’s timeout, the ALB can terminate a long-running request before the backend finishes, causing CloudFront to receive no response and return 504. Since the ALB logs show 200 for completed requests, the occasional slow requests are being cut off by this idle setting; aligning the ALB idle timeout to be greater than CloudFront’s origin timeout (or tuning backend latency) is the fix.

Why this answer

The ALB returns 200 OK for all requests, so the origin itself is not failing. However, HTTP 504 errors from CloudFront typically indicate that the origin (ALB) is not responding within CloudFront's timeout window. The ALB's idle timeout (default 60 seconds) can cause the ALB to close idle connections, while CloudFront's origin timeout (default 30 seconds) is separate.

If the ALB's idle timeout is shorter than the time CloudFront waits for a response, the ALB may close the connection before CloudFront receives the full response, leading to a 504. Option D directly addresses this mismatch.

Exam trap

The trap here is that candidates assume 504 errors always indicate an unhealthy origin, but the ALB logs show 200 OK, so they incorrectly focus on health checks or caching instead of the timeout mismatch between CloudFront and the ALB.

How to eliminate wrong answers

Option A is wrong because the ALB access logs show 200 OK for all requests, meaning the target group health check settings and instance health are not the issue—healthy instances would still produce 200 responses. Option B is wrong because examining request headers in CloudFront logs for unusual patterns would not explain a consistent 504 error when the ALB itself is responding successfully; the issue is at the transport layer, not the application layer. Option C is wrong because a low cache hit ratio would cause more origin requests but not 504 errors; optimizing caching strategies would reduce origin load but not fix a timeout mismatch between CloudFront and the ALB.

257
MCQhard

An organization uses AWS CodePipeline to orchestrate deployments to multiple environments (dev, test, prod). Each environment uses a different AWS account. The pipeline uses cross-account actions with IAM roles. Recently, the pipeline failed at the deploy stage for the prod account with the error 'Access Denied' when assuming the cross-account role. The role ARN is correct and the trust policy allows the pipeline's service role. What is the MOST likely cause?

A.The EC2 instances in the prod account do not have an appropriate instance profile.
B.The pipeline's service role lacks the `sts:AssumeRole` permission for the cross-account role.
C.The cross-account role's permissions boundary denies the deploy action.
D.The pipeline's service role does not have permission to perform the deploy action in the prod account.
AnswerB

For cross-account deployments, the pipeline service role in the source account must contain a policy that explicitly grants the `sts:AssumeRole` action on the ARN of the destination account's cross-account role. This is in addition to the trust policy on the cross-account role that allows the service role to assume it. Without this permission, CodePipeline's attempt to switch into the production account fails with a 403 AccessDenied at the AssumeRole step, which is exactly the described symptom. This is the root cause of the deployment failure.

Why this answer

The pipeline's service role must have an `sts:AssumeRole` permission on the cross-account role to perform the role assumption. Even if the trust policy on the cross-account role allows the pipeline's service role, the pipeline's service role itself needs an IAM policy granting `sts:AssumeRole` for the cross-account role ARN. Without this permission, the `AssumeRole` API call fails with 'Access Denied', which is the exact error described.

Exam trap

The trap here is that candidates often focus on the cross-account role's trust policy or permissions, forgetting that the pipeline's service role also needs explicit `sts:AssumeRole` permission, which is a separate IAM policy requirement.

How to eliminate wrong answers

Option A is wrong because the error occurs during the cross-account role assumption, not during an EC2 instance action; instance profiles are irrelevant to CodePipeline cross-account deployments. Option C is wrong because a permissions boundary on the cross-account role would limit the maximum permissions of the assumed role, but the error is 'Access Denied' at the assumption step, not during the deploy action itself. Option D is wrong because the pipeline's service role does not directly perform deploy actions in the prod account; it assumes the cross-account role, and the cross-account role's permissions govern the deploy action.

258
Multi-Selecthard

A DevOps team is using AWS CodeBuild to run integration tests against a test database. The database is an Amazon RDS instance in a private subnet. The CodeBuild project is configured to run in a VPC. Which THREE steps are required to allow CodeBuild to access the RDS instance?

Select 3 answers
A.Place the RDS instance in a public subnet with a public IP.
B.Ensure the security group attached to the RDS instance allows inbound traffic from the CodeBuild security group.
C.Attach a NAT gateway to the VPC so that CodeBuild can route to RDS.
D.Ensure the VPC's route tables have routes to allow traffic between CodeBuild subnets and RDS subnets.
E.Configure the CodeBuild project to use a VPC that has access to the RDS instance.
AnswersB, D, E

Security groups act as a virtual firewall at the instance level. By referencing the CodeBuild project's security group as a source in the RDS security group's inbound rule, you allow traffic specifically from the Elastic Network Interfaces (ENIs) that CodeBuild uses when it runs inside the VPC. This is the most direct and least-privileged way to permit the integration tests to reach RDS, because it avoids opening the database to CIDR ranges or the entire VPC. It also updates automatically if CodeBuild's IP addresses change, as long as the security group ID remains the same.

Why this answer

The security group attached to the RDS instance must explicitly allow inbound traffic from the security group associated with the CodeBuild project's elastic network interfaces. This is a fundamental network access control in AWS: security groups act as virtual firewalls, and without an inbound rule permitting traffic from the CodeBuild security group on the database port (e.g., 3306 for MySQL), the connection will be blocked regardless of other network configurations.

Exam trap

The trap here is that candidates often confuse the need for a NAT gateway (which is for internet access) with the requirement for internal VPC routing, or they mistakenly think that placing RDS in a public subnet is necessary for CodeBuild to reach it, when in fact private subnet communication via security groups and route tables is the correct approach.

259
Multi-Selectmedium

A company uses AWS Lambda with an Amazon DynamoDB trigger. Recently, the Lambda function started failing with 'ProvisionedThroughputExceededException' errors. The DevOps team needs to mitigate the issue. Which TWO actions should the team take? (Choose TWO.)

Select 2 answers
A.Increase the Lambda function's reserved concurrency
B.Disable DynamoDB Streams on the table
C.Enable DynamoDB Accelerator (DAX) for the table
D.Increase the DynamoDB table's write capacity
E.Reduce the batch size for the DynamoDB stream event source mapping
AnswersD, E

DynamoDB throttling occurs when write requests exceed the provisioned write capacity (WCUs) of the table. If the Lambda function writes processed items back to the same table, insufficient WCUs will cause ProvisionedThroughputExceededException, leading to retries and stream processing failures. Increasing the write capacity reduces throttling, allowing the stream-triggered writes to succeed and the function to make progress.

Why this answer

To mitigate 'ProvisionedThroughputExceededException' errors when a Lambda function is triggered by DynamoDB Streams, two actions are effective. Option D: Increase the DynamoDB table's write capacity to handle the write demand from the stream processing. Option E: Reduce the batch size for the DynamoDB stream event source mapping to lower the number of writes per invocation, reducing the chance of exceeding throughput.

Option A is wrong because Lambda reserved concurrency controls how many concurrent executions Lambda can run, but the issue is DynamoDB throttling, not Lambda capacity. Option B is wrong because disabling DynamoDB Streams would stop the trigger entirely, which is not a mitigation. Option C is wrong because DynamoDB Accelerator (DAX) is an in-memory cache for reads, not writes, and does not affect write throughput.

260
MCQeasy

A company is using AWS KMS to encrypt data at rest for S3 objects. The security team wants to rotate the KMS key annually. Which action should the team take to implement automatic key rotation?

A.Enable automatic key rotation when creating the KMS key
B.Create a new key manually each year and update the S3 bucket policy
C.Use AWS Certificate Manager (ACM) to rotate the KMS key
D.Use an AWS managed key, which rotates automatically every year
AnswerA

Enabling automatic key rotation on a customer-managed KMS key at creation schedules KMS to generate new cryptographic material every 365 days while preserving the same key ID, ARN, and aliases. S3 bucket policies and IAM policies still reference the same key, so no re-encryption of existing objects or policy updates are required. KMS automatically keeps all versions of the backing material available for decryption, so data encrypted before rotation remains decryptable without interruption.

Why this answer

AWS KMS supports automatic annual key rotation for customer managed keys (CMKs) when enabled at creation or via the key's rotation configuration. Once enabled, KMS automatically rotates the key material every 365 days, creating a new backing key while retaining the old one for decryption of previously encrypted data. This satisfies the security team's requirement without manual intervention.

Exam trap

The trap here is that candidates may confuse AWS managed keys (which rotate automatically but cannot be configured by the customer) with customer managed keys (which require explicit enabling of automatic rotation), leading them to select Option D incorrectly.

How to eliminate wrong answers

Option B is wrong because manually creating a new key each year and updating the S3 bucket policy is not automatic rotation and introduces operational overhead and potential misconfiguration. Option C is wrong because AWS Certificate Manager (ACM) is used for managing SSL/TLS certificates, not for rotating KMS keys; ACM has no integration with KMS key rotation. Option D is wrong because AWS managed keys (e.g., aws/s3) do rotate automatically, but they are not customer managed keys; the question implies the company is using a customer managed key (since they want to control rotation), and AWS managed keys cannot be configured for rotation by the customer.

261
MCQmedium

Refer to the exhibit. A DevOps engineer checks the CloudWatch alarm configuration and state. The alarm is in ALARM state for CPUUtilization averaging 90% over 5 minutes, but no notification was received. What is the most likely reason?

A.The SNS topic does not have any confirmed subscriptions.
B.The EC2 instance is stopped.
C.The alarm period is set to 300 seconds, which is too long.
D.The alarm has insufficient data to evaluate.
AnswerA

CloudWatch alarm actions publish notifications to the configured SNS topic, but the publish operation only succeeds if the topic has at least one active, confirmed subscriber endpoint, such as an email address that has clicked the confirmation link. When subscriptions are still in PendingConfirmation state or no subscription exists, the message is silently dropped while the alarm continues to show ALARM. Therefore, even though the alarm is firing correctly, no notification reaches the engineer because the delivery path lacks a confirmed subscriber. This is the root cause.

Why this answer

A CloudWatch alarm in ALARM state triggers its configured actions, typically publishing to an SNS topic. If no notification is received, the most likely cause is that the SNS topic has no confirmed subscriptions—meaning no endpoint (email, SMS, Lambda, etc.) has completed the subscription confirmation handshake. Without a confirmed subscription, SNS accepts the publish but delivers to no one, so the alarm fires silently.

This is a common operational oversight.

Exam trap

DOP-C02 often tests the assumption that an alarm in ALARM state automatically sends notifications, ignoring that SNS subscriptions must be confirmed and that alarm actions can be disabled or misconfigured.

How to eliminate wrong answers

Option B is wrong because if the EC2 instance were stopped, the CPUUtilization metric would either stop being published or drop to zero, not average 90%; the alarm is already in ALARM state, so the instance is running and consuming CPU. Option C is wrong because a 300-second period is a standard and valid evaluation window; it does not prevent notifications—it only affects how quickly the alarm evaluates. Option D is wrong because the alarm is explicitly in ALARM state, meaning it has sufficient data to evaluate and has breached the threshold; insufficient data would result in INSUFFICIENT_DATA state, not ALARM.

262
MCQmedium

A team uses AWS CloudFormation to manage infrastructure. They want to automatically update the stack when a new version of a Docker image is pushed to Amazon ECR. Which approach should they use?

A.Configure an Amazon EventBridge rule to detect ECR image pushes and invoke an AWS Lambda function that calls 'UpdateStack' with the new image URI.
B.Create a CodeBuild project that triggers on ECR push, and have the build execute an 'aws cloudformation update-stack' command.
C.Use AWS CodeDeploy with a trigger on ECR push to deploy the new image to a target group, and have the target group update the stack.
D.Set up an Amazon SNS topic subscribed to ECR image push events, and have the SNS topic send a notification to an AWS CloudFormation stack update endpoint.
AnswerA

This is correct because Amazon EventBridge natively captures ECR image push events (e.g., the 'ECR Image Action' event type) and can directly target an AWS Lambda function. Lambda can then call CloudFormation's UpdateStack API, passing the new image URI as a parameter to the stack, thereby updating the infrastructure in an event-driven, serverless manner. This pattern is simple, requires no persistent compute, and integrates seamlessly with CloudFormation's parameter-based updates.

Why this answer

Amazon EventBridge can detect ECR image push events and trigger an AWS Lambda function. The Lambda function can then call the CloudFormation UpdateStack API with the new image URI, automating the stack update. This approach is serverless, event-driven, and directly integrates with CloudFormation, making it the most efficient and least operational overhead solution.

Exam trap

DOP-C02 often tests the integration between AWS services for automation; candidates might incorrectly assume that SNS can directly trigger CloudFormation updates or that CodeBuild/CodeDeploy can natively respond to ECR events, overlooking the need for EventBridge and Lambda.

How to eliminate wrong answers

Option B is wrong because CodeBuild is not designed to trigger on ECR push events natively; it would require additional configuration like CloudWatch Events, and it adds unnecessary complexity. Option C is wrong because CodeDeploy is for deploying applications to EC2/on-premises or ECS, not for updating CloudFormation stacks; it does not have a direct mechanism to update a stack. Option D is wrong because SNS cannot directly invoke a CloudFormation stack update; CloudFormation does not provide an SNS endpoint for stack updates, and this approach would require additional compute to process the notification.

263
MCQhard

An organization uses AWS CloudFormation to manage infrastructure across multiple accounts using AWS Organizations. They want to enforce that all S3 buckets are encrypted with SSE-S3. A DevOps engineer creates a service control policy (SCP) to deny the creation of any S3 bucket without encryption. However, CloudFormation stack creation fails with an access denied error even when the template includes encryption. What is the most likely cause?

A.The CloudFormation template specifies SSE-KMS encryption, which is not allowed by the SCP.
B.The SCP is denying the s3:PutBucketPublicAccessBlock action, which is required for all bucket creation requests.
C.The SCP is incorrectly scoped to the management account instead of the member accounts.
D.The CloudFormation service role does not have permissions to create buckets in the target account.
AnswerA

Correct. SCPs are service control policies that set maximum permissions for all IAM principals in an account, and they can include conditions that deny S3 bucket creation when SSE-KMS is specified. If the organization's SCP only allows SSE-S3 encryption (or denies kms:GenerateDataKey or kms:CreateGrant), then any CloudFormation template that specifies SSE-KMS encryption for the bucket will be rejected with an AccessDenied error, regardless of the IAM service role or the target bucket configuration. This perfectly matches the scenario where the error occurs only when certain encryption settings are present in the template.

Why this answer

The SCP denies the creation of S3 buckets without encryption, but it specifically allows only SSE-S3 encryption. If the CloudFormation template specifies SSE-KMS encryption, the SCP will deny the request, causing an access denied error—even though encryption is present. This mismatch between the encryption type required by the SCP (SSE-S3) and what the template requests (SSE-KMS) is the most likely cause of the failure.

Exam trap

The trap is assuming that any encryption (SSE-S3 or SSE-KMS) satisfies the SCP requirement. However, SCPs can be very specific; if the SCP only allows SSE-S3, then using SSE-KMS will be denied. Candidates may overlook the distinction between encryption types.

How to eliminate wrong answers

Option A is wrong because SSE-KMS is a form of encryption; if the SCP denies bucket creation without encryption, specifying SSE-KMS would satisfy the encryption requirement, so it would not cause an access denied error. Option B is wrong because s3:PutBucketPublicAccessBlock is not required for all bucket creation requests; it is an optional action to block public access, and denying it would not prevent bucket creation—only the ability to set public access settings. Option D is wrong because the CloudFormation service role's permissions are separate from SCPs; if the role lacks permissions, the error would be an authorization failure, but the question explicitly states the SCP is the cause, and SCPs cannot be overridden by IAM roles—they act as a boundary, so the role's permissions are irrelevant if the SCP denies the action.

264
Drag & Dropmedium

Drag and drop the steps to implement a blue/green deployment using AWS CodeDeploy.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

First create the application and deployment group, then configure blue/green settings, then deploy, then validate, then reroute traffic.

265
MCQmedium

A DevOps engineer receives an alert that an EC2 instance has been compromised. The instance is part of an Auto Scaling group. What is the first step the engineer should take to isolate the instance?

A.Create a snapshot of the instance's root volume
B.Detach the instance from the Auto Scaling group and remove it from the load balancer
C.Create an AMI of the instance for analysis
D.Terminate the instance immediately
AnswerB

Detaching the instance from the Auto Scaling group and removing it from the load balancer is the correct first step because it immediately stops new traffic from reaching the instance while also preventing the ASG from automatically replacing or terminating it. This preserves the running state for forensic collection, including memory and volatile data, and allows you to investigate safely without the instance being scaled away or continuing to affect production traffic. The instance stays alive but is decoupled from both the horizontal scaling and the request path.

Why this answer

The first step to isolate a compromised EC2 instance in an Auto Scaling group is to detach it from the Auto Scaling group and remove it from the load balancer. This stops all incoming traffic to the instance, preventing further damage or data exfiltration while preserving the instance for forensic analysis. Option A (snapshot) is useful for preserving evidence but does not isolate the instance.

Option C (AMI) similarly does not provide immediate isolation. Option D (terminate) may destroy evidence and should only be done after investigation.

266
MCQmedium

An organization uses AWS CodeDeploy to deploy applications to Amazon EC2 instances. The deployment is failing consistently with the error 'ScriptMissing' for the AppSpec lifecycle hook 'ApplicationStop'. The scripts are located in the /opt/scripts directory on the instances. What is the most likely cause of this error?

A.The ApplicationStop hook is not defined in the AppSpec file.
B.The CodeDeploy agent is not the latest version.
C.The AppSpec file specifies a path to the script that does not exist on the instance.
D.The scripts have incorrect file permissions.
AnswerC

The AppSpec hooks section contains a location field that must point to a script actually present in the extracted application revision. During a lifecycle event, the agent resolves that path against the deployment archive's directory structure; if the file is not there—perhaps due to a typo, an absolute path instead of a relative one, or omitted from the bundle—the operating system returns ENOENT and CodeDeploy reports ScriptMissing. Fixing the AppSpec path to match the bundled script resolves the deployment.

Why this answer

The 'ScriptMissing' error in AWS CodeDeploy indicates that the deployment lifecycle hook (in this case, 'ApplicationStop') cannot find the script file at the path specified in the AppSpec file. Since the scripts are located in /opt/scripts, the most likely cause is that the AppSpec file's 'location' field for the ApplicationStop hook points to a path that does not match the actual file location on the instance, or the script file itself is missing from that directory.

Exam trap

The trap here is that candidates often confuse 'ScriptMissing' with permission errors or agent issues, but AWS specifically uses 'ScriptMissing' to indicate the file path is incorrect or the file is absent, not that the file exists but cannot be executed.

How to eliminate wrong answers

Option A is wrong because if the ApplicationStop hook were not defined in the AppSpec file, CodeDeploy would not attempt to run it and would not produce a 'ScriptMissing' error; instead, it would simply skip that hook. Option B is wrong because the CodeDeploy agent version does not affect script path resolution; 'ScriptMissing' is a file-not-found error, not an agent compatibility issue. Option D is wrong because incorrect file permissions would cause a 'ScriptFailed' or permission-denied error, not a 'ScriptMissing' error, which specifically indicates the script file cannot be located at the given path.

267
MCQeasy

A company uses Amazon CloudFront to serve static content from an S3 bucket. Users report that they see outdated content even after the engineer has updated the files in the S3 bucket. What should the engineer do to ensure users see the latest content?

A.Create an invalidation for the updated file paths.
B.Change the S3 bucket policy to allow public access.
C.Reduce the TTL for the CloudFront distribution.
D.Delete and recreate the CloudFront distribution.
AnswerA

CloudFront caches objects at edge locations until their TTL expires, so after you update files in Amazon S3 the edge locations continue serving the old copies. An invalidation request explicitly removes the specified file paths from every CloudFront edge cache, forcing subsequent requests to fetch the latest version from the S3 origin. You can target an individual file with /path/file.js or use a wildcard like /images/* to clear a whole directory. This is the intended, immediate mechanism for propagating content updates without waiting for natural cache expiry.

Why this answer

Creating a CloudFront invalidation for the specific file paths forces the edge locations to fetch the updated content from the S3 origin immediately, ensuring users see the latest files. Option B is incorrect because changing the bucket policy to allow public access does not affect CloudFront's cache; it only controls direct access to the bucket. Option C is incorrect because reducing the TTL affects how long new content is cached but does not clear already-cached outdated content.

Option D is incorrect because deleting and recreating the distribution is an overly disruptive solution; a simple invalidation suffices.

268
MCQmedium

Refer to the exhibit. A DevOps engineer runs the command to get the pipeline definition. The pipeline has a source stage from an S3 bucket and a build stage with CodeBuild. The CodeBuild project is configured to output artifacts to a specific S3 bucket. However, the pipeline fails at the build stage with an error: 'Artifact 'BuildArtifact' is not found'. What is the most likely cause?

A.The source stage is using CodeCommit instead of S3.
B.The source artifact is not being passed to the build stage.
C.The IAM role for CodePipeline does not have permissions to read from the S3 bucket.
D.The CodeBuild project is not configured to output the expected artifact named 'BuildArtifact'.
AnswerD

CodeBuild only produces output artifacts if the project (or buildspec `artifacts` section) is configured to do so with a matching name. The pipeline's Build action expects an output artifact literally named 'BuildArtifact'; if the CodeBuild project defines a different artifact name, or omits artifact output entirely, the build may succeed but the pipeline's build stage will fail when resolving the expected artifact. This mismatch is the most common reason for a Build stage failure when the build logs show a successful build.

Why this answer

The error 'Artifact 'BuildArtifact' is not found' indicates that CodePipeline expects an artifact named 'BuildArtifact' from the CodeBuild project, but the project's output artifact configuration does not include that name. In CodePipeline, the build stage must produce an artifact with the exact name specified in the pipeline definition; if the CodeBuild project's artifacts section (in buildspec.yml or console) omits or misnames it, the pipeline fails at that stage.

Exam trap

The trap here is that candidates often confuse permission errors (IAM) with artifact naming mismatches, or assume the source stage is misconfigured, when the real issue is a simple name mismatch between the pipeline's expected output artifact and the CodeBuild project's actual output artifact.

How to eliminate wrong answers

Option A is wrong because the source stage is explicitly configured with an S3 bucket (as per the exhibit), and using CodeCommit would not cause a 'BuildArtifact not found' error—it would affect source retrieval, not artifact naming. Option B is wrong because the source artifact is passed to the build stage by default via the pipeline's artifact store; the error is about the output artifact from CodeBuild, not the input. Option C is wrong because insufficient IAM permissions to read from the S3 bucket would produce an access denied error (e.g., 'AccessDenied' or '403 Forbidden'), not a 'not found' error for a named artifact.

269
MCQmedium

A DevOps engineer runs the above command and sees that instance i-0abcd1234efgh5678 is unhealthy with reason 'Target.Timeout'. The instance is running and the application on port 80 responds to curl from the instance itself. What is the MOST likely cause?

A.The ALB health check interval is set too high.
B.The web server process is not running on the instance.
C.The health check path returns a 404 status code.
D.The security group for the instance does not allow inbound traffic from the ALB on port 80.
AnswerD

A timeout typically indicates a network connectivity issue between ALB and instance.

Why this answer

The 'Target.Timeout' health check reason indicates that the load balancer's health check request timed out, meaning it could not establish a connection or receive a response within the timeout period. Since the instance responds to curl locally, the application is running, but the ALB cannot reach it. The most likely cause is that the instance's security group does not allow inbound traffic from the ALB on port 80, blocking the health check requests.

Exam trap

DOP-C02 often tests troubleshooting of load balancer health checks, and candidates might overlook security group rules, assuming the application is at fault because local curl works.

How to eliminate wrong answers

Option A is wrong because a high health check interval would delay checks but not cause timeouts; the interval is the time between checks, not the timeout duration. Option B is wrong because the application responds to curl, so the web server process is running. Option C is wrong because a 404 status code would result in a 'Target.ResponseCodeMismatch' reason, not a timeout.

270
MCQmedium

An IAM policy is attached to an S3 bucket to allow access from a specific VPC CIDR range. However, users from the VPC are receiving 'Access Denied' errors when trying to access objects in the bucket. What is the MOST likely reason?

A.The users are assuming an IAM role that does not have permission to access S3
B.The condition key 'aws:SourceIp' evaluates the public IP address, but the VPC uses private IP addresses
C.The policy should use 'aws:sourceVpce' instead of 'aws:SourceIp' to restrict access to a VPC endpoint
D.The bucket policy requires HTTPS and the requests are using HTTP
AnswerB

The aws:SourceIp condition key in a bucket policy compares the source IP address of the requester as seen by S3, which is the public IP address after any network address translation (NAT) or internet gateway processing. In a VPC, instances typically use private RFC 1918 IP addresses, which are not exposed to S3 when traffic traverses a NAT gateway or the internet. Thus, a policy specifying private IP ranges will never match, causing access to be denied for legitimate VPC clients.

Why this answer

The 'aws:SourceIp' condition key evaluates the source IP address of the request as seen by AWS, which is always a public IP address. When requests originate from within a VPC (e.g., from an EC2 instance), the source IP seen by S3 is the public IP of the instance or its NAT device, not the VPC's private CIDR range. Since the policy specifies a private VPC CIDR range using 'aws:SourceIp', it never matches the public source IP, resulting in 'Access Denied' errors.

Exam trap

The trap here is that candidates often confuse 'aws:SourceIp' with being able to match private IP ranges, not realizing that AWS services always see the public IP address of the request, even if the request originates from within a VPC.

How to eliminate wrong answers

Option A is wrong because the question states that the IAM policy is attached to the S3 bucket, and the error occurs despite that policy; if the users were assuming an IAM role without S3 permissions, the error would be consistent regardless of the bucket policy, but the issue here is specifically tied to the VPC CIDR condition. Option C is wrong because 'aws:sourceVpce' is used to restrict access to a specific VPC endpoint (interface or gateway), not to a VPC CIDR range; using it would require the request to originate from a VPC endpoint, which is not the scenario described. Option D is wrong because the error message 'Access Denied' is distinct from a '403 Forbidden' or '400 Bad Request' that would occur if HTTPS were required and HTTP was used; S3 bucket policies with 'aws:SecureTransport' condition would explicitly deny HTTP, but the question does not mention HTTPS enforcement.

271
Multi-Selecthard

Which TWO are correct about using AWS CloudFormation to manage infrastructure across multiple AWS accounts? (Select TWO.)

Select 2 answers
A.You can use AWS Organizations to centrally manage accounts and use StackSets with trusted access.
B.CloudFormation can automatically create new AWS accounts using a template.
C.Nested stacks can be used to deploy resources in different accounts from a single template.
D.AWS CloudFormation StackSets can deploy stacks across multiple accounts.
E.You can use cross-stack references to share resources between accounts.
AnswersA, D

Enabling trusted access for AWS CloudFormation in AWS Organizations allows StackSets to use the organization's structure (accounts and OUs) to automatically determine the set of target accounts. With service-managed permissions, any account added to the organization or an OU is automatically included in deployments, removing the need to manually maintain account lists. This centralization is a best practice for multi-account infrastructure management.

Why this answer

AWS Organizations can centrally manage accounts, and StackSets can be enabled with trusted access to deploy stacks across accounts. Option D is correct because AWS CloudFormation StackSets allow deploying stacks across multiple accounts. Option B is incorrect because CloudFormation cannot automatically create new AWS accounts.

Option C is incorrect because nested stacks operate within a single stack and cannot deploy resources across different accounts from a single template. Option E is incorrect because cross-stack references only work within the same account and region.

272
MCQeasy

A company wants to encrypt data at rest in Amazon S3 using server-side encryption with AWS Key Management Service (SSE-KMS) and enforce that all new objects are encrypted. Which bucket policy statement should be added?

A.{"Effect":"Deny","Action":"s3:PutObject","Resource":"arn:aws:s3:::bucket/*","Condition":{"StringNotEquals":{"s3:ServerSideEncryption":"awskms"}}}
B.{"Effect":"Deny","Action":"s3:PutObject","Resource":"arn:aws:s3:::bucket/*","Condition":{"StringEquals":{"s3:x-amz-server-side-encryption-aws-kms-key-id":"arn:aws:kms:us-east-1:123456789012:key/1234-5678-9012"}}}
C.{"Effect":"Deny","Action":"s3:PutObject","Resource":"arn:aws:s3:::bucket/*","Condition":{"StringNotEquals":{"s3:x-amz-server-side-encryption-aws-kms-key-id":"arn:aws:kms:us-east-1:123456789012:key/1234-5678-9012"}}}
D.{"Effect":"Deny","Action":"s3:PutObject","Resource":"arn:aws:s3:::bucket/*","Condition":{"StringNotEquals":{"s3:x-amz-server-side-encryption":"AES256"}}}
AnswerC

Using `StringNotEquals` on the KMS key ID correctly denies every `PutObject` request whose encryption key is not the specified ARN. This ensures that only objects encrypted with the designated customer-managed KMS key can be written to the bucket, enforcing SSE-KMS with that exact key.

Why this answer

Option C uses the valid condition key `s3:x-amz-server-side-encryption-aws-kms-key-id` with `StringNotEquals` to deny `s3:PutObject` if the request does not include the specified KMS key ID. This enforces that all new objects are encrypted with that specific KMS key. Option A is incorrect because `s3:ServerSideEncryption` is not a valid condition key for S3.

Option B is incorrect because it uses `StringEquals` instead of `StringNotEquals`, causing it to deny requests that include the specified KMS key, which prevents the use of SSE-KMS. Option D is incorrect because it uses `s3:x-amz-server-side-encryption` with value `AES256`, which enforces SSE-S3, not SSE-KMS.

Exam trap

Candidates might confuse the condition keys or the string comparison operators. `s3:x-amz-server-side-encryption` checks the encryption algorithm (e.g., AES256 or aws:kms), while `s3:x-amz-server-side-encryption-aws-kms-key-id` checks the specific KMS key ID. Also, `StringNotEquals` is used to deny if a required condition is not met, whereas `StringEquals` with deny can accidentally block valid requests.

273
MCQeasy

A DevOps team uses AWS CodeBuild to compile code and run unit tests. The team notices that builds are failing with a timeout error after 60 minutes. What is the most likely cause and solution?

A.The buildspec has syntax errors; validate the YAML.
B.The build environment is too small; increase compute type.
C.The source code repository is too large; use shallow clone.
D.The build timeout limit is exceeded; increase the timeout in CodeBuild project settings.
AnswerD

CodeBuild projects have a configurable 'Timeout' property (default 60 minutes) that caps the total time from build start to the last phase, including download, install, build, and post-build. When a build runs past this limit, CodeBuild forcibly terminates it and records a 'BUILD_TIMEOUT' error, which matches the scenario described. To resolve it, the DevOps team can increase the timeout in the CodeBuild project settings (via console, AWS CLI, or CloudFormation), setting it up to 480 minutes (8 hours) for standard builds. This directly addresses the root cause rather than masking a resource or syntax problem.

Why this answer

The default build timeout for AWS CodeBuild is 60 minutes. When a build runs longer than this limit, it fails with a timeout error. The correct solution is to increase the timeout value in the CodeBuild project settings, which can be set up to a maximum of 8 hours (480 minutes) for on-demand compute environments.

Exam trap

The trap here is that candidates may confuse a timeout error with resource constraints (like compute size) or source code issues, but the specific 60-minute failure point is a direct indicator of the default build timeout being exceeded.

How to eliminate wrong answers

Option A is wrong because syntax errors in the buildspec YAML would cause immediate build failures with parse errors, not a timeout after exactly 60 minutes. Option B is wrong because an undersized compute environment would manifest as slow performance or resource exhaustion errors (e.g., memory or disk), not a consistent timeout at the 60-minute mark. Option C is wrong because a large source repository would cause slow clone operations, but CodeBuild's default clone timeout is 10 minutes, and the overall build timeout is independent of clone time; shallow clone reduces clone time but does not address a 60-minute timeout on the entire build process.

274
MCQeasy

An application running on Amazon EC2 instances behind an Application Load Balancer (ALB) is experiencing intermittent 503 errors. The target group health checks are failing. The DevOps engineer checks the instance logs and finds that the application is running but taking longer than 30 seconds to respond. What is the MOST likely cause?

A.The Auto Scaling group's scaling policy is too aggressive, causing frequent instance replacements.
B.The security group for the ALB does not allow inbound traffic from the internet.
C.The health check timeout is set too low, causing the ALB to mark instances unhealthy.
D.The EC2 instances are running out of memory and the application is crashing.
AnswerC

A health check timeout set too low can cause the ALB to mark otherwise functional instances as unhealthy when the application's response time occasionally exceeds the timeout. The ALB health check settings include an interval, timeout, and unhealthy threshold; if a slow application misses the timeout a few consecutive times, the target is deregistered and the ALB returns 503 Service Unavailable when no healthy targets remain. This matches the symptom of intermittent errors under load, as the application may respond normally at times but exceed the timeout during traffic spikes.

Why this answer

The most likely cause of intermittent 503 errors and failing health checks is that the health check timeout is set too low. The application takes longer than 30 seconds to respond, but if the health check timeout is set to a value less than the application response time, the ALB will mark the instance as unhealthy, leading to 503 errors. Adjusting the health check timeout to accommodate the application's response time would resolve the issue.

Exam trap

DOP-C02 often tests the misconception that 503 errors are always due to security groups or instance failures, overlooking health check timeout misconfigurations.

How to eliminate wrong answers

Option A is wrong because an aggressive scaling policy would cause instances to be replaced frequently, but that would not directly cause health checks to fail due to response time; it might cause other issues. Option B is wrong because if the security group did not allow inbound traffic, the ALB would not be able to reach the instances at all, resulting in consistent failures, not intermittent ones. Option D is wrong because if instances were running out of memory and crashing, the application would not be running, but the logs show it is running and taking longer than 30 seconds.

275
MCQhard

Your company uses AWS CodePipeline to automate the deployment of a critical web application. The pipeline consists of a source stage (CodeCommit), a build stage (CodeBuild), and a deploy stage (CodeDeploy) that deploys to an Auto Scaling group of EC2 instances running Amazon Linux 2. The deployment strategy is 'AllAtOnce'. Recently, the team noticed that during deployments, the application becomes completely unavailable for a few minutes until the new instances are registered with the load balancer. The business requires zero downtime during deployments. You need to modify the deployment process to achieve zero downtime while minimizing cost and complexity. The Auto Scaling group currently has a minimum of 2 instances and a maximum of 4 instances. The application is stateless and sessions are stored in ElastiCache. Which solution should you implement?

A.Create a second Auto Scaling group, deploy to it, and then update Route 53 to point to the new group.
B.Change the deployment configuration to 'HalfAtATime' to update half the instances at a time.
C.Use CodeDeploy's blue/green deployment with an Application Load Balancer. Create a new Auto Scaling group for the green environment, deploy to it, and then shift traffic.
D.Increase the Auto Scaling group's minimum size to 4 so there are always extra instances.
AnswerC

CodeDeploy's blue/green deployment with an Application Load Balancer is the correct solution because it provisions a completely new Auto Scaling group (the green environment) running the new revision and registers it with a new or existing target group. The ALB only shifts traffic to the green target group after all instances pass health checks, while the blue instances continue serving requests during the entire setup and validation. After a configurable wait time, the original blue ASG is terminated, ensuring no point where users have no service. This architecture provides true zero downtime because the old fleet remains fully operational until the new fleet is verified.

Why this answer

CodeDeploy blue/green deployments with an ALB allow traffic to be shifted from the original (blue) Auto Scaling group to a newly provisioned green Auto Scaling group only after the new instances pass health checks, so the application never becomes unavailable. Because the app is stateless and sessions live in ElastiCache, no session affinity or state migration is required, making this both low-risk and low-complexity. This directly satisfies the zero-downtime requirement without over-provisioning capacity.

Exam trap

DOP-C02 often tests the misconception that changing the deployment configuration (e.g., HalfAtATime) or adding instances is sufficient for zero downtime, when in fact only blue/green (or a rolling strategy with proper connection draining) with a load balancer achieves it.

How to eliminate wrong answers

Option A is wrong because manually creating a second Auto Scaling group and repointing Route 53 DNS is a hand-rolled blue/green that introduces DNS TTL delays, lacks automated traffic shifting/rollback, and adds significant operational complexity compared to CodeDeploy's native blue/green. Option B is wrong because 'HalfAtATime' is an in-place deployment that terminates and replaces instances within the same Auto Scaling group; during replacement the remaining instances may still be deregistering or warming up, and with a min of 2 instances, taking half offline can still cause availability gaps and does not guarantee zero downtime. Option D is wrong because simply raising the minimum size to 4 does not change the AllAtOnce deployment behavior — CodeDeploy still replaces all instances simultaneously, so the app still goes down; it only adds cost without solving the availability problem.

276
Multi-Selectmedium

A security engineer is designing a secure VPC architecture for a web application. The application must be isolated from the internet and only accessible through a load balancer. Which TWO actions should the engineer take?

Select 2 answers
A.Place the EC2 instances in a private subnet with no internet gateway attachment.
B.Attach an Internet Gateway to the VPC and route the private subnet to it.
C.Configure a network ACL on the private subnet to allow inbound traffic on all ephemeral ports.
D.Configure the security group for the EC2 instances to allow traffic only from the ALB's security group.
E.Set up an AWS Direct Connect connection for the instances to access the internet.
AnswersA, D

Placing the EC2 instances in a private subnet with no internet gateway (IGW) route ensures they cannot receive unsolicited inbound traffic from the internet. The ALB, deployed in a public subnet, is the only intentional ingress point; it forwards traffic to the instances over the private VPC network, which does not require an IGW. This design prevents direct exposure of instance IPs and relies on the security group to restrict traffic to just the ALB's interface.

Why this answer

Placing EC2 instances in a private subnet without an internet gateway ensures they have no direct path to the internet, meeting the isolation requirement. This forces all traffic to and from the instances to go through the load balancer, which is the only entry point for the application.

Exam trap

The trap here is that candidates often confuse the need for a network ACL to allow ephemeral ports (Option C) as a necessary step for inbound traffic from the ALB, but security groups handle stateful filtering and the ALB's security group is the correct source, while network ACLs are stateless and require explicit rules for both inbound and outbound traffic, which is not the primary action for isolation.

277
MCQhard

An organization wants to ensure that all objects stored in the S3 bucket are encrypted at rest using server-side encryption with S3 managed keys (SSE-S3). The bucket policy above is intended to enforce this. However, a user reported that they can still upload unencrypted objects. What is the MOST likely reason?

A.The condition is applied to the GetObject action, not the PutObject action.
B.The bucket policy is not attached to the bucket because of a circular dependency.
C.The bucket policy does not apply to objects uploaded by the root user.
D.The condition should use 's3:x-amz-server-side-encryption-aws-kms-key-id' instead.
AnswerA

The bucket policy's condition key is tied to the s3:GetObject action, which only governs read requests. To enforce encryption during uploads, the policy must target s3:PutObject, because that is the action that accepts the x-amz-server-side-encryption header and where a missing encryption header should be denied. Placing the condition on GetObject only denies reading unencrypted objects, leaving uploads allowed without encryption.

Why this answer

The bucket policy condition `s3:x-amz-server-side-encryption` is applied to the `s3:GetObject` action instead of `s3:PutObject`. This means the policy only checks encryption headers when reading objects, not when uploading them. To enforce encryption at upload time, the condition must be attached to the `s3:PutObject` action, which is the operation that accepts the `x-amz-server-side-encryption` header.

Exam trap

The trap here is that candidates often focus on the condition key or value syntax and overlook the action to which the condition is attached, assuming any encryption-related condition will automatically apply to uploads.

How to eliminate wrong answers

Option B is wrong because a circular dependency would prevent the bucket policy from being attached at all, but the user can still upload objects, so the policy is attached and functional. Option C is wrong because bucket policies apply to all principals, including the root user, unless explicitly excluded; the root user is not exempt from policy conditions. Option D is wrong because `s3:x-amz-server-side-encryption-aws-kms-key-id` is used for SSE-KMS, not SSE-S3; the question specifies SSE-S3, which uses the `AES256` value with the `s3:x-amz-server-side-encryption` condition key.

278
MCQhard

A company is designing a disaster recovery (DR) strategy for a stateless web application deployed on Amazon ECS with Fargate. The application is fronted by an Application Load Balancer (ALB) and uses Amazon ElastiCache for Redis for session state. The primary region is us-east-1. The DR plan requires a Recovery Point Objective (RPO) of 15 minutes and a Recovery Time Objective (RTO) of 30 minutes. Which solution meets these requirements with the LEAST operational overhead?

A.Deploy an ALB with a warm standby ECS service in us-west-2. Use Route 53 health checks to route traffic to the secondary region if primary fails. Use ElastiCache Global Datastore for Redis to replicate data across regions.
B.Deploy an Active-Active configuration across two AWS regions using Route 53 latency routing. Use ElastiCache for Redis Global Datastore with multi-region writes.
C.Deploy a Pilot Light environment in us-west-2 with a scaled-down ECS service and Redis cluster. Use Route 53 DNS failover. On disaster, scale up the ECS service and promote the Redis cluster.
D.Use Amazon ECS with Fargate in us-east-1 only, and schedule daily snapshots of ElastiCache for Redis. In case of disaster, restore the snapshot in a new region and update DNS.
AnswerA

A warm standby architecture places a fully functional but possibly smaller ECS service behind an ALB in us-west-2, and Route 53 health checks automatically shift user traffic when us-east-1's ALB or service is unhealthy. ElastiCache Global Datastore for Redis maintains a cross-region replica with sub-second replication lag, so the secondary can serve reads and writes after promote without a point-in-time restore. This combination keeps both RPO and RTO under 30 minutes because failover is automated and data is already in the secondary region.

Why this answer

It uses ElastiCache Global Datastore for Redis, which provides cross-region replication with an RPO of seconds (well within 15 minutes) and automatic failover, minimizing operational overhead. The warm standby ECS service in us-west-2 with Route 53 health checks allows traffic to be redirected within the 30-minute RTO without manual intervention, as the ALB and ECS service are pre-provisioned.

Exam trap

The trap here is that candidates may confuse Pilot Light (Option C) as lower overhead, but it requires manual scaling and promotion steps, whereas a warm standby with Global Datastore automates failover, making it the least operational overhead for the given RPO/RTO.

How to eliminate wrong answers

Option B is wrong because an Active-Active configuration with multi-region writes for ElastiCache Global Datastore is not supported; Global Datastore only supports active-passive (one primary, one replica) to avoid write conflicts. Option C is wrong because a Pilot Light approach requires manual scaling of the ECS service and promoting the Redis cluster on disaster, which adds operational overhead and risks exceeding the 30-minute RTO due to provisioning delays. Option D is wrong because daily snapshots of ElastiCache cannot achieve a 15-minute RPO (snapshots are at most daily), and restoring a snapshot in a new region plus updating DNS would likely exceed the 30-minute RTO due to manual steps and data transfer time.

279
MCQmedium

A company uses AWS CloudTrail to monitor API activity. The security team notices that an IAM user 'dev-user' deleted an S3 bucket. They need to quickly identify the source IP address of the delete request. Which CloudTrail feature should they use to find this information?

A.Use CloudTrail Lake to query the event and extract the IP address from the userIdentity field.
B.Check S3 server access logs for the bucket deletion event.
C.Enable CloudTrail Insights to analyze unusual activity.
D.Search the CloudTrail event history for the delete event and review the sourceIPAddress field.
AnswerD

CloudTrail event history retains management events for the last 90 days, and each event record includes the sourceIPAddress field, which identifies the IP address from which the API call was made. Searching event history for the DeleteBucket event, or a similar deletion event, and expanding the event details will show the sourceIPAddress in the raw event record. This is the direct and correct way to retrieve the requester's IP for a specific CloudTrail event.

Why this answer

CloudTrail Event History retains 90 days of management events and each event record includes a sourceIPAddress field that captures the IP address from which the API call was made. Searching Event History for the DeleteBucket event and inspecting sourceIPAddress is the fastest way to identify the source IP without additional setup.

Exam trap

DOP-C02 often tests whether candidates know that S3 server access logs do not capture control-plane API calls like DeleteBucket — only CloudTrail does.

How to eliminate wrong answers

Option A is wrong because CloudTrail Lake is a paid, query-based service for long-term retention and SQL analysis; while it can return the IP, it is overkill and not the 'quick' method for a recent event. Option B is wrong because S3 server access logs record bucket-level access requests but do not capture the DeleteBucket API call itself (which is a control-plane action logged by CloudTrail, not S3 access logs). Option C is wrong because CloudTrail Insights detects anomalous API call rates and error rates — it does not provide per-event source IP details.

280
MCQeasy

A company uses Ansible for configuration management on EC2 instances. They want to ensure that only instances with a specific tag (Environment: Production) are targeted by their playbooks. What is the best way to achieve this?

A.Add a 'when' condition in the playbook to check the instance tag at runtime.
B.Maintain a static inventory file listing only Production instances.
C.Use the AWS EC2 dynamic inventory plugin to filter instances based on tags.
D.Use the ec2_tag module to assign the tag to instances.
AnswerC

The EC2 dynamic inventory plugin queries the AWS API and groups hosts by their tags, so `Environment: Production` becomes a selectable group or filterable host pattern. This satisfies the requirement to target only tagged production instances without maintaining a static inventory file that would drift as instances launch or terminate.

Why this answer

The best approach because the AWS EC2 dynamic inventory plugin allows filtering instances by tags (e.g., 'Environment: Production') at runtime, ensuring only tagged instances are targeted. Option A (adding a 'when' condition) is less efficient as it connects to all instances first. Option B (static inventory) requires manual updates and is not dynamic.

Option D (ec2_tag module) is for assigning tags, not selecting instances.

281
MCQmedium

A DevOps engineer is troubleshooting a production AWS Lambda function that occasionally times out. The function has a timeout of 30 seconds and uses a synchronous invocation. The engineer wants to capture invocation logs to identify the cause. Which approach will provide the MOST detailed diagnostic information?

A.Enable AWS CloudTrail data events for Lambda.
B.Create a CloudWatch dashboard with function duration metrics.
C.Add more logging statements to the function code and check CloudWatch Logs.
D.Enable AWS X-Ray tracing on the Lambda function.
AnswerD

AWS X-Ray tracing on the Lambda function provides an end-to-end view of the invocation, showing each subsegment's duration, including the time spent in external HTTP calls, AWS SDK operations, and service integrations. It automatically captures the trace for every invocation without requiring code changes (unless you need custom subsegments), and the trace timeline reveals exactly which downstream call or code block is consuming the most time. This makes X-Ray the correct tool to diagnose performance bottlenecks in a production Lambda, as it gives the per-invocation, subsegment-level detail that the other options lack.

Why this answer

AWS X-Ray provides end-to-end tracing for Lambda functions, capturing detailed timing information for each invocation, including subsegments for downstream calls, function initialization, and execution phases. This allows the engineer to pinpoint exactly where time is being spent, which is essential for diagnosing intermittent timeouts in synchronous invocations.

Exam trap

The trap here is that candidates often confuse CloudWatch Logs (which show custom log output) with X-Ray tracing (which provides automatic, detailed timing of every subcomponent), leading them to choose option C instead of the more diagnostic X-Ray approach.

How to eliminate wrong answers

Option A is wrong because CloudTrail data events for Lambda only record API calls (e.g., Invoke, UpdateFunctionConfiguration) and do not capture function execution logs or timing details needed to diagnose timeouts. Option B is wrong because a CloudWatch dashboard with duration metrics shows aggregated statistics (e.g., average, p99) over time, but cannot reveal per-invocation breakdowns or pinpoint the specific phase causing a timeout. Option C is wrong because adding more logging statements to the function code and checking CloudWatch Logs provides only custom log output without automatic tracing of downstream calls or sub-millisecond timing, making it insufficient for identifying intermittent timeout causes.

282
MCQeasy

A company uses Amazon CloudFront to distribute content from an S3 bucket origin. Some users report intermittent access errors. The DevOps team suspects the origin is overwhelmed. What is the MOST effective way to improve resilience?

A.Set up an origin failover with two S3 buckets behind an Application Load Balancer (ALB).
B.Reduce the CloudFront cache TTL to serve fresher content.
C.Increase the CloudFront cache TTL to reduce requests to the origin.
D.Configure CloudFront to perform health checks on the origin.
AnswerD

An origin group with health checks lets CloudFront monitor the primary origin using HTTP or HTTPS requests; when the origin returns errors or times out, CloudFront automatically fails over to a designated secondary origin. This gives near-instant recovery without manual intervention, directly addressing the overwhelmed origin scenario described in the question. Health checks are essential for the failover mechanism to know when to stop sending traffic to the failing origin.

Why this answer

Configuring CloudFront to perform health checks on the origin allows it to detect origin issues and trigger failover to a secondary origin, improving resilience. CloudFront origin groups support failover based on error rates, effectively acting as health checks. Option A is incorrect because an ALB is not required for origin failover; CloudFront can directly use multiple S3 buckets as origins.

Option B (reducing TTL) increases requests to the origin, worsening the problem. Option C (increasing TTL) reduces load but does not address origin failures or provide failover.

Exam trap

Candidates may assume an ALB is necessary for origin failover, but CloudFront can directly failover between S3 buckets using origin groups without additional load balancers.

283
MCQmedium

A DevOps engineer supports a Python application on AWS Lambda that writes structured JSON logs to CloudWatch Logs. When a request fails, the engineer must be able to search all invocations across a log group for a specific request ID and see only lines containing that ID, returning results in seconds. Which approach meets this requirement with the least effort?

A.Create a CloudWatch Logs metric filter that matches the request ID pattern and graph the resulting metric on a dashboard.
B.Subscribe the log group to a Lambda function that scans events and writes matches to an Amazon DynamoDB table.
C.Use CloudWatch Logs Insights with a query that filters on the request ID field parsed from the JSON log events.
D.Export the log group to Amazon S3 and query it with Amazon Athena using a Glue table.
AnswerC

CloudWatch Logs Insights parses JSON log events automatically, so a query can filter on the request ID field and return matching lines in seconds across the entire log group. This requires no additional infrastructure, making it the lowest-effort solution for the stated search requirement.

Why this answer

CloudWatch Logs Insights is purpose-built for interrogating log groups with a query language that understands JSON, so filtering on a request ID field returns the exact log lines quickly. It works against the existing log group without new infrastructure, which satisfies both the speed and minimal-effort constraints.

Exam trap

The trap here is confusing metric filters, which only emit numeric metrics, with query tools that can actually return matching log lines.

284
MCQhard

A company's application runs on Amazon EC2 instances in an Auto Scaling group. The application writes logs to local instance storage. The operations team needs to ensure logs are not lost during instance termination or scaling events. What should be done?

A.Increase the size of the instance store volumes.
B.Use an Amazon EFS file system and mount it to each instance for log storage.
C.Configure the Auto Scaling group to terminate instances after logs are copied to S3.
D.Install the CloudWatch Logs agent on each instance and stream logs to CloudWatch Logs.
AnswerD

Installing the CloudWatch Logs agent (or the newer unified CloudWatch agent) on each EC2 instance enables real-time delivery of log data to the CloudWatch Logs service, which stores it durably and independently of the instance lifecycle. Because logs are streamed continuously as they are generated, they survive instance termination, crashes, or scale-in events without requiring manual offload steps. This is the AWS-recommended managed solution for centralized log collection and retention.

Why this answer

The CloudWatch Logs agent streams log data in near real-time to Amazon CloudWatch Logs, ensuring logs are persisted independently of the EC2 instance lifecycle. This decouples log storage from ephemeral instance store volumes, so logs are not lost during Auto Scaling termination or scaling events. The agent handles log rotation, compression, and encryption in transit, providing a durable and centralized log solution.

Exam trap

The trap here is that candidates often assume instance store volumes are persistent or that increasing their size provides durability, when in fact instance store volumes are ephemeral and tied to the instance lifecycle, making them unsuitable for critical log data that must survive termination.

How to eliminate wrong answers

Option A is wrong because increasing instance store volume size does not address the fundamental issue that instance store volumes are ephemeral and data is lost when the instance is stopped, terminated, or replaced during scaling events. Option B is wrong because while Amazon EFS provides persistent shared storage, it introduces network latency and additional cost, and the application would need to be reconfigured to write logs to the EFS mount point; more critically, the question does not indicate that the application can handle network file system writes, and EFS is not the simplest or most AWS-native solution for log streaming. Option C is wrong because configuring the Auto Scaling group to delay termination until logs are copied to S3 is unreliable and complex; there is no native Auto Scaling lifecycle hook that guarantees logs are fully copied before termination, and this approach can cause scaling delays, race conditions, and increased operational overhead.

285
MCQmedium

A company runs a web application on EC2 instances behind an Application Load Balancer (ALB). The security team requires that all traffic to the ALB must be encrypted (HTTPS) and that the ALB must only accept traffic from CloudFront. The DevOps engineer has configured CloudFront with an origin pointing to the ALB, and the ALB has a listener on port 443 with a valid SSL certificate. The engineer also added a security group rule to the ALB that allows HTTPS traffic only from CloudFront's IP ranges. However, users are reporting intermittent 503 errors. The engineer checks CloudFront logs and sees that some requests are failing with 'Origin Connect Error'. What is the most likely cause?

A.The ALB has a Web Application Firewall (WAF) that is blocking requests from CloudFront.
B.The security group rule is using an outdated list of CloudFront IP ranges, and CloudFront has added new IP ranges that are being blocked.
C.The SSL certificate on the ALB is not trusted by CloudFront, causing handshake failures.
D.The ALB idle timeout is set too low, causing CloudFront to close connections prematurely.
AnswerB

CloudFront's origin-facing IP ranges change periodically, so a static security group rule based on a previously captured list will block newer edge locations. Those blocked connections surface as 'Origin Connect Error' and intermittent 503s, matching the stem's requirement that the ALB accept traffic only from CloudFront.

Why this answer

CloudFront publishes its origin-facing IP ranges in the managed prefix list (com.amazonaws.global.cloudfront.origin-facing) and updates it periodically. If the ALB security group was configured with a static, manually copied list of CloudFront IPs, newly added edge locations will connect from IPs not in the allow list, and the ALB will drop the TCP handshake — producing 'Origin Connect Error' in CloudFront logs and 503s to users. The correct fix is to reference the AWS-managed prefix list rather than hardcoded CIDRs.

Exam trap

DOP-C02 often tests the misconception that a static snapshot of CloudFront IP ranges is sufficient for origin lockdown, when in fact AWS updates these ranges continuously and only the managed prefix list stays current.

How to eliminate wrong answers

Option A is wrong because a WAF blocking requests would return an HTTP 403 from the ALB, not an 'Origin Connect Error' — WAF operates at layer 7 after the TCP connection succeeds. Option C is wrong because CloudFront does not validate the origin's certificate against a public CA trust chain the way a browser does; it accepts the origin certificate during the TLS handshake as long as it is valid for the origin domain, so an untrusted cert would not cause intermittent connect errors. Option D is wrong because an idle timeout mismatch would cause sporadic 504 Gateway Timeout errors on long-running requests, not 'Origin Connect Error' on connection establishment.

286
MCQmedium

A company runs a stateless web application on a fleet of EC2 instances in an Auto Scaling group. The application stores session state in a shared ElastiCache Redis cluster. During traffic spikes, the application becomes slow. Monitoring shows that the Redis cluster has high CPU utilization. Which solution is MOST cost-effective and scalable?

A.Upgrade the Redis instance to a larger node type to handle more operations
B.Enable cluster mode on the ElastiCache Redis cluster and add more shards
C.Add read replicas to offload read traffic from the primary node
D.Migrate session state to DynamoDB with DAX for caching
AnswerB

Enabling cluster mode on ElastiCache for Redis partitions the keyspace across multiple shards, with each shard having its own primary node to handle writes independently. Adding shards increases aggregate write throughput and lets you scale beyond the limits of a single node, while also supporting larger datasets by distributing memory. This is the correct answer because it directly addresses a write-heavy workload by horizontally scaling the data plane and maintains Redis' low-latency session state access.

Why this answer

Read replicas offload GET/SMEMBERS-type traffic from the primary, but every write must still be processed by the primary node, so if the workload has a significant write component, the primary's CPU and write path remain unchanged. Note: ElastiCache Redis read replicas do NOT require cluster mode — a cluster-mode-disabled replication group already supports up to 5 read replicas on a single shard. Cluster mode is only needed to add additional shards for horizontal write scaling, which is why option B (enabling cluster mode) is the more complete, scalable fix when the bottleneck may include writes.

Exam trap

The trap here is that candidates often confuse read replicas with horizontal scaling for write-heavy workloads, not realizing that replicas only help with read scaling and cannot reduce CPU from write operations, while cluster mode directly addresses both read and write scaling by splitting the data set.

How to eliminate wrong answers

Option A is wrong because upgrading to a larger node type (vertical scaling) is less cost-effective and has an upper limit; it does not provide the linear scalability of horizontal sharding and can lead to over-provisioning during low traffic. Option C is wrong because adding read replicas offloads only read traffic, but session state in Redis involves both reads and writes, and the high CPU is likely from write-heavy operations (e.g., SET/GET) that replicas cannot offload; replicas also introduce eventual consistency issues for session data. Option D is wrong because migrating to DynamoDB with DAX introduces unnecessary complexity and cost for session state that is already well-served by Redis; DAX is a separate caching layer that adds latency and cost, and DynamoDB's throughput pricing can be less predictable than ElastiCache for bursty traffic.

287
MCQeasy

A company uses AWS CodeBuild to run unit tests and package a Node.js application. The buildspec.yml file includes commands to install dependencies using npm. The build is failing with the error: 'npm ERR! code EACCES'. How should a DevOps engineer resolve this issue?

A.Configure the CodeBuild project to use a custom VPC with a NAT gateway for internet access
B.Configure the buildspec to run npm install with sudo
C.Add a command to change the ownership of the node_modules directory to the current user
D.Use 'npm ci' instead of 'npm install' and ensure a package-lock.json is present
AnswerD

Running `npm ci` instead of `npm install` is the recommended fix because it performs a clean, lockfile-driven installation and never attempts to modify the global package cache, thus eliminating the permission conflict. The presence of a valid `package-lock.json` ensures that exact dependency versions are resolved and installed without npm calculating a new dependency tree, which can trigger hidden writes to system locations. This approach is purpose-built for CI/CD pipelines and greatly reduces the chance of inconsistent state or permission failures.

Why this answer

The EACCES error in CodeBuild indicates a permission issue when npm tries to write to the node_modules directory. The default CodeBuild user (usually 'codebuild-user') lacks write permissions to the project root, which is owned by root. Using 'npm ci' (clean install) is the correct resolution because it bypasses the permission issue by using the package-lock.json to install dependencies deterministically, and it does not attempt to modify the lock file or run lifecycle scripts that may require elevated permissions.

Additionally, 'npm ci' is faster and more reliable in CI/CD environments like CodeBuild.

Exam trap

The trap here is that candidates often assume the EACCES error is a network or VPC issue (Option A) or that sudo is a quick fix (Option B), but the exam tests knowledge of npm's behavior in CI/CD and the deterministic install method 'npm ci' as the proper solution.

How to eliminate wrong answers

Option A is wrong because the EACCES error is a filesystem permission issue, not a network connectivity issue; a custom VPC with a NAT gateway addresses internet access for private subnets, not file write permissions. Option B is wrong because running npm install with sudo would escalate privileges unnecessarily and is considered a security anti-pattern in CI/CD; it also may not resolve the issue if the underlying user context still lacks proper ownership. Option C is wrong because changing ownership of node_modules does not fix the root cause—the directory may not exist yet, and the error occurs during the initial creation of node_modules, not after it exists.

288
MCQeasy

A developer is setting up AWS CodeBuild to compile a Go application. The build fails with the error: 'go: command not found'. What is the MOST likely cause?

A.The environment variables in buildspec.yml are incorrectly set
B.The build project does not have enough memory to compile Go code
C.The build environment image does not have the Go runtime installed
D.The CodeBuild service role does not have permission to access the S3 bucket for artifacts
AnswerC

The CodeBuild managed build image is version-specific and does not include every language runtime; for example, the Amazon Linux 2 x86_64 standard image version 3.0 has Java, Python, Node.js, and Ruby, but no Go binary on its PATH. When the buildspec runs 'go build', the shell returns 'go: command not found' because the executable is simply not installed in the selected image. The correct remedy is to select an image that includes Go (such as a newer standard image with the Go runtime), add a runtime-versions entry for golang, or provide a custom Dockerfile that installs Go.

Why this answer

The error 'go: command not found' indicates that the Go executable is not available in the build environment's PATH. CodeBuild uses a managed or custom build environment image (e.g., Ubuntu, Amazon Linux 2) to run build commands. If the image does not include the Go runtime, the shell cannot locate the 'go' binary, causing the build to fail.

The most direct fix is to select a build environment image that has Go pre-installed or to install Go in the install phase of the buildspec.

Exam trap

The trap here is that candidates often blame environment variables (Option A) or permissions (Option D) because they are common build failures, but the root cause is the missing runtime in the build environment image — a fundamental prerequisite that CodeBuild does not automatically provide.

How to eliminate wrong answers

Option A is wrong because environment variables in buildspec.yml control runtime behavior (e.g., GOPATH, GO111MODULE) but do not cause a 'command not found' error; the shell would still find the 'go' binary if it exists in PATH. Option B is wrong because insufficient memory leads to out-of-memory (OOM) kills or build timeouts, not a 'command not found' error; the Go compiler itself would be invoked before memory limits become an issue. Option D is wrong because S3 bucket permissions affect artifact uploads or cache retrieval, not the execution of build commands; the 'go' binary would still be found and run regardless of S3 access.

289
MCQeasy

A company uses AWS CodeBuild to run unit tests as part of a CI pipeline. The buildspec.yaml file is located in the root of the source repository. The build takes 30 minutes to complete. The team wants to speed up the build by caching dependencies. Which approach should they take?

A.Download dependencies from the internet each time the build runs.
B.Mount an Amazon EFS file system to the build container and store dependencies there.
C.Configure the buildspec.yaml to enable local caching and specify the paths to cache.
D.Store dependencies in an AWS CodeCommit repository and clone it during the build.
AnswerC

Local caching in CodeBuild persists declared dependency directories between builds on the same host, so package managers such as npm or Maven skip re-downloading. Specifying cache paths in buildspec.yaml directly cuts the 30-minute runtime without altering the pipeline structure.

Why this answer

CodeBuild supports caching dependencies to speed up builds. The simplest native approach is to enable caching in the buildspec.yaml using the cache configuration, specifying local caching paths (e.g., /root/.m2 for Maven or /root/.cache/pip for pip) so dependencies are preserved between builds on the same build environment. This avoids re-downloading dependencies from the internet each time and can significantly reduce build time.

Note that CodeBuild also supports Amazon S3 caching, and EFS can be mounted for shared storage, but local caching is the most direct and commonly recommended option for dependency caching within the buildspec.

Exam trap

Candidates may assume any persistent storage (EFS, S3, CodeCommit) will speed up builds, but the question asks for the approach configured in buildspec.yaml. Local caching is the native, buildspec-level caching mechanism for dependencies, though S3 caching is also valid in other contexts.

How to eliminate wrong answers

Option A is wrong because downloading dependencies from the internet each time is the current slow behavior the team wants to avoid, and it does not implement any caching mechanism. Option B is wrong because mounting an Amazon EFS file system introduces network latency and I/O overhead, which can actually slow down the build, and EFS is not designed for high-frequency dependency caching in ephemeral build containers; it also requires additional VPC configuration and security group management. Option D is wrong because storing dependencies in a CodeCommit repository and cloning them during the build adds unnecessary network transfer and storage costs, and does not provide the incremental caching benefit that local caching offers.

290
MCQeasy

A DevOps engineer wants to automate the creation and cleanup of temporary development environments on AWS. Each environment consists of an Amazon EC2 instance and an Amazon RDS database. The environments should be isolated and cost-effective. Which AWS service is best suited for this?

A.AWS Elastic Beanstalk
B.AWS CloudFormation
C.AWS OpsWorks
D.Amazon ECS
AnswerB

AWS CloudFormation is the definitive Infrastructure-as-Code service that uses declarative JSON/YAML templates to provision and manage all resources in a stack, including EC2 instances, RDS databases, VPCs, and security groups as a single unit. Stack creation handles dependency ordering and rollback on failure, while stack deletion removes the complete set of resources, enabling reliable automation of ephemeral environments. The ability to set a deletion policy on resources like RDS allows you to control whether the database is retained or deleted, making cleanup predictable and safe.

Why this answer

AWS CloudFormation is best suited because it allows you to define the entire temporary environment—EC2 instance and RDS database—as infrastructure as code (IaC) in a single template. You can create and delete the entire stack with one API call, ensuring isolation via separate stacks and cost-effectiveness by automating cleanup when the stack is deleted.

Exam trap

The trap here is that candidates confuse AWS Elastic Beanstalk's environment management (which does create EC2 and RDS) with the ability to easily and cost-effectively tear down temporary environments, but Elastic Beanstalk environments are designed for persistent applications and lack the simple, automated stack deletion that CloudFormation provides for ephemeral workloads.

How to eliminate wrong answers

Option A is wrong because AWS Elastic Beanstalk is a PaaS service that abstracts underlying infrastructure, but it does not provide native, automated stack-level cleanup for temporary environments; it manages long-running applications and lacks the granular, one-shot creation/deletion lifecycle needed for ephemeral dev environments. Option C is wrong because AWS OpsWorks is a configuration management service based on Chef/Puppet, designed for managing long-lived server configurations and deployments, not for automating the creation and teardown of isolated, temporary infrastructure stacks. Option D is wrong because Amazon ECS is a container orchestration service for running Docker containers, not for provisioning EC2 instances and RDS databases as a coordinated, isolated environment; it focuses on container lifecycle, not infrastructure provisioning and cleanup.

291
MCQeasy

Refer to the exhibit. The above IAM policy is attached to an IAM role used by a CI/CD pipeline. Which action is this policy allowing?

A.Start builds for any CodeBuild project in the account.
B.View details of any build in the account.
C.Start and view builds for the specified CodeBuild project.
D.Create and manage CodeBuild projects.
AnswerC

This is exactly what the policy authorizes: it includes the StartBuild action to begin a build and BatchGetBuilds to retrieve detailed information about those builds, with the Resource set to the specific CodeBuild project ARN shown in the exhibit. The policy therefore grants the minimum permissions needed to start and observe builds for only that one project.

Why this answer

The IAM policy grants `codebuild:StartBuild` and `codebuild:BatchGetBuilds` actions, which allow starting a build and viewing build details respectively. The `Resource` element restricts these permissions to the specific CodeBuild project `arn:aws:codebuild:us-east-1:123456789012:project/my-project`. Therefore, the policy allows starting and viewing builds for that single project, not any project in the account.

Exam trap

The trap here is that candidates see `codebuild:StartBuild` and `codebuild:BatchGetBuilds` and assume they apply to all projects, overlooking the resource ARN restriction that limits the policy to a single project.

How to eliminate wrong answers

Option A is wrong because `codebuild:StartBuild` is allowed only for the specified project ARN, not for all projects (`*`). Option B is wrong because `codebuild:BatchGetBuilds` is also scoped to the single project ARN, so it does not grant viewing details of any build in the account. Option D is wrong because the policy does not include actions like `codebuild:CreateProject` or `codebuild:UpdateProject`; it only covers starting and viewing builds.

292
MCQmedium

A company hosts a static website on Amazon S3 with a CloudFront distribution. The website is critical for business operations and must be available even if the primary AWS Region fails. Currently, the S3 bucket is in us-east-1, and CloudFront uses that bucket as the origin. The company has a secondary bucket in us-west-2 with a replica of the data. The company wants to use CloudFront to automatically fail over to the secondary bucket if the primary becomes unavailable. The DevOps engineer needs to implement a solution that requires minimal operational overhead. What should the engineer do?

A.Use an Application Load Balancer in front of both S3 buckets and point CloudFront to the ALB.
B.Create a second CloudFront distribution pointing to the secondary bucket and use Route 53 failover routing between the two distributions.
C.Modify the application to switch the CloudFront origin URL using Lambda@Edge when health checks fail.
D.Configure CloudFront Origin Failover by adding both buckets as origins, with the primary in us-east-1 and secondary in us-west-2.
AnswerD

CloudFront Origin Failover is the native, minimal-configuration solution: you create an origin group containing two S3 buckets as the primary and secondary origins, and attach that group to your cache behavior. When the primary origin returns a configurable HTTP error code (commonly 5xx) or a connection timeout, CloudFront automatically retries the request against the secondary bucket in us-west-2, all within the same edge location and without involving DNS or custom code. This requires no additional compute, no Route 53 policies, and no application modifications—only enabling Cross-Region Replication between the two buckets so content stays consistent. It is the only option that leverages a built-in CloudFront feature designed specifically for this use case, meeting both the fault-tolerance and low-operational-overhead requirements.

Why this answer

CloudFront Origin Failover is a native feature that allows you to designate a primary and secondary origin within a single distribution; CloudFront automatically routes requests to the secondary origin when the primary returns specific error codes (e.g., 500, 502, 503, 504) or fails health checks. This requires no additional infrastructure, no DNS changes, and no custom code, making it the lowest-operational-overhead solution. It directly meets the requirement of automatic failover to the us-west-2 bucket.

Exam trap

The trap here is that candidates often reach for Route 53 failover or Lambda@Edge because they are familiar multi-region patterns, missing that CloudFront Origin Failover is a built-in, zero-overhead feature designed exactly for this scenario.

How to eliminate wrong answers

Option A is wrong because Application Load Balancers cannot target S3 buckets as origins — ALBs route to EC2 instances, Lambda functions, or IP addresses, not S3 static website endpoints. Option B is wrong because creating a second CloudFront distribution and using Route 53 failover routing adds significant operational overhead (managing two distributions, DNS health checks, and failover policies) and is unnecessary when Origin Failover exists. Option C is wrong because Lambda@Edge cannot dynamically switch origin URLs based on health checks in a supported way — origin selection is determined at distribution configuration time, and Lambda@Edge runs at edge locations for request/response manipulation, not origin failover.

293
Multi-Selectmedium

A DevOps team is using AWS Elastic Beanstalk to deploy a web application. They need to customize the software configuration on the EC2 instances that are part of the Elastic Beanstalk environment. Which THREE methods can they use? (Choose THREE.)

Select 3 answers
A.Use saved configurations to create reusable environment templates
B.Create a custom AMI with the desired configuration
C.Use OpsWorks to manage the instances
D.Use .ebextensions configuration files
E.Use platform hooks to run custom scripts during deployment
AnswersA, D, E

Saved configurations serialize an environment's settings—such as instance type, autoscaling limits, environment variables, and security group rules—into a reusable template. A DevOps team can save a tuned production environment's configuration and then apply that exact template when creating a new environment, either through the Elastic Beanstalk console or the eb config save command. This ensures consistency across environments, but it should not be used for deploying code; it only captures environment-level options.

Why this answer

Saved configurations in Elastic Beanstalk allow you to capture the environment's settings, including custom software configurations, and reuse them as templates for future environments. This enables consistent deployment without manual reconfiguration, directly addressing the need to customize software on EC2 instances.

Exam trap

The trap here is that candidates may confuse saved configurations (option A) with a method for runtime customization, when in fact they are for environment template reuse, not direct software configuration on instances.

294
MCQeasy

A company uses AWS CloudFormation to deploy infrastructure. A stack update fails with the error 'UPDATE_ROLLBACK_FAILED'. What should the engineer do to resolve this?

A.Retry the stack update with the same parameters.
B.Delete the stack and recreate it.
C.Ignore the error and continue using the stack.
D.Use the 'ContinueUpdateRollback' operation to fix the resource that caused the failure.
AnswerD

The ContinueUpdateRollback operation is the correct recovery mechanism specifically designed for stacks that fail during an update and subsequently fail to roll back automatically, landing in UPDATE_ROLLBACK_FAILED. It instructs CloudFormation to resume the rollback to the last known good state, optionally skipping resources that are causing repeated failures via the ResourcesToSkip parameter after manually fixing them. This restores the stack to a stable, updatable condition without destroying the underlying infrastructure.

Why this answer

When a CloudFormation stack update fails and rollback also fails, the stack enters UPDATE_ROLLBACK_FAILED, a terminal state where CloudFormation cannot automatically recover. The ContinueUpdateRollback API operation tells CloudFormation to resume rolling back the stack, optionally skipping specific resources that are stuck, so the engineer can manually fix the problematic resource and then complete the rollback. This is the documented recovery path for this state.

Exam trap

The trap is thinking a failed rollback can be fixed by simply retrying or deleting the stack, when the correct action is the ContinueUpdateRollback recovery operation.

How to eliminate wrong answers

Option A is wrong because retrying the same update with identical parameters will fail again for the same underlying reason and does not address the stuck rollback. Option B is wrong because deleting and recreating the stack destroys all resources and state, causing data loss and downtime, and is a last resort rather than the correct recovery procedure. Option C is wrong because ignoring the error leaves the stack in a non-operational, inconsistent state where further updates are blocked and resources may be partially configured.

295
MCQmedium

An application running on Amazon ECS Fargate is experiencing intermittent 'CannotPullContainerError' errors. The task definition references a Docker image in a private Amazon ECR repository. The task execution role has the 'AmazonECSTaskExecutionRolePolicy' policy attached. What is the most likely cause?

A.The Fargate task is in a private subnet without a NAT gateway or VPC endpoint
B.The task execution role does not have sufficient permissions
C.The ECS service is not configured with Auto Scaling
D.The ECR repository is not in the same region as the ECS cluster
AnswerA

Fargate tasks provisioned in a private subnet have no route to the internet unless a NAT gateway is configured in a public subnet. Because ECR's API and Docker Hub need outbound HTTPS access, and image layers are fetched from Amazon S3, the task's image pull fails without a NAT gateway or VPC endpoints for ECR (both API and DKR) and S3. This manifests as a 'CannotPullContainerError' or 'ResourceInitializationError' in the task's stopped reason. Adding a NAT gateway or the appropriate VPC endpoints resolves the issue.

Why this answer

The 'CannotPullContainerError' occurs when the ECS task cannot retrieve the container image from ECR. Since the task execution role has the 'AmazonECSTaskExecutionRolePolicy' attached, which includes the necessary permissions (ecr:GetAuthorizationToken, ecr:BatchCheckLayerAvailability, ecr:BatchGetImage, ecr:GetDownloadUrlForLayer), the issue is not permissions. The most likely cause is that the Fargate task is running in a private subnet that lacks a route to the internet (via NAT gateway) or a VPC endpoint for ECR.

Without either, the task cannot reach the ECR API to pull the image. Option A is correct. Option B is wrong because the policy provides sufficient permissions.

Option C is irrelevant; Auto Scaling does not affect image pulling. Option D is less likely because ECR repositories are typically in the same region, and cross-region pulls would still be possible with proper permissions and networking.

296
Multi-Selecthard

Which THREE actions can be performed using the AWS CLI for CodeDeploy? (Choose three.)

Select 3 answers
A.create-deployment
B.get-deployment
C.push-revision
D.list-deployment-groups
E.register-instance
AnswersA, B, D

The `aws deploy create-deployment` command is a valid, top-level CodeDeploy API operation that starts a new deployment for an application. You provide the application name, deployment group name, and a revision (for example, `--s3-location` or `--github-location`), and it returns a `deploymentId` that you can later query with `get-deployment`. This is the standard way to trigger a deployment from the CLI.

Why this answer

The `create-deployment` command is correct because it is the primary AWS CLI operation used to start a new deployment in CodeDeploy. It triggers the deployment of a revision (application code and configuration) to a specified deployment group, making it a core action for automating SDLC pipelines.

Exam trap

The trap here is that candidates may confuse `push-revision` with the valid `aws deploy push` command, or assume `register-instance` is a valid command when the actual command uses the longer `register-on-premises-instance` syntax, leading to incorrect selections.

297
MCQeasy

A DevOps engineer needs to monitor the number of 4xx and 5xx HTTP errors returned by an Application Load Balancer (ALB). They want to set up a dashboard that shows the error count over the last 24 hours. Which CloudWatch metrics should they use?

A.Use the 'HTTPCode_Target_4XX_Count' and 'HTTPCode_Target_5XX_Count' metrics.
B.Use the 'RequestCount' metric with a statistic of 'ErrorCount'.
C.Use the 'HTTPCode_ELB_4XX_Count' and 'HTTPCode_ELB_5XX_Count' metrics.
D.Use the 'TargetResponseTime' metric and count the number of responses above 4 seconds.
AnswerA

These CloudWatch metrics are emitted specifically for responses returned by the registred targets, such as EC2 instances or ECS tasks, after the load balancer forwards the request. By applying a Sum statistic over your monitoring window, you get the exact number of 4xx and 5xx status codes produced by your backend application, which is precisely what you need to detect target-side errors. Unlike ELB-level metrics, these don't include errors generated by the load balancer itself.

Why this answer

The correct metrics are 'HTTPCode_Target_4XX_Count' and 'HTTPCode_Target_5XX_Count', which track HTTP errors returned by the targets behind the Application Load Balancer. Option B is incorrect because 'RequestCount' does not have an 'ErrorCount' statistic. Option C is incorrect because 'HTTPCode_ELB_4XX_Count' and 'HTTPCode_ELB_5XX_Count' are load balancer-level metrics that track errors generated by the ALB itself (e.g., due to misconfiguration), not the errors returned by targets.

Option D is incorrect because 'TargetResponseTime' measures latency, not error counts.

298
MCQhard

An organization uses AWS CodeBuild to compile a Java application. The buildspec.yml includes a pre_build phase that runs unit tests. Recently, the build started failing with 'NoClassDefFoundError' for certain test dependencies, even though the pom.xml includes them. The build environment uses an Amazon Linux 2 Docker image. What is the MOST likely cause?

A.The CodeBuild project has a cache that is corrupted or out of sync. Clear the build cache.
B.The S3 bucket for artifacts has incorrect permissions. Update the bucket policy.
C.The CodeCommit repository is not pulling the latest code. Add a webhook to trigger builds on push.
D.The build environment does not have Maven installed. Install Maven in the buildspec.
AnswerA

If the CodeBuild project is configured with a local dependency cache or an S3 cache that stores the Maven ~/.m2/repository, an interrupted download or a stale snapshot can leave cached JARs missing classes or mismatched with the declared POM versions. This directly produces NoClassDefFoundError at runtime because a class that existed during compilation is absent from the cached artifact on the classpath. Clearing the cache forces CodeBuild to re-resolve dependencies from their remote repositories, eliminating the corrupted or out-of-sync state.

Why this answer

The 'NoClassDefFoundError' for test dependencies that are declared in pom.xml indicates that the required JAR files are missing from the build environment at runtime. A corrupted or out-of-sync build cache in CodeBuild can cause previously downloaded Maven dependencies to be reused incorrectly, leading to missing classes even though the pom.xml specifies them. Clearing the cache forces a fresh download of all dependencies, resolving the mismatch.

Exam trap

The trap here is that candidates may assume the issue is a missing tool (Maven) or a code-pull problem, but the error specifically points to a dependency resolution failure, which is a classic symptom of a corrupted build cache in CI/CD pipelines.

How to eliminate wrong answers

Option B is wrong because incorrect S3 bucket permissions for artifacts would cause upload failures, not runtime class-loading errors like NoClassDefFoundError. Option C is wrong because the CodeCommit repository not pulling the latest code would result in building an outdated version of the source, not a missing dependency class; a webhook triggers builds on push but does not affect dependency resolution. Option D is wrong because the Amazon Linux 2 Docker image for CodeBuild includes Maven by default, and the build was previously succeeding, so Maven installation is not the issue.

299
MCQmedium

A DevOps engineer is troubleshooting a CloudFormation stack creation failure. The stack includes an EC2 instance with a UserData script that installs software. The stack creation fails with the error: 'The following resource(s) failed: EC2Instance (AWS::EC2::Instance) – Resource creation cancelled'. What is the most likely cause?

A.The EC2 instance type is not supported in the region
B.The IAM role attached to the instance lacks permissions
C.The stack creation timed out or was manually cancelled
D.The UserData script failed due to a syntax error
AnswerC

When a CloudFormation stack has a TimeoutInMinutes value and that timeout is exceeded, or when a user calls CancelUpdateStack or moves to cancel the operation in the console, all resources that are still in a CREATE_IN_PROGRESS state are marked with status 'Resource creation cancelled'. This is an orchestration-level interruption: CloudFormation abandons the pending resource creation and begins rolling back the stack. Because the EC2 instance creation was still in progress when the stack operation was aborted, the event correctly indicates this is the most plausible explanation.

Why this answer

The error message 'Resource creation cancelled' indicates that the CloudFormation stack creation was either timed out or manually cancelled, not that the EC2 instance itself failed. CloudFormation has a default timeout of 60 minutes for stack creation, and if the UserData script takes longer than this or the user manually cancels the operation, CloudFormation will mark the resource as cancelled. This is distinct from a resource-specific failure like an invalid instance type or IAM permissions issue, which would produce a different error message such as 'Resource failed to stabilize' or 'API error'.

Exam trap

The trap here is that candidates often confuse a UserData script failure with a resource creation failure, but CloudFormation does not monitor UserData execution; it only tracks the EC2 instance's state transition to 'running', so a script error would not cause a cancellation error.

How to eliminate wrong answers

Option A is wrong because an unsupported instance type would cause a specific API error like 'InvalidParameterValue: Instance type not supported in this region', not a 'Resource creation cancelled' message. Option B is wrong because insufficient IAM permissions would result in an 'AccessDenied' or 'UnauthorizedOperation' error during resource creation, not a cancellation error. Option D is wrong because a UserData script syntax error would not cause the EC2 instance creation to be cancelled; the instance would still be created successfully, and the script failure would be logged in the instance's system logs, but CloudFormation would not cancel the resource creation due to a script error.

300
Multi-Selectmedium

A company is using Amazon CloudWatch Logs to store application logs. The security team requires that logs are encrypted at rest using a customer-managed KMS key. Which TWO steps must be taken to achieve this?

Select 2 answers
A.Recreate the log group after associating the key.
B.Add a statement to the KMS key policy that allows CloudWatch Logs to use the key.
C.Create a KMS grant to allow CloudWatch Logs to use the key.
D.Specify the KMS key ARN when creating each log stream.
E.Use the put-log-group-encryption API to associate the KMS key with the log group.
AnswersB, E

To use a customer-managed KMS key, the key policy must explicitly grant the CloudWatch Logs service principal (logs.<region>.amazonaws.com) the required cryptographic permissions: kms:Encrypt, kms:Decrypt, kms:ReEncrypt, kms:GenerateDataKey, and kms:DescribeKey. Without an Allow statement, CloudWatch Logs cannot decrypt the key for the log group, and the association or ingest operations will fail. The key policy is the sole authorization mechanism for this integration.

Why this answer

To encrypt CloudWatch Logs at rest with a customer-managed KMS key, you must add a statement to the KMS key policy granting CloudWatch Logs permission to use the key (option B). Then, use the put-log-group-encryption API to associate the KMS key with the log group (option E). Option C is incorrect because CloudWatch Logs uses key policies, not grants.

Option D is incorrect because you specify the key ARN at the log group level, not per log stream. Option A is incorrect because you do not need to recreate the log group; you can associate the key with an existing log group.

Page 3

Page 4 of 18

Page 5