Courseiva

AWS Certified DevOps Engineer Professional DOP-C02 (DOP-C02) — Questions 376–450

1298 questions total · 18pages · All types, answers revealed

Page 5

Page 6 of 18

Page 7
376
MCQhard

A company runs a critical application on a fleet of EC2 instances managed by an Auto Scaling group. The application generates logs that are sent to CloudWatch Logs using the CloudWatch agent. Recently, the operations team noticed that some instances are missing logs for certain periods. The CloudWatch agent is configured to batch log events and send them every 5 seconds. The instances have high CPU utilization (90%+) during the missing periods. The DevOps engineer suspects that the agent is being throttled or failing. Which of the following is the MOST likely cause and the BEST course of action?

A.The network bandwidth is saturated, causing log delivery to fail. Increase instance network performance.
B.The CloudWatch Logs retention policy is set to 1 day, so older logs are deleted. Increase retention.
C.The CloudWatch agent is being starved of CPU resources, causing it to drop logs. Increase the CPU credits or instance size.
D.The instances are running out of disk space, preventing log buffering. Add more EBS volume space.
AnswerC

The CloudWatch agent runs as a separate user-space daemon that periodically reads log files and sends them to the CloudWatch Logs API. When the host's CPU is saturated — especially on T-series instances with exhausted CPU credits — the agent's log collection and flush loop can be delayed or preempted for long enough that it begins dropping buffered events to avoid creating an ever-growing backlog. Increasing instance size or CPU credits gives the agent the scheduling time it needs to reliably process and upload log batches, directly resolving the observed missing periods.

Why this answer

When CPU utilization is sustained at 90%+, the CloudWatch agent competes for CPU and can be starved, causing it to drop or delay log batches. The agent buffers events in memory and on disk; under CPU starvation, the buffer may overflow or the agent may fail to flush within the 5-second interval. Increasing CPU credits (for burstable instances) or moving to a larger instance size gives the agent the resources it needs.

Exam trap

DOP-C02 often tests resource contention — candidates blame network or disk because logs are I/O, but the question explicitly states high CPU, pointing to agent starvation.

How to eliminate wrong answers

Option A is wrong because network saturation would affect all traffic, not just logs, and the symptom is CPU-related (90%+ utilization). Option B is wrong because a 1-day retention policy deletes logs after one day, not during the missing periods, and retention does not cause gaps in delivery. Option D is wrong because disk space issues would produce agent errors about buffer overflow, and the question points to CPU as the stressor.

377
MCQmedium

A security audit reveals that an IAM user has long-term access keys that have not been rotated in over 90 days. What is the most secure way to enforce key rotation?

A.Use an AWS Lambda function to automatically rotate keys.
B.Manually rotate keys every 90 days.
C.Use IAM roles instead of long-term access keys.
D.Delete the user and create a new one.
AnswerC

IAM roles are the AWS-recommended alternative because they do not require long-term secrets at all: a principal assumes a role through the AWS Security Token Service (STS) and receives temporary credentials with a configurable lifetime (up to 12 hours for role sessions) that are automatically rotated and expire. These credentials can be further constrained by session policies, preventing privilege escalation, and eliminate the operational burden of rotating static access keys, making them far more secure and auditable.

Why this answer

A custom Lambda function to rotate keys requires you to build and maintain the rotation logic, permission grants, and secret distribution yourself - AWS IAM has no native, built-in automatic key-rotation feature for access keys, so this approach is at best a partial mitigation with real engineering overhead, and it still leaves long-term static credentials in play between rotations. IAM roles are the superior fix because they remove the need for any long-term key or rotation logic at all.

Exam trap

The trap here is that candidates focus on 'rotation' as a process (automated or manual) rather than recognizing that the most secure solution is to eliminate the need for rotation entirely by using IAM roles with temporary credentials.

How to eliminate wrong answers

Option A is wrong because while a Lambda function can automate key rotation, it still relies on long-term access keys (the user still exists with keys that must be rotated), and implementing such a solution introduces complexity, potential security gaps (e.g., Lambda execution role permissions), and does not eliminate the fundamental risk of long-term credentials. Option B is wrong because manual rotation every 90 days is error-prone, relies on human compliance, and does not address the underlying security issue that long-term keys can be exfiltrated and used for extended periods before detection. Option D is wrong because deleting and recreating the user does not solve the problem—the new user would still have long-term access keys that require rotation, and this approach disrupts workflows without addressing the root cause.

378
MCQhard

A DevOps engineer is designing a CI/CD pipeline for a microservices architecture on AWS. They want to use AWS CodeDeploy to deploy applications to an Auto Scaling group. The pipeline must ensure that only a small percentage of instances are updated at a time, and if health checks fail, the deployment is automatically rolled back. Which deployment configuration should be used?

A.Blue/green deployment with a fixed number of instances.
B.In-place deployment with 'CodeDeployDefault.AllAtOnce' configuration.
C.In-place deployment with 'CodeDeployDefault.HalfAtATime' configuration.
D.In-place deployment with 'CodeDeployDefault.OneAtATime' configuration and automatic rollback enabled.
AnswerD

In-place deployment with 'CodeDeployDefault.OneAtATime' configuration is the built-in CodeDeploy strategy that stages the deployment to a single instance at a time, waiting for the instance to pass health checks before proceeding to the next. When the fleet size is reasonably large, one instance represents a small percentage of total traffic, satisfying the gradual rollout requirement. Enabling automatic rollback ensures that if any instance fails its health check or a deployment hook returns a non-zero exit code, CodeDeploy immediately redeploys the previous revision and stops the deployment, minimizing impact.

Why this answer

`CodeDeployDefault.OneAtATime` deploys to one instance at a time, which is the smallest possible batch and satisfies the 'small percentage of instances' requirement. Enabling automatic rollback ensures that if health checks fail, CodeDeploy reverts to the last known good revision, meeting the rollback requirement.

Exam trap

The trap is equating 'small percentage' with HalfAtATime (50% is not small) or assuming blue/green is always safer — but blue/green does not meet the explicit 'small percentage of instances updated at a time' wording in the question.

How to eliminate wrong answers

Option A is wrong because blue/green deployment shifts traffic between two environments and does not inherently limit updates to a small percentage of instances in an Auto Scaling group — it replaces the whole fleet, and 'fixed number of instances' is not a standard CodeDeploy configuration name. Option B is wrong because `AllAtOnce` updates every instance simultaneously, maximizing blast radius and violating the small-percentage requirement. Option C is wrong because `HalfAtATime` updates 50% of instances at once, which is not a 'small percentage' and would cause significant capacity loss during deployment.

379
Multi-Selectmedium

A company is designing a multi-region disaster recovery strategy for a stateless web application. They want to minimize RTO and RPO. Which TWO of the following should they implement? (Choose TWO.)

Select 2 answers
A.Use cross-region replication for data stores.
B.Use a passive standby in a single Availability Zone.
C.Perform periodic backups and restore in the DR region.
D.Configure cross-region read replicas for the database.
E.Deploy an active-active workload using Route 53 weighted routing.
AnswersA, E

Cross-region replication for data stores, such as Amazon S3 CRR, DynamoDB global tables, or Aurora Global Database, synchronizes data between regions automatically, typically with sub-minute RPO. This ensures that when a regional failure occurs, the application can fail over to the DR region with minimal data loss, making it a strong foundation for a multi-region DR strategy. Because replication is continuous rather than point-in-time, it provides a much lower RPO than backup and restore, and it often pairs with per-region compute stacks to achieve low RTO.

Why this answer

Cross-region replication for data stores ensures that data is continuously synchronized to a secondary AWS region, minimizing Recovery Point Objective (RPO) to near-zero and reducing Recovery Time Objective (RTO) as the data is already available in the DR region. This approach avoids the need to restore from backups, which would increase RTO and potentially lose recent transactions.

Exam trap

The trap here is that candidates often confuse cross-region read replicas (which are read-only and not suitable for active-active writes) with true multi-region replication solutions, leading them to select Option D instead of Option A or E.

380
Multi-Selecthard

A DevOps engineer needs to set up a monitoring solution for an application running on Amazon EKS. The application emits custom metrics that need to be stored in Amazon CloudWatch and visualized on a dashboard. Which THREE steps should the engineer take? (Choose THREE.)

Select 3 answers
A.Configure the CloudWatch agent to emit custom metrics to CloudWatch.
B.Use CloudWatch Logs Insights to analyze the custom metrics.
C.Create a CloudWatch dashboard to visualize the collected metrics.
D.Install the CloudWatch agent on the EKS cluster using a DaemonSet.
E.Use Amazon Managed Service for Prometheus to scrape the metrics.
AnswersA, C, D

The CloudWatch agent is the correct mechanism for collecting custom metrics from an EKS cluster and emitting them to CloudWatch. By configuring the agent with a metrics collection interval and a custom namespace, you can send application and cluster-level metrics (e.g., pod CPU, memory, or custom application counters) via the PutMetricData API. This is the foundational step that makes those metrics available for dashboards, alarms, and further analysis within CloudWatch, so it is a required part of a CloudWatch-centric monitoring solution.

Why this answer

The CloudWatch agent can be configured to emit custom application metrics to Amazon CloudWatch, which is the required destination for storing the metrics. The agent uses the CloudWatch PutMetricData API to send these metrics, enabling centralized monitoring and alerting within CloudWatch.

Exam trap

The trap here is that candidates may confuse CloudWatch Logs Insights (for logs) with CloudWatch Metrics (for numeric data), or assume Amazon Managed Service for Prometheus is a direct replacement for CloudWatch metrics, when the question specifically requires storing custom metrics in CloudWatch.

381
MCQhard

A DevOps engineer creates a CloudFormation stack with the above template. After creation, they want to update the Lambda function code by uploading a new zip file to the S3 bucket and updating the S3Key property. However, the stack update fails because the Lambda function is published as a version and the alias points to that version. What is the most likely reason for the update failure?

A.The AWS::Lambda::Version resource is immutable and cannot be updated.
B.The alias must be deleted before updating the function code.
C.The IAM role does not have permission to update the function.
D.The function code cannot be updated because the S3 bucket is in a different region.
AnswerA

The AWS::Lambda::Version resource is immutable; once a Lambda version is published, its code and configuration are fixed and cannot be modified in place. CloudFormation therefore rejects any update attempt that alters existing AWS::Lambda::Version properties, reporting that the resource cannot be updated. To change the version, you must create a new version (e.g., by changing the logical ID or using SAM's AutoPublishAlias) and then redirect the alias. This is exactly why the stack update fails.

Why this answer

The AWS::Lambda::Version resource is immutable by design; once created, it cannot be updated or replaced. When a CloudFormation stack includes a Lambda version and an alias pointing to that version, any attempt to update the function code (e.g., by changing the S3Key property) triggers an update to the AWS::Lambda::Version resource, which CloudFormation cannot perform because versions are immutable. This causes the stack update to fail.

Exam trap

The trap here is that candidates assume CloudFormation can update any resource, but they overlook the immutable nature of Lambda versions, which causes the update to fail even when the alias is present.

How to eliminate wrong answers

Option B is wrong because the alias does not need to be deleted; the alias can be updated to point to a new version after the function code is updated, but the core issue is the immutability of the version resource. Option C is wrong because the IAM role permissions are not the cause of the failure; the error occurs at the CloudFormation resource level, not due to missing permissions. Option D is wrong because the S3 bucket region does not affect the ability to update the function code; cross-region S3 buckets are supported as long as the bucket name and key are correctly specified.

382
MCQhard

A DevOps engineer is troubleshooting a CloudFormation stack that is in UPDATE_ROLLBACK_FAILED state. The stack attempted to update an Auto Scaling group but failed due to insufficient capacity in the Availability Zone. What is the recommended next step?

A.Execute a new stack update with the same template
B.Manually increase the Auto Scaling group capacity in the affected AZ
C.Use the ContinueUpdateRollback operation to skip the resources that failed
D.Delete the Auto Scaling group and recreate it
AnswerC

The ContinueUpdateRollback operation is the correct remediation because it explicitly resumes the failed rollback, allowing CloudFormation to finish reverting the stack to its last known good state. By specifying ResourcesToSkip, you can instruct CloudFormation to skip the specific resources that caused the rollback to fail, such as resources with external dependencies that cannot be rolled back automatically. This is the documented, supported approach to recover a stack stuck in UPDATE_ROLLBACK_FAILED. Once the rollback completes, the stack returns to a usable state, and you can then address the skipped resource manually.

Why this answer

When a CloudFormation stack is in UPDATE_ROLLBACK_FAILED state, the recommended next step is to use the ContinueUpdateRollback operation with the 'ResourcesToSkip' parameter to skip the resources that failed during rollback. This allows CloudFormation to complete the rollback of the remaining resources and move the stack to a stable state, after which you can investigate and fix the underlying issue (e.g., insufficient capacity in the AZ) before attempting the update again. Option C directly aligns with AWS documentation for handling this specific stack state.

Exam trap

The trap here is that candidates may think manually fixing the underlying issue (e.g., increasing capacity) is sufficient to resolve the rollback failure, but they overlook that CloudFormation requires an explicit ContinueUpdateRollback API call to exit the UPDATE_ROLLBACK_FAILED state, even after the root cause is addressed.

How to eliminate wrong answers

Option A is wrong because executing a new stack update with the same template will likely fail again due to the same insufficient capacity issue in the AZ, and CloudFormation will not bypass the failed resource without explicit skip instructions. Option B is wrong because manually increasing the Auto Scaling group capacity in the affected AZ does not resolve the rollback failure; the stack remains in UPDATE_ROLLBACK_FAILED state and requires the ContinueUpdateRollback operation to proceed. Option D is wrong because deleting the Auto Scaling group and recreating it is an overly destructive action that does not address the stack's rollback state and may cause data loss or service disruption; CloudFormation provides the ContinueUpdateRollback operation specifically to handle such scenarios without manual resource deletion.

383
MCQeasy

A company wants to design a resilient architecture for a web application using AWS services. Which of the following is a best practice for improving resilience?

A.Deploy EC2 instances in multiple Availability Zones.
B.Use an Auto Scaling group in a single AZ.
C.Use a single AZ with RDS Multi-AZ.
D.Use one large EC2 instance to handle all traffic.
AnswerA

Placing EC2 instances in multiple Availability Zones (AZs) and fronting them with an Elastic Load Balancer and an Auto Scaling group ensures that if an entire AZ becomes unavailable, the load balancer can route traffic only to healthy instances in the remaining AZs. This pattern provides fault tolerance at the AZ granularity, which is the foundation of a resilient web architecture. It also allows the application to absorb a single-AZ failure without requiring any manual intervention.

Why this answer

Deploying EC2 instances across multiple Availability Zones (AZs) is a fundamental best practice for resilience because it eliminates a single point of failure at the data center level. If one AZ experiences an outage, traffic can be automatically routed to healthy instances in other AZs via an Elastic Load Balancer (ELB), ensuring application availability. This approach aligns with the AWS Well-Architected Framework's Reliability Pillar, which mandates distributing workloads across multiple AZs to achieve high availability.

Exam trap

The trap here is that candidates often confuse database-level high availability (RDS Multi-AZ) with full application resilience, mistakenly thinking that a single-AZ compute layer is acceptable as long as the database is redundant.

How to eliminate wrong answers

Option B is wrong because using an Auto Scaling group in a single AZ creates a single point of failure; if that AZ becomes unavailable, all instances are lost, and the application goes down. Option C is wrong because RDS Multi-AZ provides high availability for the database layer, but the compute layer (EC2) remains in a single AZ, meaning an AZ failure still takes down the web application. Option D is wrong because relying on one large EC2 instance violates the principle of horizontal scaling and introduces a single point of failure; if the instance fails or the AZ fails, the entire application becomes unavailable.

384
Multi-Selectmedium

Which TWO AWS services can be used as sources in an AWS CodePipeline? (Choose two.)

Select 2 answers
A.AWS Lambda
B.AWS CodeCommit
C.AWS CloudFormation
D.Amazon S3
E.AWS CodeBuild
AnswersB, D

AWS CodeCommit is a fully managed source control service hosting private Git repositories, making it a first-class source provider in CodePipeline. When you attach a CodeCommit repository as the source stage of a pipeline, the pipeline automatically triggers on new commits to a selected branch, using CloudWatch Events or polling. Because CodeCommit stores versioned source code and supports branches and tags, it serves as a reliable source from which CodePipeline can pull artifacts for subsequent build and deploy stages.

Why this answer

AWS CodePipeline supports AWS CodeCommit and Amazon S3 as source stages. CodeCommit is a fully managed source control service that integrates natively with CodePipeline, allowing automatic pipeline execution on code changes. Amazon S3 can serve as a source when you upload a source bundle (e.g., a ZIP file) to an S3 bucket, and CodePipeline can poll the bucket for changes or use S3 event notifications to trigger the pipeline.

Exam trap

The trap here is that candidates often confuse services that can be used as actions (like Lambda, CodeBuild, or CloudFormation) with services that can be used as sources, leading them to select Lambda or CodeBuild as source options.

385
MCQmedium

A company runs a microservices application on Amazon ECS with Fargate launch type. The application experiences intermittent failures when calling an external API. The errors are transient and usually resolve within a few seconds. How should the company improve resilience?

A.Increase the timeout of the external API call to 60 seconds.
B.Implement retry logic with exponential backoff in the application code.
C.Increase the number of tasks in the ECS service to handle failures.
D.Use an Amazon SQS queue to decouple the API call from the application.
AnswerB

Transient external API failures resolve within seconds, so retrying with exponential backoff lets the application recover without overwhelming the dependency. This directly addresses the intermittent, short-lived errors described, unlike circuit breakers or timeouts that would abort calls rather than succeed on retry.

Why this answer

Transient errors from an external API are best handled with retry logic using exponential backoff and jitter. This allows the application to recover from temporary failures without overwhelming the downstream API, improving resilience without architectural changes.

Exam trap

DOP-C02 often tests the difference between scaling (more tasks) and resilience patterns (retries, backoff) — candidates pick 'increase tasks' because it sounds like high availability, but it does nothing for transient downstream errors.

How to eliminate wrong answers

Option A is wrong because increasing the timeout to 60 seconds does not address transient failures — it just makes the caller wait longer and can exhaust resources. Option C is wrong because adding more ECS tasks increases capacity but does not fix the underlying call failures; each task would still fail. Option D is wrong because SQS decouples asynchronous processing but the application needs a synchronous response from the external API; queuing does not solve transient call failures and adds complexity.

386
Multi-Selecthard

A DevOps team is troubleshooting an application that occasionally throws 'Connection reset by peer' errors when connecting to an RDS MySQL instance. The errors are intermittent and seem to correlate with high traffic. Which TWO steps should the team take to diagnose the issue?

Select 2 answers
A.Check the RDS error log for messages about connection timeouts or aborted connections.
B.Increase the max_connections parameter in the DB parameter group.
C.Enable Multi-AZ deployment to provide a standby instance.
D.Enable Performance Insights to analyze database load and find bottlenecks.
E.Review the security group rules to ensure the application can connect.
AnswersA, D

The RDS error log is the first place to look because it records 'Aborted connection' entries with explicit reasons such as 'Got timeout reading communication packets' or 'Client has exceeded the max_user_connections' when a reset occurs. These entries include the source host, user, and exact error code, which lets you determine whether the reset originates from the database server (e.g., idle timeout) or from the client side. This evidence directly confirms or rules out connection-level failures before you change any configuration or architecture.

Why this answer

Option A is correct because the RDS error log records server-side events such as 'Aborted connection' and connection timeout messages that directly explain why the server is resetting client connections under load. Option D is correct because Performance Insights captures database load (DBLoad) broken down by wait events and top SQL, letting the team pinpoint the bottleneck causing intermittent resets during high traffic. Option B is not a diagnostic step and blindly raising max_connections may not address the root cause.

Option C is a high-availability configuration change, not a troubleshooting action, and Multi-AZ does not resolve connection resets. Option E is unlikely to be the cause since security group misconfiguration would produce consistent connection failures, not intermittent resets correlated with traffic.

Exam trap

DOP-C02 often tests the tendency to jump to configuration changes (like increasing max_connections) instead of first diagnosing via logs and performance tools.

387
MCQeasy

A DevOps engineer needs to rotate database credentials stored in AWS Secrets Manager automatically every 30 days. What is the simplest way to achieve this?

A.Enable automatic rotation in Secrets Manager with a rotation interval of 30 days.
B.Store the credentials in Systems Manager Parameter Store and use a scheduled automation to update them.
C.Create a CloudWatch Events rule that triggers a Lambda function to rotate the secret.
D.Write a custom Lambda function that rotates the secret and schedule it with CloudWatch Events.
AnswerA

Secrets Manager automatically rotates the secret on a schedule you define (e.g., 30 days) using a built-in Lambda template tailored to your database engine, such as RDS for PostgreSQL or MySQL. Because the rotation orchestration—including updating the secret and testing it against the database—is handled entirely by the service, you avoid writing or maintaining custom code and scheduling components. This gives you a fully managed, auditable rotation process with no operational overhead beyond configuring the interval.

Why this answer

AWS Secrets Manager has built-in automatic rotation: you enable it on the secret, specify a rotation interval (e.g., 30 days), and provide a Lambda rotation function (AWS provides templates for RDS, Redshift, DocumentDB). This is the simplest, fully managed approach and requires no custom scheduling infrastructure.

Exam trap

The trap is over-engineering — candidates pick the custom Lambda + CloudWatch Events option because it sounds more 'controlled', but the exam rewards the managed, built-in rotation feature when the requirement is simply periodic rotation.

How to eliminate wrong answers

Option B is wrong because Systems Manager Parameter Store does not natively rotate secrets — you would have to build and schedule the rotation logic yourself, which is more complex than Secrets Manager's built-in feature. Option C is wrong because creating a CloudWatch Events rule to trigger a Lambda is exactly what Secrets Manager does internally; doing it manually adds unnecessary components and management overhead. Option D is wrong because writing a custom Lambda and scheduling it with CloudWatch Events reimplements functionality Secrets Manager already provides out of the box, violating the 'simplest way' requirement.

388
MCQeasy

A DevOps engineer needs to securely store database credentials for an application running on Amazon ECS. Which AWS service should be used to manage the credentials and provide them to the ECS tasks?

A.AWS Secrets Manager
B.Amazon S3 with server-side encryption
C.AWS Systems Manager
D.AWS Systems Manager Parameter Store
AnswerA

AWS Secrets Manager is the correct choice because it is a purpose-built service for securely storing and managing database credentials, API keys, and other secrets. It provides native automatic rotation of database credentials via Lambda, fine-grained IAM-based access control, and built-in audit integration with AWS CloudTrail and Amazon EventBridge. Unlike generic storage services, Secrets Manager caches secrets securely and enforces resource-based policies, making it the only fully managed solution that directly addresses the requirements.

Why this answer

AWS Secrets Manager is purpose-built for storing, rotating, and retrieving secrets such as database credentials, and it integrates natively with ECS so tasks can inject secrets as environment variables or in log configuration. It supports automatic rotation via Lambda, which is ideal for database credentials. ECS task definitions reference Secrets Manager secrets using the secrets block with valueFrom.

Exam trap

DOP-C02 often tests the distinction between Secrets Manager and Parameter Store — candidates pick Parameter Store for database credentials, but Secrets Manager is preferred when automatic rotation and native secret management are required.

How to eliminate wrong answers

Option B is wrong because S3 with SSE stores objects but is not designed for secret management, lacks native rotation, and requires custom code to retrieve and inject credentials. Option C is wrong because AWS Systems Manager is a broad operational service (patch management, run command, session manager); while it includes Parameter Store, the service itself is not the credential-management answer. Option D is wrong because Systems Manager Parameter Store can store SecureString parameters, but it lacks built-in automatic rotation and is better suited for configuration data; Secrets Manager is the recommended service for database credentials with rotation.

389
MCQeasy

A development team uses AWS CodeCommit to store their application code. They want to enforce that all code changes to the main branch are reviewed and approved by at least one other team member before being merged. Which AWS service or feature should they use to implement this requirement?

A.AWS Identity and Access Management (IAM) policies
B.AWS CodePipeline manual approval action
C.AWS CodeCommit approval rule templates
D.AWS CodeCommit repository triggers
AnswerC

Approval rule templates in CodeCommit allow you to define the number of approvals required for a pull request to be merged. You can apply a template to a repository to enforce that at least one approval is needed before merging into the main branch. This directly meets the requirement for mandatory code review.

Why this answer

CodeCommit approval rule templates define the number of approvals required for a pull request to be merged. By applying a template to the repository, you can enforce that at least one approval is needed before merging into the main branch. This is the native feature for implementing mandatory code reviews in CodeCommit.

Exam trap

The trap here is confusing deployment approval (CodePipeline manual approval) with code review approval (CodeCommit approval rules), which serve different purposes.

390
MCQmedium

A company uses AWS Lambda to process messages from an Amazon SQS queue. The Lambda function occasionally times out after 15 seconds. To improve resilience, the team wants to ensure messages are not lost and are retried. Which configuration is MOST appropriate?

A.Reduce the Lambda timeout to 5 seconds to fail fast and retry quickly.
B.Set the SQS queue visibility timeout to less than the Lambda timeout.
C.Increase the batch size and remove the DLQ to speed up processing.
D.Increase the Lambda timeout to 30 seconds and configure a dead-letter queue (DLQ) for the SQS queue.
AnswerD

Increasing the Lambda function timeout to 30 seconds gives the SQS-triggered processor enough wall-clock time to complete network calls, database writes, or third-party integrations that were failing under a shorter timeout, while still remaining within a reasonable operational bound. Configuring a dead-letter queue (DLQ) on the SQS queue ensures that messages that still repeatedly fail after retries are moved to a separate queue for later inspection and redrive, preventing poison messages from consuming the main queue indefinitely. Together, these changes address the immediate failure mode and give you a durable, observable mechanism for handling the few messages that cannot be processed successfully.

Why this answer

Increasing the Lambda timeout to 30 seconds accommodates the occasional processing delays that cause the current 15-second timeout, preventing premature failures. Configuring a dead-letter queue (DLQ) for the SQS queue ensures that messages that repeatedly fail after all retries are exhausted are preserved for analysis and manual reprocessing, rather than being lost. This combination directly addresses the requirement to not lose messages and to allow retries, as Lambda will automatically retry failed invocations up to the function's configured retry count (default 2) before sending the message to the DLQ.

Exam trap

The trap here is that candidates may think reducing the timeout or adjusting the visibility timeout alone improves resilience, but they overlook the critical need for a DLQ to prevent message loss and the necessity of matching the visibility timeout to the function's execution window to avoid duplicate processing.

How to eliminate wrong answers

Option A is wrong because reducing the Lambda timeout to 5 seconds would cause even more frequent timeouts, increasing failures without solving the underlying processing issue, and does not preserve messages for retry. Option B is wrong because setting the SQS visibility timeout to less than the Lambda timeout would cause messages to become visible again in the queue while the Lambda function is still processing them, leading to duplicate processing and potential data inconsistency. Option C is wrong because increasing the batch size would increase the processing load per invocation, likely worsening timeouts, and removing the DLQ would cause messages that exceed the maximum retries to be silently discarded, violating the requirement to not lose messages.

391
MCQmedium

A DevOps engineer wants to use AWS CodeDeploy to deploy an application to an Auto Scaling group. The deployment must ensure that only a certain percentage of instances are taken out of service at a time. Which deployment configuration supports this requirement?

A.CodeDeployDefault.OneAtATime
B.CodeDeployDefault.LambdaCanary10Percent5Minutes
C.CodeDeployDefault.AllAtOnce
D.CodeDeployDefault.LambdaLinear10PercentEvery1Minute
AnswerA

This is the right choice for minimizing risk during deployment to an EC2/ASG. It shifts traffic to one instance at a time, waiting for a successful deployment health check on that instance before proceeding to the next, so if something fails, only that instance is affected and deployment can be stopped before impacting the rest of the fleet. This is akin to a rolling update with a batch size of one, preserving overall availability.

Why this answer

CodeDeployDefault.OneAtATime is the correct deployment configuration because it ensures that only one instance in the Auto Scaling group is taken out of service at a time, which directly satisfies the requirement of limiting the percentage of instances removed during deployment. This configuration is designed for EC2/On-Premises deployments and uses a fixed number of instances (one) rather than a percentage, making it ideal for gradual, safe rollouts.

Exam trap

The trap here is that candidates often confuse deployment configurations designed for Lambda functions (like Canary and Linear) with those for EC2/On-Premises, or they mistakenly think AllAtOnce limits the percentage of instances taken out of service, when in fact it takes all instances out at once.

How to eliminate wrong answers

Option B is wrong because CodeDeployDefault.LambdaCanary10Percent5Minutes is a deployment configuration for AWS Lambda functions, not for EC2/On-Premises deployments to an Auto Scaling group; it shifts 10% of traffic to the new version and then waits 5 minutes before shifting the remaining 90%. Option C is wrong because CodeDeployDefault.AllAtOnce deploys to all instances simultaneously, which would take all instances out of service at once, violating the requirement to limit the percentage removed at a time. Option D is wrong because CodeDeployDefault.LambdaLinear10PercentEvery1Minute is also a Lambda-specific configuration that increments traffic by 10% every minute, and it does not apply to EC2/On-Premises deployments with Auto Scaling groups.

392
MCQeasy

A company is running a critical application on Amazon RDS for PostgreSQL. The DevOps team needs to set up monitoring to detect when database connections exceed 80% of the maximum connections for more than 5 minutes. Which CloudWatch metric should be used to create an alarm?

A.DatabaseConnections
B.FreeableMemory
C.CPUUtilization
D.DiskQueueDepth
AnswerA

DatabaseConnections reports the count of client sessions currently connected to the RDS PostgreSQL instance, so an alarm threshold set at 80% of max_connections with a five-minute evaluation period detects sustained connection saturation, exactly the condition the DevOps team must monitor.

Why this answer

Amazon RDS publishes the DatabaseConnections metric to CloudWatch, representing the number of client connections currently open to the DB instance. To alarm at 80% of max connections, you compare DatabaseConnections against the instance's max_connections parameter. This is the only listed metric that directly measures connection count.

Exam trap

The trap here is assuming CPUUtilization or FreeableMemory reflects connection saturation; candidates must recognize DatabaseConnections as the direct metric for connection limits.

How to eliminate wrong answers

Option B is wrong because FreeableMemory measures available RAM, which can correlate with connection pressure but does not count connections. Option C is wrong because CPUUtilization measures processor load and may stay low even when connections are near the limit. Option D is wrong because DiskQueueDepth measures pending disk I/O operations, unrelated to connection saturation.

393
MCQeasy

A company uses Amazon RDS for MySQL with Multi-AZ deployment. The application experiences increased latency during peak hours. The DevOps engineer investigates and notices that the Read Replicas are not being utilized effectively. The application is configured to use the primary database endpoint. The engineer wants to offload read traffic to the Read Replicas without changing the application code. What is the BEST solution?

A.Increase the instance size of the primary database to handle the load.
B.Modify the application to use separate endpoints for read and write operations.
C.Create a new Multi-AZ cluster with a read-only endpoint.
D.Configure Amazon RDS Proxy in front of the database and enable read/write splitting.
AnswerD

Amazon RDS Proxy sits between the application and the database, pooling and reusing connections to reduce connection overhead, and when read/write splitting is enabled it can automatically route read queries to one or more Read Replicas while sending write transactions to the primary. This transparently offloads read traffic from the primary without requiring application code changes or manual endpoint management. The proxy also maintains session state and preserves transaction semantics, making it the correct solution for reducing load on a Multi-AZ RDS MySQL instance.

Why this answer

Amazon RDS Proxy with read/write splitting allows the application to offload read traffic to Read Replicas without code changes. It automatically directs read queries to Read Replicas and write queries to the primary, using a single endpoint. Option A (increasing instance size) does not leverage Read Replicas.

Option B (modifying application) requires code changes. Option C (Multi-AZ cluster with read-only endpoint) is invalid because Multi-AZ clusters use a writer and reader endpoint, but the question requires no code changes, and the existing configuration uses a primary endpoint; also, creating a new cluster may not be the simplest solution compared to RDS Proxy.

394
MCQhard

A company runs a critical batch processing workload on Amazon EMR that must complete within a 2-hour window each night. The workload is fault-tolerant but must be resilient to instance failures. Currently, the EMR cluster uses instance fleets with Spot Instances. Recently, Spot Instance interruptions caused the cluster to take over 3 hours to complete. Which change will MOST effectively ensure the workload completes within the 2-hour window despite Spot interruptions?

A.Increase the number of core nodes to 20 to improve parallelism.
B.Switch to using On-Demand instances for all nodes.
C.Use a mixed instances policy that includes multiple instance types across different Availability Zones.
D.Configure the cluster to terminate idle nodes after 5 minutes to reduce costs.
AnswerC

Using a mixed instances policy with multiple instance types across different Availability Zones is the recommended way to reduce Spot interruption risk in EMR. This approach makes the cluster's Spot capacity pool more diverse, so a capacity reclamation event in one pool is unlikely to affect all nodes simultaneously. EMR's instance fleets can automatically provision from the specified pools, improving both initial capacity acquisition and fault tolerance during interruptions.

Why this answer

A mixed instances policy across multiple Availability Zones increases the diversity of Spot capacity pools. When one instance type or zone experiences interruptions, the cluster can fall back to other pools, reducing the likelihood of prolonged delays. This approach directly addresses Spot interruption risk without sacrificing cost efficiency, as On-Demand instances would.

Exam trap

The trap here is that candidates may assume increasing parallelism (Option A) or using On-Demand instances (Option B) are the only ways to handle Spot interruptions, overlooking the cost-effective and resilient design of mixed instances across zones.

How to eliminate wrong answers

Option A is wrong because simply increasing core nodes to 20 does not mitigate Spot interruptions; it only adds parallelism, which may not help if all nodes are interrupted simultaneously. Option B is wrong because switching entirely to On-Demand instances eliminates Spot interruption risk but significantly increases cost, which is not the most effective solution given the fault-tolerant nature of the workload. Option D is wrong because terminating idle nodes after 5 minutes reduces cost but does not address the root cause of Spot interruptions causing delays; it may even worsen performance by removing nodes that could be reused.

395
MCQmedium

Your company runs a multi-tier web application on AWS. The application consists of an Application Load Balancer (ALB) that distributes traffic to a fleet of Amazon EC2 instances running a web server. The web servers write access logs to a shared Amazon EFS filesystem. The operations team needs to monitor the web server logs in real-time to detect and alert on 5xx error spikes. Currently, the team manually SSHes into instances to tail logs, which is inefficient and doesn't provide real-time alerting. The team wants a centralized, near-real-time logging solution with minimal operational overhead. They have asked you to design a solution that ingests logs from the EFS filesystem into a centralized log analytics platform. Which solution would you recommend?

A.Enable AWS CloudTrail data events for the EC2 instances to capture log file modifications.
B.Configure an Amazon EventBridge scheduled rule to invoke an AWS Lambda function that reads new log lines from EFS and publishes them to Amazon CloudWatch Logs.
C.Stream the log files to Amazon Kinesis Data Streams using a custom producer, then use a Lambda function to analyze and alert on 5xx errors.
D.Install and configure the Amazon CloudWatch Logs agent on each EC2 instance to tail the log files from the EFS mount and send them to CloudWatch Logs. Create a metric filter and alarm for 5xx errors.
AnswerD

The Amazon CloudWatch Logs agent (now part of the unified CloudWatch agent) can be installed on each EC2 instance to monitor the EFS-mounted log file and push new lines to CloudWatch Logs in near-real-time. After the log group receives the entries, a metric filter can extract the '5xx' HTTP status code pattern to create a custom metric, and a CloudWatch alarm on that metric will page the team when the error rate breaches a threshold. This is the purpose-built, low-overhead solution that supports tailing, rotation, and automatic delivery.

Why this answer

Installing the CloudWatch Logs agent on each EC2 instance allows it to tail the log files from the shared EFS mount point and stream them to CloudWatch Logs in near real-time. This provides centralized log ingestion with minimal operational overhead, and you can create a metric filter and alarm to detect and alert on 5xx error spikes without manual SSH access.

Exam trap

The trap here is that candidates may overcomplicate the solution by choosing Kinesis or Lambda-based approaches (Options B and C) when a simple agent-based solution (Option D) is sufficient, or they may confuse CloudTrail data events (Option A) with log file monitoring, not realizing CloudTrail captures API activity, not file content changes.

How to eliminate wrong answers

Option A is wrong because CloudTrail data events for EC2 instances capture API calls (e.g., RunInstances, TerminateInstances), not log file modifications on EFS; they cannot ingest or analyze web server log content. Option B is wrong because an EventBridge scheduled rule with a Lambda function that reads new log lines from EFS would introduce latency (scheduled intervals) and complexity in tracking file offsets, making it unsuitable for near-real-time monitoring. Option C is wrong because streaming logs to Kinesis Data Streams requires a custom producer to be deployed and managed, adding significant operational overhead compared to the agent-based approach, and it does not directly integrate with CloudWatch Logs for metric filtering and alerting without additional Lambda processing.

396
MCQmedium

A company needs to store audit logs for 7 years to meet compliance requirements. Which S3 storage class is the most cost-effective for long-term archival?

A.S3 Glacier Deep Archive
B.S3 Intelligent-Tiering
C.S3 Standard
D.S3 Glacier Flexible Retrieval
AnswerA

S3 Glacier Deep Archive is the lowest-cost storage class in Amazon S3, designed for long-term retention of data accessed at most once or twice per year. With a per-GB storage price that is significantly lower than both Standard and Glacier Flexible Retrieval, it is the most cost-effective choice for a 7-year compliance archive where retrieval latency of up to 12 hours is acceptable. This makes it the correct answer for the stated requirement.

Why this answer

S3 Glacier Deep Archive is the lowest-cost S3 storage class, designed for data accessed less than once per year, with a standard retrieval time of 12 hours and a minimum storage duration of 180 days. For 7-year audit log retention where retrieval is rare and compliance-driven, it is the most cost-effective choice.

Exam trap

DOP-C02 often tests the confusion between Glacier Flexible Retrieval and Glacier Deep Archive, so candidates who pick Flexible Retrieval for 'archival' miss the cost-effectiveness requirement for rarely accessed data.

How to eliminate wrong answers

Option B is wrong because S3 Intelligent-Tiering adds monitoring and automation charges and is optimized for unknown or changing access patterns, not predictable long-term archival. Option C is wrong because S3 Standard is the most expensive class and is meant for frequently accessed data, making 7-year retention prohibitively costly. Option D is wrong because S3 Glacier Flexible Retrieval costs more than Deep Archive and offers faster retrieval (minutes to hours), which is unnecessary for compliance logs that are rarely accessed.

397
MCQeasy

A company wants to centralize logging of all API calls made within their AWS account for auditing. Which service should they use?

A.Amazon S3 access logs
B.AWS CloudTrail
C.VPC Flow Logs
D.Amazon CloudWatch Logs
AnswerB

AWS CloudTrail is the native AWS service that records all management-plane API calls made by IAM users, roles, and AWS services, along with data-plane events for supported services. Each event includes the identity of the caller, the source IP address, the request parameters, and the response, enabling a complete audit trail. CloudTrail can be configured with a multi-region trail or an organization trail to centralize logging across all accounts and regions, making it the correct choice for centralized API call logging.

Why this answer

AWS CloudTrail is the service designed to log all API calls made within an AWS account, providing a detailed audit trail of actions taken by users, roles, and services. It records API activity across the AWS Management Console, SDKs, CLI, and other services. CloudTrail logs can be delivered to S3 and CloudWatch Logs for analysis.

Exam trap

DOP-C02 often tests the difference between CloudTrail and other logging services, and candidates might confuse VPC Flow Logs or S3 access logs with API auditing, but the trap is not recognizing CloudTrail's scope.

How to eliminate wrong answers

Option A is wrong because Amazon S3 access logs only record requests made to S3 buckets, not all API calls across the account. Option C is wrong because VPC Flow Logs capture IP traffic metadata for network interfaces, not API calls. Option D is wrong because Amazon CloudWatch Logs is a log storage and analysis service, but it does not itself capture API calls; it can receive CloudTrail logs but is not the source.

398
Multi-Selectmedium

A DevOps engineer is designing a CI/CD pipeline for a microservices architecture. The pipeline must ensure that only code that passes security scanning can proceed to deployment. Which TWO actions should the engineer take? (Choose TWO.)

Select 2 answers
A.Use Amazon CloudWatch Events to trigger a rollback if vulnerabilities are found after deployment.
B.Add a security scanning stage in the pipeline after the build stage and before the deploy stage.
C.Use AWS CodeDeploy to perform security scanning during deployment.
D.Configure the pipeline to run security scanning only in the deploy stage.
E.Configure the pipeline to fail if the security scanning stage returns a non-zero exit code.
AnswersB, E

Placing a security scanning stage immediately after the build stage and before the deploy stage makes the scan a quality gate on the immutable artifact itself, so deployment only proceeds when the artifact is clean. Tools like Amazon Inspector, Trivy, SonarQube, or Snyk can run in a CodeBuild action within CodePipeline to produce reports and block promotion. This 'shift-left' approach catches vulnerabilities while remediation is cheapest and before any environment is provisioned or exposed.

Why this answer

Integrating a security scanning stage after the build stage and before the deploy stage ensures that only code that has passed security checks proceeds to deployment. This aligns with the principle of shifting security left in the CI/CD pipeline, preventing vulnerable artifacts from reaching production environments.

Exam trap

The trap here is that candidates often confuse post-deployment monitoring (CloudWatch Events) with pre-deployment gating, or mistakenly think AWS CodeDeploy can perform security scanning, when in fact it only handles deployment orchestration.

399
MCQhard

A company uses AWS Config to evaluate compliance of their AWS resources. They have a custom rule that checks whether EC2 instances have a specific tag. They notice that the rule is not triggering on existing instances. What is a possible reason?

A.The rule is not configured with a trigger type of 'Configuration changes' or 'Periodic'
B.AWS Config does not support custom rules
C.The Lambda function does not have permission to describe EC2 instances
D.The EC2 instances are not in the resource types being recorded by AWS Config
AnswerA

An AWS Config rule, whether managed or custom, must be associated with at least one trigger type — `Configuration changes` or `Periodic` — to initiate evaluations. Without such a trigger, the rule never runs, leaving all resources in a `Not evaluated` state because AWS Config has no basis to invoke the rule's evaluation logic. In the AWS Management Console, the rule would show no compliance results, which exactly matches the described symptom. Adding a configuration-change trigger (for EC2 instances) or a periodic schedule (e.g., every 24 hours) is the necessary fix.

Why this answer

AWS Config custom rules require a trigger type to evaluate resources. If a rule is not configured with either 'Configuration changes' (triggered when a resource changes) or 'Periodic' (triggered on a schedule), it will never evaluate resources, including existing instances. Without a trigger, the rule remains inactive and cannot perform compliance checks.

Exam trap

The trap here is that candidates often assume a custom rule automatically evaluates all resources upon creation, but AWS Config requires an explicit trigger type to initiate evaluation, and without it, the rule remains dormant.

How to eliminate wrong answers

Option B is wrong because AWS Config fully supports custom rules via AWS Lambda functions, which can evaluate resources against custom logic. Option C is wrong because insufficient Lambda permissions would cause evaluation failures or errors, not prevent the rule from triggering on existing instances; the rule would still attempt to run but fail. Option D is wrong because if EC2 instances are not in the resource types being recorded, AWS Config would not track them at all, but the question states the rule is not triggering on existing instances, implying they are recorded; the core issue is the missing trigger type.

400
MCQmedium

A company manages its infrastructure using AWS CloudFormation. They have a production stack that includes an Amazon RDS Multi-AZ DB instance. The stack was created using the 'aws cloudformation create-stack' command with default settings. The DB instance uses a custom DB parameter group. A DevOps engineer needs to modify a parameter in the DB parameter group and update the stack. The engineer updates the template to change the parameter value and runs 'aws cloudformation update-stack'. The update fails with a 'ROLLBACK_IN_PROGRESS' status. The engineer checks the CloudFormation console and sees that the DB instance was successfully modified, but the stack is rolling back. The rollback fails because the DB instance cannot be reverted to the original parameter value. The stack is now in 'UPDATE_ROLLBACK_FAILED' state. What should the engineer do to resolve this situation and apply the desired parameter change?

A.Run 'aws cloudformation update-stack' again with the original template to revert the changes.
B.Use the 'aws cloudformation continue-update-rollback' command with the '--resources-to-skip' parameter to skip the DB instance, allowing the stack to reach 'UPDATE_ROLLBACK_COMPLETE'. Then apply a change set with the desired parameter change.
C.Revert the parameter value manually in the RDS console and then resume the rollback.
D.Delete the stack and recreate it with the updated template.
AnswerB

The continue-update-rollback command is the designed mechanism to recover from a failed rollback. By specifying --resources-to-skip, you tell CloudFormation to ignore the RDS DB instance that is blocking the rollback, allowing the stack to transition to UPDATE_ROLLBACK_COMPLETE. Once stable, you can use a change set to reapply the desired parameter change in a controlled, reversible way, ensuring the stack's template and live resources are aligned.

Why this answer

When a CloudFormation stack is in UPDATE_ROLLBACK_FAILED state, the `continue-update-rollback` command with `--resources-to-skip` allows you to skip the resource that cannot be rolled back (the RDS DB instance with the custom parameter group). This moves the stack to UPDATE_ROLLBACK_COMPLETE, after which you can apply a change set with the desired parameter change. This approach avoids manual intervention or stack deletion while preserving the modified DB instance.

Exam trap

The trap here is that candidates may think manual reversion or stack deletion is required, but CloudFormation provides a built-in recovery mechanism (`continue-update-rollback`) that avoids downtime and data loss.

How to eliminate wrong answers

Option A is wrong because running `update-stack` with the original template would attempt to revert the DB parameter group, which already failed to roll back, and would likely fail again or cause further issues. Option C is wrong because manually reverting the parameter value in the RDS console does not resolve the CloudFormation stack's failed rollback state; the stack remains in UPDATE_ROLLBACK_FAILED and cannot resume rollback without CloudFormation's `continue-update-rollback` command. Option D is wrong because deleting and recreating the stack would cause downtime and data loss for the production RDS DB instance, and is unnecessarily destructive when a non-disruptive recovery path exists.

401
Multi-Selectmedium

A company is using AWS CloudFormation to deploy a multi-tier application. The DevOps team wants to ensure that the database password is not exposed in the template or the console. Which two methods should they use to securely manage the password? (Choose TWO.)

Select 2 answers
A.Hardcode the password in the template and use a condition to only apply it in production.
B.Use a CloudFormation parameter with the NoEcho property set to true.
C.Store the password in AWS Systems Manager Parameter Store and reference it with {{resolve:ssm:...}}
D.Use a dynamic reference to AWS Secrets Manager secret in the CloudFormation template.
E.Pass the password as user data to the EC2 instance and encrypt the user data.
AnswersB, D

A CloudFormation parameter with the NoEcho property set to true is a valid way to pass a password because the value is supplied at stack creation or update time—either via the console, CLI, or API—and is never displayed in the AWS Management Console, returned by DescribeStackResources, or emitted in CloudFormation event logs. However, you must still reference the parameter in the template resource properties so that the actual secret value is used during provisioning, and the secret remains visible to anyone with permission to view the resulting resource configurations (e.g., EC2 instance metadata if you pass it to user data). NoEcho only prevents the value from being echoed back in API responses; it does not encrypt the template or the secret at rest, so you must also ensure IAM policies restrict access to the stack's parameters. This is the simplest built-in mechanism for basic secret handling, but for more robust secret rotation and management, a dynamic reference or service like Secrets Manager is preferred.

Why this answer

Setting the `NoEcho` property to `true` on a CloudFormation parameter masks the password from console outputs and `DescribeStack` calls, preventing exposure in logs or the AWS Management Console. This is a straightforward way to handle sensitive input without external services, though it does not encrypt the value at rest in the template.

Exam trap

The trap here is that candidates often confuse `NoEcho` with encryption, thinking it secures the value at rest, when it only masks output; or they incorrectly assume that Systems Manager Parameter Store's `{{resolve:ssm:...}}` syntax works universally in CloudFormation, when it requires the `ssm-secure` variant and has property-specific limitations.

402
MCQmedium

A company uses an Application Load Balancer (ALB) to distribute traffic to EC2 instances. The ALB is in us-east-1a and us-east-1b. They want to ensure that if one AZ fails, traffic is routed only to healthy instances in the other AZ. What configuration is necessary?

A.Enable sticky sessions (session affinity)
B.Configure health checks on the target group
C.Add more subnets in additional AZs
D.Enable cross-zone load balancing on the ALB
AnswerD

Cross-zone load balancing on the ALB enables each load balancer node to distribute incoming traffic evenly across all healthy targets in all enabled AZs, rather than limiting each node to its own AZ. With this enabled, if an AZ becomes unhealthy or unreachable, the remaining healthy ALB nodes can seamlessly route traffic to healthy targets in other AZs, ensuring continuous availability. This directly addresses the root cause of an ALB failing to route to instances in other AZs during an outage.

Why this answer

Cross-zone load balancing must be enabled on the ALB so that traffic can be distributed across instances in all AZs. By default, an ALB routes requests only to targets in the same Availability Zone as the requesting client. Enabling cross-zone load balancing allows the ALB to distribute traffic evenly across all registered targets in all enabled AZs, ensuring that if one AZ fails, traffic can be routed to healthy instances in other AZs.

Option B is incorrect because health checks are already enabled by default and do not affect cross-AZ routing. Option C is incorrect because adding more AZs does not change the default AZ-affinity behavior; cross-zone load balancing must be explicitly enabled to utilize multiple AZs for failover.

403
Multi-Selecthard

A company runs a web application on Amazon EC2 instances behind an Application Load Balancer (ALB). The application logs show that some requests are timing out. The team needs to identify the source of the issue. Which TWO steps should they take?

Select 2 answers
A.Enable ALB access logs and analyze them.
B.Enable VPC Flow Logs to capture network traffic.
C.Enable AWS WAF logs to inspect HTTP requests.
D.Review CloudWatch metrics for the ALB, such as 'RequestCount' and 'TargetResponseTime'.
E.Enable AWS CloudTrail to log all API calls.
AnswersA, D

ALB access logs are the authoritative source for request-level diagnostics because they capture every HTTP request processed by the load balancer, including the target processing time, request processing time, HTTP status, and client/user-agent data. Analyzing these logs lets you identify exactly which requests experienced slow responses or timeouts, isolate problematic targets by IP or URL pattern, and correlate with backend health. Unlike aggregated metrics, access logs provide per-request granularity that is essential for root-causing intermittent timeout issues.

Why this answer

Option A is correct because ALB access logs capture detailed per-request information such as request processing time, target response time, and the specific error codes (e.g., 504 Gateway Timeout) returned by the load balancer, which directly helps pinpoint whether timeouts originate at the ALB or the backend targets. Option D is correct because CloudWatch metrics for the ALB, particularly 'TargetResponseTime' and 'RequestCount', reveal latency trends and traffic patterns that indicate whether targets are slow or overloaded, helping isolate the source of the timeouts. Option B is not appropriate because VPC Flow Logs only capture IP-level metadata (source/destination, ports, accept/reject) and cannot show HTTP-layer timing or application-level errors.

Option C is not appropriate because AWS WAF logs only record requests that match or are blocked by WAF rules, and the scenario does not indicate WAF is in use or that requests are being blocked. Option E is not appropriate because CloudTrail records AWS API calls for auditing, not application request behavior or latency.

Exam trap

DOP-C02 often tests the distinction between logging services: candidates may confuse VPC Flow Logs (network-level) with ALB access logs (application-level) or assume CloudTrail captures performance data, leading to wrong selections.

404
MCQmedium

A development team uses AWS CodeCommit as a source control repository. A developer accidentally pushed a commit that contains sensitive information (e.g., AWS access keys) to the main branch. The team wants to remove the sensitive data from the repository history completely. Which action should the engineer take?

A.Use 'git filter-branch' to rewrite the repository history and remove the sensitive file
B.Delete the repository and create a new one, then force push the remaining branches
C.Use 'git revert' to create a new commit that undoes the changes
D.Create a new branch from the commit before the sensitive data was added and merge it to main
AnswerA

git filter-branch rewrites every commit in the repository's DAG, eliminating the sensitive blob from historical snapshots and changing commit SHAs. Once the rewritten history is force-pushed to CodeCommit, the file is unreachable via any prior commit, but all branch tips must be updated and team members must re-clone or rebase to avoid propagating the old history. This is the standard, targeted purging technique for leaked credentials.

Why this answer

'git filter-branch' (or the modern 'git filter-repo') rewrites the repository history by removing or replacing the sensitive file in every commit, effectively purging it from the entire Git history. This is the only native Git method that completely eliminates the sensitive data from all past commits, preventing anyone from retrieving it via 'git log' or by cloning the repository. After rewriting history, a force push to the remote CodeCommit repository is required to overwrite the remote branches.

Exam trap

The trap here is that candidates confuse 'git revert' (which adds a new commit but leaves the sensitive data in history) with 'git filter-branch' (which actually rewrites history to remove the data), leading them to choose a non-destructive but ineffective option.

How to eliminate wrong answers

Option B is wrong because deleting the repository and creating a new one, then force pushing remaining branches, does not remove the sensitive data from the existing repository's history on the remote; the old repository would still exist in CodeCommit's trash or backup, and the sensitive data would remain accessible. Option C is wrong because 'git revert' creates a new commit that undoes the changes of a previous commit, but the sensitive data remains in the commit history and can still be viewed with 'git log' or by checking out the old commit. Option D is wrong because creating a new branch from the commit before the sensitive data was added and merging it to main does not remove the commit containing the sensitive data from the history; the merge will still include the sensitive commit in the ancestry, and the data remains accessible.

405
MCQhard

Refer to the exhibit. An alarm is configured as shown. The CPU utilization averages 85% for 10 minutes, then spikes to 95% for the next 5 minutes, and returns to 80%. How many times will the SNS topic receive a notification?

A.0
B.1
C.2
D.3
AnswerA

The CloudWatch alarm is configured with `EvaluationPeriods: 3` and `DatapointsToAlarm: 2`. During the monitoring window, the CPU utilization never produces two breaching datapoints within any three consecutive periods, so the alarm condition remains unsatisfied. Consequently, the alarm never leaves the OK state and thus never enters ALARM.

Why this answer

Per the exhibit: Period=300s (5 min), EvaluationPeriods=2, Threshold=90.0, ComparisonOperator=GreaterThanThreshold, no DatapointsToAlarm override (so 2 consecutive breaching 5-minute periods are required to enter ALARM). CPU averages 85% for two 5-minute periods (not breaching, since 85 is not > 90), then 95% for one 5-minute period (breaching), then returns to 80% (not breaching). The single breaching period is immediately preceded and followed by non-breaching periods, so the alarm never accumulates 2 consecutive breaches and never transitions into ALARM -- it stays in OK the entire time.

Since neither AlarmActions nor OKActions fire without an actual state transition, and none occurs, the SNS topic receives zero notifications.

Exam trap

The trap here is that candidates assume the alarm triggers immediately when the metric exceeds the threshold, but CloudWatch requires a specified number of consecutive evaluation periods (datapoints) to breach before changing state, and the threshold comparison is strict (greater than, not greater than or equal).

How to eliminate wrong answers

Option A (0) is wrong because the alarm does trigger after 3 consecutive periods above the threshold, sending one notification. Option C (2) is wrong because only one notification is sent when the alarm enters ALARM state; no notification is sent when it returns to OK because the scenario ends before the 3-period requirement for OK is met. Option D (3) is wrong because there is no repeated flapping or multiple state changes; the alarm transitions only once from OK to ALARM.

406
MCQmedium

An organization has a compliance requirement to automatically detect and alert on any IAM user creation in all AWS accounts. Which combination of services should be used to meet this requirement?

A.Amazon GuardDuty and Amazon SNS
B.Amazon S3 server access logs and Amazon Athena
C.AWS Config and AWS Lambda
D.AWS CloudTrail and Amazon CloudWatch Events
AnswerD

CloudTrail is the authoritative audit service that records every management API call, including the IAM CreateUser event, along with details such as the requesting IAM principal, source IP address, and timestamps. You can configure CloudWatch Events (or Amazon EventBridge) with an event pattern that matches `"eventSource": "iam.amazonaws.com"` and `"eventName": "CreateUser"`, and route matching events to an SNS topic or Lambda function for immediate notification. This provides the real-time, event-driven alerting that directly satisfies a compliance requirement to automatically detect and respond to IAM user creation.

Why this answer

AWS CloudTrail captures all IAM user creation events as `CreateUser` API calls. Amazon CloudWatch Events (now Amazon EventBridge) can be configured with a rule that matches this specific event pattern and triggers an alert via Amazon SNS. This combination provides real-time detection and notification without custom code.

Exam trap

The trap here is that candidates often confuse AWS Config (which evaluates resource configurations) with CloudTrail (which records API activity), leading them to select Option C, but AWS Config cannot trigger alerts on API call events like `CreateUser`; it only reacts to configuration changes after they have occurred.

How to eliminate wrong answers

Option A is wrong because Amazon GuardDuty is a threat detection service that analyzes VPC flow logs, DNS logs, and CloudTrail management events for malicious activity, but it does not provide a native mechanism to trigger custom alerts on specific IAM user creation events; it focuses on anomalies and threats, not compliance-driven event monitoring. Option B is wrong because Amazon S3 server access logs record requests made to an S3 bucket, not IAM user creation events, and using Athena to query them would require a separate mechanism to capture CloudTrail logs into S3, adding latency and complexity; this approach is not designed for real-time alerting on IAM actions. Option C is wrong because AWS Config evaluates resource configurations against rules and can detect changes, but it is not designed for real-time event-driven alerting on API calls; it operates on configuration snapshots and compliance evaluations, not on streaming API events like `CreateUser`.

407
MCQmedium

An organization uses AWS CodePipeline to deploy a web application. The pipeline includes a test stage that runs integration tests using AWS CodeBuild. The tests are flaky and sometimes fail due to external dependencies. The team wants to automatically retry failed tests before marking the stage as failed. How should this be achieved?

A.Add a manual approval step after the test stage.
B.Use Amazon CloudWatch Events to listen for test failures and trigger a new pipeline execution.
C.Configure the CodeBuild project to automatically retry the build on failure.
D.Create a second pipeline that triggers only on test failures.
AnswerC

CodeBuild projects support an automatic retry configuration that re-runs a failed build when the build action is executed by CodePipeline. By setting a retry limit (up to 10 attempts) and an optional execution timeout, the CodeBuild service itself retries the exact same build spec, allowing flaky integration tests to succeed on a subsequent attempt without any pipeline-level changes. The build action in CodePipeline reports a single success/failure status based on the final attempt, so a successful retry lets the stage pass without re-running earlier source or build stages. This is the only option that provides an automatic, stage-local retry mechanism for the build, directly addressing the flaky tests.

Why this answer

AWS CodeBuild natively supports automatic retries on build failure through the 'auto retry limit' configuration. By setting this limit (e.g., 3), CodeBuild will automatically re-run the build if it fails, which directly addresses flaky tests without requiring additional pipeline stages or external event handling. This keeps the retry logic within the same pipeline execution, ensuring that the test stage only fails after exhausting all retry attempts.

Exam trap

The trap here is that candidates may overcomplicate the solution by thinking they need external event-driven retries (Option B) or separate pipelines (Option D), when CodeBuild's built-in retry configuration directly solves the problem with minimal overhead.

How to eliminate wrong answers

Option A is wrong because adding a manual approval step after the test stage does not retry failed tests; it only pauses the pipeline for human intervention, which does not automate the retry process and introduces unnecessary delay. Option B is wrong because using Amazon CloudWatch Events to trigger a new pipeline execution on test failures would start a completely separate pipeline run, not retry the current stage within the same execution, leading to duplicate work and potential race conditions. Option D is wrong because creating a second pipeline that triggers only on test failures is overly complex and redundant; it does not retry the test stage in the original pipeline and would require additional orchestration to manage state between pipelines.

408
Drag & Dropmedium

Drag and drop the steps to troubleshoot a failed deployment in AWS CodeDeploy into the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

Troubleshooting starts with console, then agent logs, then AppSpec, then instance configuration, then redeploy.

409
Multi-Selectmedium

A company is implementing a CI/CD pipeline using AWS CodePipeline. The pipeline has a source stage from GitHub, a build stage using AWS CodeBuild, and a deploy stage using AWS Elastic Beanstalk. The team wants to ensure that the pipeline only proceeds if the code quality checks pass and unit tests are successful. Which TWO actions should be taken?

Select 2 answers
A.Add a test stage in the pipeline with a CodeBuild action that runs code quality and unit tests.
B.Add a manual approval step before the deploy stage.
C.Modify the buildspec file of the build stage to include test commands and fail on test failures.
D.Configure the source stage to use an S3 bucket and add a test action.
E.Use AWS CloudFormation to create a test environment and run tests.
AnswersA, C

Adding a dedicated test stage in CodePipeline with a CodeBuild action is the canonical pattern for automated validation: after the source is retrieved and built, the pipeline invokes a CodeBuild project whose buildspec runs code quality checks and unit tests. Because the stage fails the pipeline if any test exits non-zero, this creates an explicit quality gate that must pass before the build artifact proceeds to deployment, and CodePipeline's stage boundaries give you clear visibility into test results and execution history.

Why this answer

Adding a dedicated test stage with a CodeBuild action allows the pipeline to explicitly run code quality checks and unit tests as a separate, visible step. This ensures the pipeline only proceeds to deployment if these tests pass, as CodeBuild can be configured to fail the action on non-zero exit codes from test commands. Option C is also correct because modifying the buildspec file in the existing build stage to include test commands and setting the build to fail on test failures integrates quality gates directly into the build process, which is a common and valid approach for enforcing test success before deployment.

Exam trap

The trap here is that candidates may think a manual approval step (Option B) or infrastructure provisioning (Option E) can enforce test quality gates, but neither actually executes or validates test results automatically within the pipeline.

410
MCQeasy

A company uses AWS KMS to encrypt data in S3. They want to audit who used which KMS key and when. Which AWS service should they use?

A.Amazon CloudWatch
B.Amazon GuardDuty
C.AWS CloudTrail
D.AWS Config
AnswerC

AWS CloudTrail is the correct answer because it records KMS API calls as data events, providing the principal, key ID, source IP address, and timestamp for each operation. When data events are enabled for a customer master key, CloudTrail captures every Encrypt, Decrypt, GenerateDataKey, and ScheduleKeyDeletion call, which is exactly what is needed to audit encryption usage. This KMS activity is delivered as a JSON event to an S3 bucket (and optionally to CloudWatch Logs), forming a durable, tamper-evident audit trail for compliance and security investigations.

Why this answer

AWS CloudTrail is the correct service because it records all AWS KMS API calls, including the key ID, the principal who made the request, the time of the request, and the source IP address. These logs are delivered to an S3 bucket and can be queried using CloudTrail Insights or Athena to audit KMS key usage for S3 decryption events.

Exam trap

The trap here is that candidates often confuse CloudWatch Logs (which can store logs) with CloudTrail (which captures the API audit trail), leading them to pick CloudWatch because they think 'audit logs' are just logs, but only CloudTrail records the specific KMS API calls needed for key usage auditing.

How to eliminate wrong answers

Option A is wrong because Amazon CloudWatch is a monitoring service for metrics, alarms, and logs, but it does not natively capture the detailed API-level audit trail of KMS key usage; it can only visualize CloudTrail events if they are streamed to it. Option B is wrong because Amazon GuardDuty is a threat detection service that analyzes DNS, VPC flow logs, and CloudTrail events for malicious activity, but it is not designed to provide a direct audit log of who used which KMS key and when. Option D is wrong because AWS Config is a resource inventory and compliance service that tracks configuration changes to AWS resources, not the API calls that use KMS keys for encryption or decryption operations.

411
MCQeasy

A CloudFormation template snippet is shown. An engineer attempts to create a stack with this template and receives an error: 'Bucket my-unique-bucket-name already exists'. What is the most likely cause?

A.The bucket policy has a syntax error that prevents the bucket from being created.
B.The S3 bucket name 'my-unique-bucket-name' is already taken by another AWS account.
C.The versioning configuration is incompatible with the bucket policy.
D.The bucket policy references the bucket name incorrectly, causing a circular dependency.
AnswerB

S3 bucket names exist in a single global namespace shared by every AWS account and region. Once any account has registered 'my-unique-bucket-name', a CloudFormation CreateBucket call in your account fails with HTTP 409 BucketAlreadyExists, matching the symptom described. Because the error is explicitly tied to the bucket name, the most probable root cause is that another AWS account has already claimed that globally unique string, not a template configuration issue.

Why this answer

S3 bucket names must be globally unique across all AWS accounts and regions. The error 'Bucket my-unique-bucket-name already exists' indicates that the name is already taken by another AWS account, not that the bucket already exists in the current account. CloudFormation cannot create the bucket because the name is not available in the global S3 namespace.

Exam trap

The trap here is that candidates may assume the error refers to a bucket already existing in their own account, but AWS S3 enforces global uniqueness, so the error always means the name is taken by any account in the entire AWS ecosystem.

How to eliminate wrong answers

Option A is wrong because a syntax error in the bucket policy would cause a different validation error (e.g., 'Malformed policy') during stack creation, not a 'Bucket already exists' error. Option C is wrong because versioning configuration and bucket policy are independent settings; incompatibility between them would not produce a 'Bucket already exists' error—it would cause a separate validation or update failure. Option D is wrong because a circular dependency would cause a stack creation failure with a 'Circular dependency' error message, not a 'Bucket already exists' error.

412
MCQeasy

A company's application runs on EC2 instances in a single Availability Zone. The operations team wants to improve resilience without redesigning the application. Which action is the MOST effective?

A.Use a larger instance type to handle more traffic.
B.Enable EC2 Auto Recovery to automatically restart the instance if it fails.
C.Deploy EC2 instances across multiple Availability Zones using an Auto Scaling group.
D.Place the instance in a placement group to ensure low latency.
AnswerC

Deploying EC2 instances across multiple Availability Zones with an Auto Scaling group is the standard pattern for high availability. If one AZ becomes unavailable, the load balancer routes traffic to instances in the remaining healthy AZs, and the Auto Scaling group maintains instance count across AZs to replace any that are terminated. This eliminates the single-AZ point of failure and ensures application availability during an AZ outage.

Why this answer

Deploying EC2 instances across multiple Availability Zones (AZs) using an Auto Scaling group is the most effective action because it eliminates the single point of failure at the AZ level. If one AZ experiences an outage, the Auto Scaling group automatically launches replacement instances in the remaining healthy AZs, ensuring application availability without requiring any application-level changes. This directly addresses the goal of improving resilience by leveraging AWS's fault-isolated infrastructure.

Exam trap

The trap here is that candidates often confuse instance-level recovery (Auto Recovery) with infrastructure-level resilience (multi-AZ deployment), mistakenly thinking that restarting a failed instance in the same AZ provides sufficient protection against the most common cause of downtime—an AZ outage.

How to eliminate wrong answers

Option A is wrong because using a larger instance type only increases compute capacity, not resilience; a single AZ failure still takes down all instances regardless of size. Option B is wrong because EC2 Auto Recovery only recovers an instance within the same AZ if the underlying hardware fails, but it does not protect against an entire AZ outage, which is the primary risk. Option D is wrong because a placement group is designed to reduce network latency by ensuring instances are in close proximity, but it actually increases the risk of correlated failures and does not improve resilience against AZ-level failures.

413
MCQeasy

A company's production environment uses an Amazon ElastiCache Redis cluster for session caching. The operations team reports that the cache hit ratio has dropped significantly, causing increased load on the backend database. What is the MOST likely cause?

A.The cache is under memory pressure and evicting keys to make room.
B.The cluster was resized from a single node to a cluster mode.
C.The encryption in transit was enabled, adding latency.
D.There is a network partition between the application and the cache.
AnswerA

When an ElastiCache cluster reaches its maxmemory limit, the eviction policy (e.g., allkeys-lru or volatile-lru) kicks in to free space by deleting keys. Evicted keys disappear before they can serve valid reads, so subsequent lookups for those keys become cache misses, directly reducing the CacheHitRate metric. This is the most consistent explanation for a sustained drop in hit ratio: the cache is actively discarding data to accommodate new writes, and the eviction is a continuous symptom of memory pressure, not a one-time event.

Why this answer

A drop in cache hit ratio with increased backend load is most commonly caused by memory pressure leading to eviction of keys. When ElastiCache Redis reaches maxmemory, it evicts keys according to the eviction policy (e.g., LRU), causing more cache misses and higher database load.

Exam trap

DOP-C02 often tests the difference between symptoms of memory pressure (evictions, low hit ratio) and other issues like network partitions or encryption, tricking candidates into selecting configuration changes that do not address the root cause.

How to eliminate wrong answers

Option B is wrong because resizing to cluster mode generally increases capacity and does not inherently reduce hit ratio; it might require reconfiguration but not eviction. Option C is wrong because enabling encryption in transit adds minor latency but does not cause cache misses or reduce hit ratio. Option D is wrong because a network partition would cause complete connectivity loss, not a gradual drop in hit ratio; it would result in errors, not just misses.

414
Multi-Selecteasy

A company uses AWS Lambda for data processing. The operations team wants to be alerted when a function fails. Which TWO methods can they use?

Select 2 answers
A.Configure S3 event notifications to trigger on Lambda errors.
B.Enable AWS CloudTrail to log Lambda invocations.
C.Configure a dead-letter queue (DLQ) for the Lambda function and monitor the queue.
D.Create a CloudWatch alarm on the 'Errors' metric for the Lambda function.
E.Use AWS Config to detect Lambda function failures.
AnswersC, D

For asynchronous invocations, Lambda can be configured with a dead-letter queue (an SQS queue or SNS topic) to receive event payloads that could not be processed after a function fails or exhausts retry attempts. The DLQ preserves the exact original event data, allowing you to inspect, replay, or archive failed records in a durable buffer. Monitoring the DLQ (for example, with a CloudWatch alarm on SQS ApproximateNumberOfMessagesVisible) directly alerts you to the presence of undelivered events. This is why the correct answer combines a DLQ with active monitoring of that queue.

Why this answer

Option D is correct because Lambda automatically publishes the 'Errors' metric to Amazon CloudWatch, and a CloudWatch alarm can be created on that metric to trigger an Amazon SNS notification when the error threshold is breached, directly alerting the operations team. Option C is correct because configuring a dead-letter queue (an Amazon SQS queue or SNS topic) for the Lambda function captures failed asynchronous invocations, and monitoring that queue (e.g., via CloudWatch metrics like ApproximateNumberOfMessagesVisible) provides an alerting mechanism for failures. Option A is incorrect because S3 event notifications only trigger Lambda invocations on object events; they do not detect or report Lambda execution errors.

Option B is incorrect because CloudTrail records API activity and management events, not Lambda function invocation failures, and it is not an alerting service. Option E is incorrect because AWS Config evaluates resource configuration compliance, not runtime function failures.

Exam trap

DOP-C02 often tests the confusion between monitoring services (CloudWatch) and logging/auditing services (CloudTrail, AWS Config), so candidates must recognize that only CloudWatch alarms and DLQ monitoring provide direct alerting on Lambda failures.

415
MCQmedium

A company wants to automate the rotation of IAM user access keys every 90 days. Which AWS service should be used to implement this rotation?

A.AWS Lambda with custom rotation logic
B.AWS Config
C.AWS Secrets Manager
D.AWS Systems Manager Parameter Store
AnswerA

Correct. AWS Lambda with custom rotation logic is the required approach since no other AWS service natively rotates IAM user access keys. You can write a Lambda function to generate new keys, update the IAM user, and handle the rotation schedule.

Why this answer

AWS Lambda with custom rotation logic is the correct service to automate IAM user access key rotation because AWS does not provide a native service that automatically rotates IAM access keys. AWS Secrets Manager can rotate secrets for databases and other services, but it does not support rotating IAM user access keys. Therefore, the recommended approach is to implement a custom Lambda function that generates new keys, updates the user, and manages the rotation lifecycle.

Option B (AWS Config) only monitors compliance, not credentials. Option D (Systems Manager Parameter Store) stores secrets but lacks rotation capabilities for IAM keys.

416
Multi-Selecteasy

A company wants to ensure that its application running on AWS can withstand the failure of an entire AWS Region. Which TWO strategies should the company implement?

Select 2 answers
A.Deploy the application in multiple AWS Regions using an active-active or active-passive pattern
B.Deploy the application across multiple Availability Zones in a single Region
C.Replicate data across Regions using services like DynamoDB global tables or RDS cross-Region replication
D.Use a single CloudFront distribution with multiple origins in the same Region
E.Configure RDS read replicas in the same Region
AnswersA, C

Running the application in multiple AWS Regions using an active-active or active-passive pattern is the foundational disaster recovery strategy for a regional outage. Active-active routes live traffic across Regions for automatic failover, while active-passive keeps a warm standby in another Region that can be promoted via Route 53 health checks and failover policies. This approach directly addresses the failure domain of an entire Region.

Why this answer

Option A is correct because deploying the application in multiple AWS Regions using an active-active or active-passive pattern ensures that if an entire Region fails, traffic can be served from another Region, providing true Region-level fault tolerance. Option C is correct because replicating data across Regions with services like DynamoDB global tables or RDS cross-Region replication ensures the application's data is available in the secondary Region, which is essential for the failover strategy in Option A to work. Option B is incorrect because multiple Availability Zones within a single Region only protect against AZ-level failures, not the failure of an entire Region.

Option D is incorrect because a single CloudFront distribution with multiple origins in the same Region still depends on that one Region and does not survive a Region-wide outage. Option E is incorrect because RDS read replicas in the same Region remain within that Region and would be lost if the entire Region failed.

Exam trap

DOP-C02 often tests the distinction between multi-AZ (single-Region HA) and multi-Region (disaster recovery); candidates select multi-AZ answers because they sound resilient but do not survive a Region outage.

417
MCQhard

A DevOps engineer is troubleshooting a CloudFormation stack that fails to create. The error message indicates a 'circular dependency' between two resources: a security group and an EC2 instance. The security group contains an ingress rule that references the instance's private IP address, which is not known until the instance is created. The instance's network interface uses the security group. What change should the engineer make to resolve the circular dependency?

A.Create the EC2 instance first without a security group, then attach the security group after creation.
B.Add an AWS::EC2::SecurityGroupIngress rule that references the instance's network interface using Fn::GetAtt on the network interface resource.
C.Hardcode the instance's private IP address in the security group rule.
D.Use the Ref function on the EC2 instance to get its private IP address.
AnswerB

It breaks the circular dependency by using `Fn::GetAtt` on the `AWS::EC2::NetworkInterface` resource to reference the private IP address. This creates a dependency on the network interface, which is created before the instance, while the security group ingress rule depends on the network interface, breaking the cycle.

Why this answer

It resolves the circular dependency by creating an explicit dependency on the network interface resource rather than the EC2 instance. The `AWS::EC2::SecurityGroupIngress` rule can use `Fn::GetAtt` on the `AWS::EC2::NetworkInterface` resource to retrieve the private IP address of the instance's primary network interface, which is known after the network interface is created but before the instance is fully launched. This breaks the cycle because the security group ingress rule depends on the network interface, and the network interface depends on the security group (via association), but the instance itself is not directly referenced in the ingress rule, allowing CloudFormation to resolve the dependencies in the correct order.

Exam trap

The trap here is that candidates often assume `Ref` on an EC2 instance returns its private IP address, but `Ref` actually returns the physical instance ID (e.g., i-1234567890abcdef0), not the IP, leading them to incorrectly choose Option D or attempt hardcoding in Option C.

How to eliminate wrong answers

Option A is wrong because creating the EC2 instance without a security group and attaching it later does not resolve the circular dependency in the CloudFormation template; it merely shifts the problem to a post-creation step that still requires the private IP address, and the template itself would still fail due to the unresolved dependency during creation. Option C is wrong because hardcoding the instance's private IP address defeats the purpose of Infrastructure as Code (IaC) and is not dynamic; it would break if the instance is recreated or if the IP changes, and it does not solve the circular dependency in the template logic. Option D is wrong because using the `Ref` function on the EC2 instance returns the logical ID (e.g., the instance ID), not the private IP address; `Ref` on an EC2 instance does not expose the private IP, and even if it did, it would create the same circular dependency because the security group ingress rule would still depend on the instance's creation.

418
MCQhard

A company runs a critical e-commerce platform on AWS. The architecture includes an Application Load Balancer (ALB) that distributes traffic to a fleet of EC2 instances in an Auto Scaling group across three Availability Zones. The instances run a Java application that connects to an Amazon RDS Multi-AZ MySQL database. The application also uses Amazon ElastiCache for Redis for session caching. The company recently experienced a severe outage where the ALB's 5xx error rate spiked to 100% for 45 minutes. The root cause was a combination of a slow-running query on the RDS primary instance and a subsequent failover that caused the application to lose connections to the database. The failover happened because the slow query caused the primary to become unresponsive, triggering a Multi-AZ failover. During the failover, the application's connection pool exhausted, and new connections failed. The application logs show a high rate of 'java.sql.SQLTimeoutException' and 'com.mysql.cj.exceptions.CJCommunicationsException'. The DevOps team needs to implement a long-term solution that minimizes the impact of similar incidents. The solution must be cost-effective and require minimal application changes. Which combination of actions should the DevOps team take?

A.Implement Amazon RDS Proxy to manage database connections and add read replicas to offload read traffic.
B.Use an Auto Scaling policy for EC2 based on RDS connection count and implement a read replica for the primary.
C.Configure Multi-AZ RDS with a synchronous standby and use Amazon RDS for MySQL with enhanced monitoring.
D.Increase the instance size of the RDS primary and enable Performance Insights to identify slow queries.
AnswerA

Amazon RDS Proxy sits between the application and the database, maintaining a warm connection pool that absorbs the spike in connection requests when EC2 instances reconnect during a failover. Because the proxy keeps connections to the RDS instance open and multiplexes client sessions, the primary no longer gets overwhelmed by thousands of short-lived connections. Adding read replicas moves read-heavy queries off the primary, reducing CPU/IO contention that can cause slow queries and cascading failovers. This directly addresses the root cause of connection exhaustion while preserving write consistency on the primary.

Why this answer

Amazon RDS Proxy is the correct solution because it efficiently manages database connection pooling, reducing the likelihood of connection exhaustion during failovers. By maintaining a warm connection pool and automatically reconnecting to the new primary after a Multi-AZ failover, RDS Proxy minimizes application-side connection timeouts and errors like SQLTimeoutException and CJCommunicationsException. Adding read replicas offloads read traffic, reducing the load on the primary and mitigating the risk of slow queries causing unresponsiveness.

This combination requires minimal application changes and is cost-effective compared to scaling the primary instance.

Exam trap

The trap here is that candidates often focus on scaling the database (e.g., increasing instance size or adding read replicas) to fix performance issues, but overlook the critical connection management problem that causes application-level timeouts during failover, which RDS Proxy directly addresses.

How to eliminate wrong answers

Option B is wrong because using an Auto Scaling policy based on RDS connection count does not address the root cause of connection exhaustion during failover; it only scales EC2 instances reactively, which may not prevent timeouts and adds complexity without solving the connection management issue. Option C is wrong because simply configuring Multi-AZ RDS with a synchronous standby and enhanced monitoring does not prevent connection pool exhaustion during failover; the application still needs to manage connections, and enhanced monitoring only provides visibility, not mitigation. Option D is wrong because increasing the instance size of the RDS primary and enabling Performance Insights addresses performance but does not solve the connection management problem during failover; it may delay the issue but does not prevent connection timeouts or exhaustion.

419
MCQhard

A company runs a critical microservice on Amazon ECS with AWS Fargate. The service must be highly available across multiple Availability Zones. The DevOps engineer configured the service with a desired count of 4 tasks spread across 2 Availability Zones. During a deployment, a new task fails to start due to a missing environment variable. The deployment fails, but the old tasks continue to run. What is the most likely cause of the deployment failure and how can the engineer ensure future deployments are resilient?

A.The deployment failed because the ECS service was using the rolling update deployment controller. Change to blue/green deployment.
B.The deployment failed because the ECS service did not have the deployment circuit breaker enabled. Enable the circuit breaker with rollback.
C.The deployment failed because the desired count was too low. Increase the desired count to 6.
D.The deployment failed because the health check grace period was too short. Increase the grace period.
AnswerB

The correct fix is to enable the ECS deployment circuit breaker with rollback. This feature monitors the deployment for indicators such as repeated task launch failures, container exits, and health check failures; when it detects that the deployment is failing beyond a configured threshold, it stops the deployment and automatically restores the service to the most recent successful task set definition. This preserves service availability by keeping the existing stable tasks running and eliminates the need for you to manually roll back a bad task definition. Without the circuit breaker, a deployment with continuously crashing tasks will eventually time out and leave the service in a FAILED state with no automatic recovery.

Why this answer

The deployment failed because the new task could not start due to a missing environment variable, and the ECS service did not have the deployment circuit breaker enabled. Without the circuit breaker, ECS continues to attempt the deployment indefinitely or until a timeout, but it does not automatically roll back to the previous stable task set. Enabling the deployment circuit breaker with rollback ensures that if a specified number of tasks fail to start (e.g., due to health checks or runtime errors), ECS automatically rolls back to the last successful deployment, maintaining service availability.

Exam trap

The trap here is that candidates may focus on the deployment controller type (rolling vs. blue/green) or task count, but the real issue is the lack of automatic rollback capability provided by the deployment circuit breaker, which is specifically designed to handle task startup failures during deployments.

How to eliminate wrong answers

Option A is wrong because the rolling update deployment controller is not the cause of the failure; it is the default and works correctly here by keeping old tasks running. Changing to blue/green deployment would not inherently fix the missing environment variable issue and adds complexity. Option C is wrong because the desired count of 4 tasks is sufficient for high availability across 2 AZs; increasing it to 6 does not address the root cause of task startup failure.

Option D is wrong because the health check grace period only delays the start of health checks, but the task failed to start entirely due to a missing environment variable, not because health checks failed prematurely.

420
MCQhard

A company runs a Stateful application on EC2 that requires sticky sessions. They use an ALB with duration-based stickiness. During a deployment, they want to drain existing connections gracefully before terminating instances. Which step is necessary?

A.Increase the deregistration delay on the target group.
B.Reduce the stickiness duration to zero.
C.Configure health checks to mark instances unhealthy.
D.Enable connection draining on the target group.
AnswerA

ALB implements graceful connection draining through the target group's deregistration delay. Increasing this delay allows in-flight sticky-session requests to complete before the instance is terminated.

Why this answer

ALB target groups use a 'deregistration delay' (formerly called connection draining) to allow in-flight requests to complete before an instance is terminated. This setting is configured on the target group, not the load balancer, and it works with sticky sessions by waiting for the delay period (default 300 seconds) for existing connections to finish, even if the stickiness cookie would otherwise route new requests to the same instance. During a deployment, increasing this delay or ensuring it is set appropriately is the necessary step to gracefully drain connections.

Exam trap

The trap here is that candidates confuse 'connection draining' (which is the deregistration delay on ALB target groups) with 'sticky session duration' or 'health check settings,' thinking that reducing stickiness or marking instances unhealthy alone will gracefully terminate connections, when in fact the deregistration delay is the specific mechanism that waits for in-flight requests to complete.

How to eliminate wrong answers

Option A is wrong because increasing the deregistration delay is not the step that enables draining; the deregistration delay is already the mechanism for connection draining (it is the same setting), but the question asks which step is necessary, and the correct answer is to enable connection draining, which is already the default behavior of the deregistration delay. Option B is wrong because reducing the stickiness duration to zero would disable sticky sessions entirely, breaking the application requirement for sticky sessions, and it does not drain existing connections—it simply stops new sessions from being sticky. Option C is wrong because configuring health checks to mark instances unhealthy would cause the ALB to stop sending new traffic to the instance, but it does not wait for existing connections to complete; the deregistration delay is what handles in-flight requests, not health check status.

421
MCQeasy

A DevOps engineer needs to grant cross-account access to an S3 bucket in Account A for a user in Account B. Which combination of policies is required?

A.An S3 bucket policy in Account A and an IAM policy in Account B.
B.An IAM role in Account A and an IAM policy in Account B.
C.Only an IAM policy in Account B.
D.Only an S3 bucket policy in Account A.
AnswerA

For cross-account S3 access, both an S3 bucket policy in the owning account (Account A) that explicitly allows the external principal (e.g., an IAM user or role in Account B) to perform the desired actions, and an IAM policy in Account B that grants that same principal permission to call S3 on the bucket ARN are required. AWS evaluates identity-based policies and resource-based policies together, so a request succeeds only if both sides explicitly allow it. Without either, the request is implicitly denied.

Why this answer

Cross-account access to an S3 bucket requires both a resource-based policy (the S3 bucket policy in Account A) that grants the necessary permissions to the principal from Account B, and an identity-based policy (an IAM policy in Account B) attached to the user that allows the S3 actions. The bucket policy explicitly authorizes the external user (by ARN), while the IAM policy ensures the user has the required permissions to initiate the request. Without both, the request will fail due to the lack of either resource-side authorization or identity-side authorization.

Exam trap

The trap here is that candidates often assume a bucket policy alone is sufficient for cross-account access, forgetting that the external user must also have an IAM policy that explicitly allows the S3 actions, leading them to incorrectly select Option D.

How to eliminate wrong answers

Option B is wrong because an IAM role in Account A would require the user in Account B to assume the role, which is a different mechanism (role-based cross-account access) and does not directly grant access via a bucket policy; the question specifically asks for granting access to an S3 bucket, not using role assumption. Option C is wrong because an IAM policy in Account B alone cannot grant access to a resource in Account A; the resource owner (Account A) must also authorize the access via a bucket policy or ACL. Option D is wrong because an S3 bucket policy in Account A alone is insufficient; the user in Account B must also have an IAM policy that allows the S3 actions, as the bucket policy only authorizes the external principal but does not grant the user permission to make the request.

422
MCQeasy

A company runs a web application on Amazon EC2 instances behind an Application Load Balancer. The DevOps team wants to receive an alert when the number of HTTP 5xx errors from the load balancer exceeds 100 in a 5-minute period. The team wants to use the most direct and least complex method. What should they do?

A.Install the CloudWatch agent on the EC2 instances to collect web server logs, and create a custom metric for 5xx errors with an alarm.
B.Configure AWS CloudTrail to log load balancer API calls and create an EventBridge rule to trigger an SNS notification when 5xx errors occur.
C.Create a CloudWatch alarm on the HTTPCode_ELB_5XX_Count metric for the load balancer with a threshold of 100 over 5 minutes, and configure an Amazon SNS topic for notifications.
D.Enable access logs on the load balancer, then use CloudWatch Logs metric filters to count 5xx errors and create an alarm on the resulting metric.
AnswerC

The Application Load Balancer publishes the HTTPCode_ELB_5XX_Count metric to CloudWatch. Creating an alarm directly on this metric with a 5-minute period and a threshold of 100, then attaching an SNS topic for email or SMS notifications, is the simplest and most direct way to meet the requirement. No additional agents or complex configurations are needed.

Why this answer

Application Load Balancers automatically publish the HTTPCode_ELB_5XX_Count metric to CloudWatch, so creating an alarm on that metric with an SNS notification is the most direct and least complex solution. The other options involve unnecessary complexity or use services that do not monitor HTTP errors.

Exam trap

The trap here is overcomplicating the solution by using access logs or CloudTrail, when a built-in CloudWatch metric already provides the needed data.

423
Multi-Selectmedium

Which TWO AWS services can be used to implement a blue/green deployment for an application running on Amazon EC2 instances?

Select 2 answers
A.AWS CloudFormation
B.AWS CodeDeploy
C.AWS Elastic Beanstalk
D.AWS OpsWorks
E.AWS CodeBuild
AnswersB, C

AWS CodeDeploy is the correct choice because it provides a native blue/green deployment model for EC2/On-Premises and AWS Lambda workloads. During a blue/green deployment, CodeDeploy automatically provisions a new replacement fleet (green), registers it with the load balancer, and shifts traffic incrementally from the original fleet (blue) to the green fleet, with options for automatic rollback if deployment fails. Its built-in deployment lifecycle hooks and traffic-routing controls make it a fully managed service purpose-built for this strategy.

Why this answer

AWS CodeDeploy (Option B) is correct because it natively supports blue/green deployments for Amazon EC2 instances by allowing you to provision a new set of instances (the green environment), deploy the new application revision to them, and then shift traffic from the old (blue) environment to the new one. This is achieved through integration with an Elastic Load Balancer (ELB) or Auto Scaling groups, where CodeDeploy manages the lifecycle of instances and traffic routing during the deployment process.

Exam trap

The trap here is that candidates often confuse AWS CloudFormation (which can define the infrastructure for blue/green deployments) with the actual deployment service that orchestrates the traffic shift, leading them to select CloudFormation instead of CodeDeploy or Elastic Beanstalk.

424
MCQeasy

Refer to the exhibit. A DevOps engineer created a CloudFormation stack that includes a Lambda function. The stack creation failed and rolled back. The error message for the Lambda function says 'Resource creation cancelled'. What is the most likely cause?

A.The Lambda function's IAM role does not have sufficient permissions.
B.The Lambda function's code is missing from the S3 bucket.
C.The stack rollback was triggered due to a failure in another resource, and the Lambda creation was cancelled.
D.The Lambda function creation timed out.
AnswerC

CloudFormation creates stack resources in parallel where no dependencies exist, and a failure in any resource triggers an automatic rollback that cancels all other in-progress resource creations. When the Lambda function's AWS::Lambda::Function resource is cancelled in this way, CloudFormation marks it as CREATE_FAILED with the reason 'Resource creation cancelled' — even though the Lambda service never reported an error. The root cause lies in a different resource that failed first, so you must inspect the full stack event list to identify the original failure.

Why this answer

CloudFormation creates resources in a specific order, and if a dependency fails or another resource fails, CloudFormation cancels the creation of subsequent resources and initiates a rollback. The error 'Resource creation cancelled' indicates that the Lambda function was never actually attempted to be created; instead, its creation was aborted due to a failure in a preceding or parallel resource. This is a standard CloudFormation behavior when a stack operation is interrupted by a failure in another resource.

Exam trap

Candidates often confuse 'Resource creation cancelled' with a resource-specific failure (e.g., insufficient permissions), but this error indicates CloudFormation cancelled the resource due to a failure in another resource during stack creation.

How to eliminate wrong answers

Option A is wrong because insufficient IAM permissions would cause a different error, such as 'API: iam:PassRole' or 'AccessDenied', not 'Resource creation cancelled'. Option B is wrong because missing code in the S3 bucket would result in an error like 'Unable to fetch code from S3' or '404 Not Found', not a cancellation message. Option D is wrong because a timeout would produce a 'Resource timed out' error, not 'Resource creation cancelled', and CloudFormation would wait for the timeout period before failing.

425
MCQhard

A financial services company uses Chef for configuration management. They need to enforce security compliance across thousands of EC2 instances. The compliance requirements include specific file permissions, firewall rules, and user account settings. They want to automatically remediate non-compliant instances. Which approach is MOST effective?

A.Use AWS Config rules to detect non-compliance and send notifications.
B.Use AWS CloudWatch Events to trigger a Lambda function that runs remediation scripts.
C.Use AWS Systems Manager Patch Manager to apply patches.
D.Use Chef recipes to define desired state and enforce compliance on each client run.
AnswerD

Chef recipes define a declarative desired state for system resources, and the chef-client runs on a schedule (typically every 30 minutes) to converge each node to that state. During each run, the client queries current system state, compares it to the recipe definitions, and automatically remediates any drift—this is true continuous enforcement, not just reporting. Because chef-client is idempotent, repeated runs are safe, which is essential in a regulated financial environment where config drift must be corrected promptly and consistently.

Why this answer

Chef is already in use for configuration management, and its core strength is enforcing desired state through recipes. By running Chef client on each EC2 instance (via cron or Systems Manager), non-compliant instances are automatically remediated on each convergence cycle without needing separate detection or notification tools. This approach directly addresses the requirement for automated remediation using the existing toolchain.

Exam trap

The trap here is that candidates often overcomplicate the solution by adding AWS services (Config, Lambda) when the existing Chef tool already provides built-in, continuous compliance enforcement through its converge cycle.

How to eliminate wrong answers

Option A is wrong because AWS Config rules only detect and evaluate compliance but do not automatically remediate; they require additional services like Systems Manager Automation or Lambda for remediation. Option B is wrong because CloudWatch Events (now Amazon EventBridge) can trigger Lambda, but this is an event-driven, reactive approach that adds complexity and latency compared to Chef's built-in continuous enforcement. Option C is wrong because Patch Manager is specifically for OS patch management, not for enforcing file permissions, firewall rules, or user account settings.

426
MCQeasy

A startup is using AWS CodePipeline to deploy a Python web application to AWS Elastic Beanstalk. The pipeline has a source stage (CodeCommit), a build stage (CodeBuild), and a deploy stage (Elastic Beanstalk). The build stage runs unit tests and creates a deployable zip file. The deploy stage uses the Elastic Beanstalk deploy provider. Recently, the deploy stage started failing with the error: 'The API call 'elasticbeanstalk:CreateApplicationVersion' failed with status 403.' The CodePipeline service role has the following permissions: 'elasticbeanstalk:DescribeApplications', 'elasticbeanstalk:DescribeEnvironments', 'elasticbeanstalk:UpdateEnvironment'. What should the DevOps engineer do to resolve the issue?

A.Change the deploy provider to use CodeDeploy instead of Elastic Beanstalk
B.Add 's3:PutObject' permission to the CodePipeline service role to allow it to upload the zip file to S3
C.Add the 'elasticbeanstalk:CreateApplicationVersion' and 'elasticbeanstalk:DeleteApplicationVersion' permissions to the CodePipeline service role
D.Update the Elastic Beanstalk environment's service role to allow CodePipeline to deploy
AnswerC

When CodePipeline deploys to Elastic Beanstalk, its service role must be allowed to call the Elastic Beanstalk APIs `CreateApplicationVersion` and `DeleteApplicationVersion`. The `CreateApplicationVersion` action registers a new application version from the source artifact, while `DeleteApplicationVersion` lets the pipeline remove old versions—both are essential for a successful deployment. Without these permissions, the deploy action receives an `AccessDenied` (403) error. Adding them to the CodePipeline service role directly resolves the failure because the role is the identity CodePipeline assumes when making API calls on your behalf.

Why this answer

The 403 error on elasticbeanstalk:CreateApplicationVersion indicates the CodePipeline service role lacks that specific IAM action. Elastic Beanstalk's deploy provider calls CreateApplicationVersion to register the new application revision, and it also calls DeleteApplicationVersion during cleanup of old revisions. Adding both permissions to the CodePipeline service role resolves the authorization failure without changing the pipeline architecture.

Exam trap

DOP-C02 often tests whether candidates can distinguish between the CodePipeline service role (the caller) and the Elastic Beanstalk environment service role (the resource) — candidates frequently pick the environment role fix when the 403 is actually about the caller's permissions.

How to eliminate wrong answers

Option A is wrong because switching to CodeDeploy does not address the missing IAM permission and would require re-architecting the deployment target, which is unnecessary. Option B is wrong because the error is an Elastic Beanstalk API authorization failure, not an S3 upload failure — CodeBuild already handles artifact upload, and adding s3:PutObject would not grant the missing elasticbeanstalk action. Option D is wrong because the Elastic Beanstalk environment's service role governs what the environment can do (e.g., access EC2, S3), not what CodePipeline can call on the Elastic Beanstalk API; the caller's identity is the CodePipeline service role.

427
MCQhard

A company uses AWS Systems Manager to manage hybrid servers. They want to automate the patching of Windows servers using Patch Manager. However, some servers are not showing up in the compliance reporting. What should the DevOps engineer check first?

A.Ensure the SSM Agent is installed and running on the servers
B.Verify that the servers have the correct patch baseline tags
C.Check that the Patch Baseline is configured to include the missing servers
D.Confirm that the servers have an IAM service role for Systems Manager
AnswerA

For a hybrid server to be managed by Systems Manager, the SSM Agent must be installed and actively running on the operating system. The agent is the on-premises component that establishes the communication channel with the Systems Manager service, handles requests for Run Command, Patch Manager, and Inventory, and reports the instance's status back to the service. If the agent is absent, stopped, or in an unhealthy state, the server will not appear in the inventory or compliance views, and no Systems Manager operation can target it. Reinstalling or restarting the agent, and periodically verifying its health, is the first-line remediation for 'missing' hybrid nodes.

Why this answer

The SSM Agent is the core component that enables a server to communicate with AWS Systems Manager. Without the agent installed and running, the server cannot register with the service, receive patch commands, or report its compliance status. Therefore, this is the most fundamental prerequisite to check first when servers are missing from compliance reporting.

Exam trap

The trap here is that candidates often jump to IAM roles or tag-based configurations first, forgetting that the SSM Agent is the absolute prerequisite for any Systems Manager functionality, including Patch Manager compliance reporting.

How to eliminate wrong answers

Option B is wrong because patch baseline tags are used to associate a server with a specific patch baseline, but they do not affect whether the server appears in compliance reporting at all; a server must first be managed by Systems Manager via the SSM Agent. Option C is wrong because the Patch Baseline configuration defines which patches to apply, not which servers are included in reporting; server visibility is determined by agent connectivity and instance registration. Option D is wrong because while an IAM instance profile (not a service role) is required for the SSM Agent to call AWS APIs, the agent must still be installed and running first; without the agent, no IAM role can make the server appear in compliance reporting.

428
MCQeasy

A company is using Amazon S3 to store sensitive data. The security team mandates that all data must be encrypted at rest using server-side encryption with AWS Key Management Service (SSE-KMS). The DevOps engineer must ensure that any new objects uploaded to the bucket are automatically encrypted. What should the engineer do?

A.Enable CORS on the bucket to allow encrypted uploads.
B.Apply a bucket policy that denies PutObject unless the request includes the x-amz-server-side-encryption header with aws:kms.
C.Enable default encryption on the S3 bucket and select AWS-KMS as the encryption method.
D.Enable S3 Versioning to protect encrypted objects.
AnswerC

Enabling default encryption on an S3 bucket with AWS-KMS selected ensures that every new object is automatically encrypted at rest using SSE-KMS, regardless of whether the upload request includes any encryption headers. This uses envelope encryption with a customer-managed KMS key, offering centralized key management, separate permissions for key access, and an audit trail of key usage. It is the correct approach because it is a direct, bucket-level setting that protects sensitive data without requiring client changes or policy-based request rejections.

Why this answer

Enabling default encryption on the S3 bucket with SSE-KMS ensures that all objects uploaded to the bucket are automatically encrypted at rest with AWS KMS. Option A is incorrect because CORS is for cross-origin requests and does not affect encryption. Option B is incorrect because while a bucket policy can enforce encryption headers, it does not provide default encryption; it only denies requests without the header, and default encryption is a simpler and more reliable method.

Option D is incorrect because versioning does not encrypt data.

429
MCQhard

A DevOps engineer is designing a CI/CD pipeline for a microservices application running on Amazon ECS with Fargate. The team wants to use a blue/green deployment strategy to minimize downtime. Which combination of AWS services and configurations should be used to implement this?

A.Use Amazon ECS service with a rolling update deployment controller
B.Create two separate ECS services and use Route 53 weighted routing to shift traffic
C.Use AWS CloudFormation with a custom resource to swap target group weights
D.Use CodeDeploy with an ECS compute platform and an Application Load Balancer
AnswerD

CodeDeploy with an ECS compute platform natively manages blue/green deployments by creating a new task set, installing it into a preconfigured green target group, and then incrementally shifting the ALB's production listener weight from the original to the new target group based on a Canary or Linear deployment configuration. The service integrates with a specified AppSpec file to run pre- and post-traffic validation hooks, and you can attach CloudWatch alarms that trigger automatic rollback if the new version misbehaves. This is the purpose-built mechanism that owns the entire traffic-shifting lifecycle, from creating the replacement task set to deregistering the old one, without resorting to custom code.

Why this answer

CodeDeploy with an ECS compute platform natively supports blue/green deployments for ECS services by orchestrating traffic shifting between two target groups behind an Application Load Balancer. This approach minimizes downtime by gradually routing traffic from the 'blue' (current) task set to the 'green' (new) task set, with built-in rollback capabilities and lifecycle hooks for validation.

Exam trap

The trap here is that candidates often confuse blue/green with rolling updates (Option A) or assume that manual traffic routing via Route 53 (Option B) or CloudFormation custom resources (Option C) can achieve the same orchestrated, automated deployment with health checks and rollback that CodeDeploy provides natively.

How to eliminate wrong answers

Option A is wrong because a rolling update deployment controller in ECS replaces tasks incrementally without creating a separate environment for validation, which does not provide the zero-downtime traffic shifting characteristic of blue/green deployments. Option B is wrong because managing two separate ECS services with Route 53 weighted routing introduces DNS caching delays and lacks orchestrated traffic shifting, health checks, and rollback automation that CodeDeploy provides. Option C is wrong because AWS CloudFormation custom resources are not designed for real-time traffic shifting or deployment orchestration; they are intended for provisioning custom infrastructure logic, and swapping target group weights manually would not integrate with ECS deployment lifecycle hooks or automatic rollback.

430
MCQhard

Refer to the exhibit. A Lambda function uses the IAM role with the above policy. The function is configured to access a DynamoDB table MyTable and an RDS instance in a VPC. When invoked, the function fails with an error indicating it cannot describe VPC subnets. What is the MOST likely cause?

A.The Lambda function is missing permissions to describe VPC subnets and security groups.
B.The Lambda function does not have permission to write to DynamoDB.
C.The Lambda function cannot create network interfaces in the VPC.
D.The DynamoDB table's resource policy denies access from Lambda.
AnswerA

The Lambda execution role must explicitly allow ec2:DescribeSubnets and ec2:DescribeSecurityGroups so that the Lambda service can inspect your VPC and locate the subnets and security groups you specified. These read-only actions are prerequisites for Lambda to create an elastic network interface (ENI); without them, the invocation fails during the VPC provisioning step, even if the role allows creating ENIs. This is why the error specifically mentions missing subnet or security group describe permissions.

Why this answer

The Lambda execution role policy shown in the exhibit grants DynamoDB and RDS-related actions but omits the ec2:DescribeSubnets, ec2:DescribeSecurityGroups, and ec2:DescribeNetworkInterfaces permissions required when a function is attached to a VPC. When Lambda is configured with VPC access, the service must call these EC2 APIs to validate and place the ENIs, so the invocation fails with a subnet-description error. Adding the missing ec2:Describe* permissions to the role resolves the failure.

Exam trap

DOP-C02 often tests the misconception that VPC-attached Lambda only needs ENI creation permissions, when in fact DescribeSubnets, DescribeSecurityGroups, and DescribeNetworkInterfaces are also mandatory and their absence produces the exact error described.

How to eliminate wrong answers

Option B is wrong because DynamoDB write failures would surface as AccessDeniedException on PutItem/UpdateItem, not as a VPC subnet description error, and the exhibit policy already grants DynamoDB actions. Option C is wrong because ENI creation failures produce a different error (e.g., 'The provided execution role does not have permissions to call CreateNetworkInterface'), and the reported error is specifically about describing subnets. Option D is wrong because a DynamoDB resource policy denial would return an access-denied error on the DynamoDB call itself, unrelated to VPC subnet enumeration.

431
MCQeasy

A company runs a web application on Amazon EC2 instances behind an Application Load Balancer (ALB). The application uses a custom health check endpoint '/health'. The DevOps team notices that the ALB is marking some instances as unhealthy even though the application is running fine. The team checks the security groups and network ACLs and confirms they allow traffic. What should the team check next?

A.Ensure the health check path is case-insensitive.
B.Increase the health check interval and timeout values.
C.Confirm that the health check path is correctly configured to '/health' on the target group.
D.Verify that the health check port matches the application port.
AnswerC

The target group health check path must exactly match the application endpoint that is designed to return an HTTP 200 OK when healthy. If the path is incorrectly set to something like '/' instead of '/health', the application may return a different status code, such as a redirect or a default page, which can cause intermittent health check failures as the application's behavior changes under load. Confirming the path resolves the root cause because the health check will then probe a purpose-built endpoint that consistently returns 200.

Why this answer

If security groups and NACLs are confirmed correct, the next most likely cause is a misconfigured health check path on the target group — for example, the path is set to '/' or '/healthz' instead of '/health', so the ALB receives a 404 and marks the instance unhealthy. Verifying the exact path string in the target group's health check settings is the correct next diagnostic step. The other options either do not apply or are premature tuning steps.

Exam trap

DOP-C02 often tests whether candidates jump to tuning health check timing parameters when the actual root cause is a simple configuration mismatch like the wrong path or success code, so candidates pick interval/timeout adjustments instead of verifying the path.

How to eliminate wrong answers

Option A is wrong because HTTP paths are case-sensitive by convention and ALB health checks do not treat paths as case-insensitive; this is not a real configuration knob. Option B is wrong because increasing interval and timeout values only delays failure detection — it does not fix a wrong path or a failing endpoint, and the team has no evidence of slow responses. Option D is wrong because the question states the application is running fine and security groups allow traffic; while port mismatch is a possible cause, the scenario points to path configuration, and the correct answer addresses the path specifically.

432
MCQeasy

A company runs a static website on Amazon S3 with public read access. The website content is stored in an S3 bucket and served through an Amazon CloudFront distribution for better performance and security. Recently, the company noticed that some users are accessing the S3 bucket directly via the S3 endpoint, bypassing CloudFront. This increases costs and exposes the bucket to potential attacks. The company wants to ensure that all access to the website goes through CloudFront only. Which solution should the company implement?

A.Set the S3 bucket policy to deny all requests that do not come from the CloudFront distribution's IP addresses.
B.Configure the S3 bucket to use AWS WAF to block requests that do not have a custom header set by CloudFront.
C.Create an origin access identity (OAI) in CloudFront and update the S3 bucket policy to allow only the OAI to read objects.
D.Change the S3 bucket to be private and use presigned URLs for all requests.
AnswerC

Creating an origin access identity (OAI) in CloudFront and updating the bucket policy to permit only that OAI to read objects is the correct approach. The OAI is a special CloudFront user that validates requests to the S3 origin with AWS Signature Version 4, and the bucket policy grants s3:GetObject permission exclusively to this principal. This blocks any request that does not come through the CloudFront distribution, while still allowing the static content to be publicly served to end users via CloudFront.

Why this answer

To restrict access to the S3 bucket only through CloudFront, use an origin access identity (OAI) and a bucket policy that allows only the OAI. This way, direct access via S3 URL is denied.

433
Multi-Selecteasy

A company wants to ensure that all changes to its Amazon S3 bucket policies are logged for auditing purposes. Which TWO AWS services should be enabled to capture these changes?

Select 2 answers
A.Amazon CloudWatch
B.AWS Config
C.Amazon GuardDuty
D.VPC Flow Logs
E.AWS CloudTrail
AnswersB, E

AWS Config is the correct service for this requirement because it continuously records and evaluates the configuration of AWS resources, including S3 bucket policies, and maintains a detailed configuration timeline. It can compare the recorded configuration against desired compliance rules (e.g., ensuring a bucket is not public). With AWS Config, you can see exactly when a bucket policy was last changed and what the previous configuration was, making it ideal for auditing and compliance.

Why this answer

Options B and E are correct because AWS Config records resource configuration changes, including S3 bucket policies, and AWS CloudTrail logs API calls such as PutBucketPolicy. Option A is incorrect because Amazon CloudWatch monitors operational metrics and logs, not auditing of policy changes. Option C is incorrect because Amazon GuardDuty provides threat detection, not audit logging.

Option D is incorrect because VPC Flow Logs capture network traffic, not configuration changes.

434
MCQhard

A company is running a critical application on an Amazon EC2 instance that needs to access an S3 bucket. The application must use temporary credentials that automatically rotate. The DevOps engineer must ensure that the credentials are never stored on disk. Which approach meets these requirements?

A.Store the credentials in AWS Secrets Manager and retrieve them at application startup.
B.Attach an IAM role to the EC2 instance and use the instance profile to obtain temporary credentials from the instance metadata service.
C.Use AWS Systems Manager Parameter Store to store the credentials and retrieve them using the EC2 instance's IAM role.
D.Generate an access key and secret key for an IAM user and store them in a configuration file on the EC2 instance.
AnswerB

The best practice is to attach an IAM role to the EC2 instance; the instance profile exposes temporary security credentials via the Instance Metadata Service (IMDSv2), which the AWS SDKs automatically load and refresh. These credentials are short-lived, rotated automatically, and never written to disk, so no secret material is present in the file system, environment variables, or configuration files. This eliminates the need to manage access keys manually and reduces the risk of exposure if the instance is compromised.

Why this answer

Attaching an IAM role to the EC2 instance and using the instance profile allows the application to obtain temporary credentials from the EC2 instance metadata service (IMDS). These credentials are automatically rotated by AWS before they expire, and they are never stored on disk—they are fetched on-demand from the metadata endpoint (http://169.254.169.254/latest/meta-data/iam/security-credentials/). This satisfies both the requirement for automatic rotation and the prohibition against disk storage.

Exam trap

The trap here is that candidates may confuse AWS Secrets Manager or Parameter Store with a solution for automatic credential rotation, not realizing that those services store static secrets unless explicitly configured with rotation via Lambda, whereas an IAM instance profile inherently provides automatically rotating temporary credentials without any disk storage.

How to eliminate wrong answers

Option A is wrong because while AWS Secrets Manager can store and rotate credentials, the application would still need to retrieve and hold them in memory, and the credentials stored there are long-term IAM user keys or secrets, not automatically rotating temporary credentials from an instance profile. Option C is wrong because AWS Systems Manager Parameter Store can store credentials, but it does not inherently rotate them; the stored credentials would be static unless manually updated, and the application would still need to handle them in memory, not leveraging the automatic rotation of instance metadata service credentials. Option D is wrong because storing access keys and secret keys in a configuration file on disk directly violates the requirement that credentials never be stored on disk, and these static credentials do not automatically rotate.

435
MCQhard

A company uses AWS CodeCommit as a source repository and wants to enforce that all commits are signed using GPG keys. The DevOps team configures a pre-receive hook in CodeCommit to validate commit signatures. However, the hook rejects all commits even when valid GPG signatures are present. What is the most likely cause?

A.The GPG key is not registered with the IAM user's profile.
B.CodeCommit does not support pre-receive hooks.
C.The hook script has a syntax error.
D.The repository is not configured to require signed commits.
AnswerB

This is correct. CodeCommit is a managed Git service that does not expose a file system or a server-side Git hooks directory, so Git's pre-receive hook scripts are not supported. Instead, CodeCommit offers repository triggers (SNS/Lambda) and notification rules, which are evaluated after the push is accepted. Therefore, any apparent pre-receive hook failure cannot actually occur in CodeCommit.

Why this answer

AWS CodeCommit does not support pre-receive hooks. Pre-receive hooks are a feature of self-managed Git repositories (e.g., GitHub Enterprise, GitLab, or on-premises Git servers) that run on the server before accepting a push. CodeCommit uses IAM policies and repository-level settings (such as requiring signed commits via the 'git push --signed' flag) to enforce commit signing, not server-side hooks.

Therefore, any attempt to configure a pre-receive hook in CodeCommit will fail, causing all commits to be rejected.

Exam trap

The trap here is that candidates confuse CodeCommit with self-managed Git platforms (like GitHub or GitLab) that support pre-receive hooks, leading them to assume CodeCommit also supports this feature, when in fact CodeCommit uses a different enforcement mechanism (repository-level settings and IAM policies).

Why the other options are wrong

A

While GPG key must be associated with the IAM user, the issue is that CodeCommit doesn't support pre-receive hooks.

C

Even if the script is correct, CodeCommit does not execute pre-receive hooks.

D

CodeCommit does not have a built-in setting to require signed commits; the hook is the intended mechanism, but it's not supported.

436
Multi-Selecteasy

Which TWO AWS services can be used to automate the configuration of EC2 instances at launch? (Choose two.)

Select 2 answers
A.Amazon CloudWatch
B.EC2 user data
C.AWS CloudFormation
D.Amazon Inspector
E.AWS Config
AnswersB, C

EC2 user data is a feature that lets you pass a script or cloud-init directives to an instance at launch; the script runs automatically the first time the instance boots, enabling package installation, file creation, and service configuration. This provides imperative, instance-level configuration automation for initial setup, though it does not manage ongoing changes or coordinate multiple resources. It is a direct and valid way to automate EC2 configuration.

Why this answer

EC2 user data allows you to specify scripts or cloud-init directives that run automatically during the first boot cycle of an EC2 instance. This enables automated software installation, configuration, and post-launch tasks without manual intervention, making it a core tool for instance configuration at launch.

Exam trap

The trap here is that candidates often confuse AWS Config (which audits existing configurations) or CloudWatch (which monitors) with services that can actively configure instances at launch, when in fact only user data and infrastructure-as-code tools like CloudFormation can perform that initial automation.

437
Multi-Selecteasy

A DevOps engineer is writing an AWS CloudFormation template to create a VPC with public and private subnets. The engineer wants to ensure that the private subnets can access the internet through a NAT gateway. Which resources must be included in the template? (Choose TWO.)

Select 2 answers
A.AWS::EC2::RouteTable
B.AWS::EC2::VPNGateway
C.AWS::EC2::InternetGateway
D.AWS::EC2::NatGateway
E.AWS::EC2::VPCEndpoint
AnswersA, D

An AWS::EC2::RouteTable is correct because a private subnet cannot reach the internet on its own; you must create a route table and explicitly associate it with the private subnet. Within that route table, a route with destination 0.0.0.0/0 pointing to the NAT gateway is required, and the association between the route table and the subnet is what makes the NAT gateway the effective next hop for all outbound internet traffic.

Why this answer

A is correct because an AWS::EC2::RouteTable resource is required to define the routing rules for the private subnets. Specifically, you must create a route table for the private subnets and add a default route (0.0.0.0/0) that points to the NAT gateway, enabling outbound internet traffic from instances in the private subnets while blocking inbound traffic from the internet.

Exam trap

The trap here is that candidates often think an Internet Gateway is required for private subnet internet access, but the NAT gateway itself uses the internet gateway; the template only needs the NAT gateway and a route table for the private subnets, not the internet gateway resource directly.

438
MCQeasy

A company uses AWS Systems Manager to manage its EC2 instances at scale. The DevOps team wants to ensure that all instances are patched with the latest security updates. Which Systems Manager capability should they use to automate patching?

A.Run Command
B.Patch Manager
C.Automation
D.State Manager
AnswerB

Patch Manager is the dedicated Systems Manager capability for automating the entire patching workflow, including scan and install operations using patch baselines, and it integrates with maintenance windows for controlled deployments. It also generates patch compliance data that can be queried across your fleet, which is exactly what is required when you need to keep EC2 instances patched automatically.

Why this answer

Patch Manager is the correct choice because it is the dedicated AWS Systems Manager capability designed to automate the process of patching managed instances with security updates and other types of updates. It uses patch baselines to define which patches should be installed and can schedule patching on a recurring basis, ensuring compliance across the EC2 fleet.

Exam trap

The trap here is that candidates often confuse Patch Manager with State Manager because both can enforce desired states, but State Manager lacks the patching-specific logic and patch baseline integration that Patch Manager provides.

How to eliminate wrong answers

Option A is wrong because Run Command is used to execute ad-hoc scripts or commands on instances, not to manage or schedule recurring patching operations. Option C is wrong because Automation is a general-purpose capability for automating common maintenance and deployment tasks, but it does not provide the built-in patch baseline management and compliance reporting that Patch Manager offers. Option D is wrong because State Manager is used to define and maintain consistent configuration state of instances (e.g., ensuring a specific software is installed or a service is running), but it lacks the native patching workflows and patch-specific features like approval rules and auto-approval delays.

439
MCQmedium

A company uses AWS Organizations with multiple accounts. The security team wants to ensure that all accounts automatically forward their CloudWatch Logs to a central logging account. Which solution should the team implement?

A.Enable AWS Config aggregator in the central account
B.Use AWS Service Catalog to create a product for log forwarding
C.Use AWS CloudFormation StackSets to deploy a subscription filter and Lambda function in each account
D.Configure AWS Organizations to automatically forward logs
AnswerC

AWS CloudFormation StackSets extends CloudFormation to deploy stacks across multiple accounts and regions within AWS Organizations, automatically and in a single operation. You can define a template containing a CloudWatch Logs subscription filter and a Lambda function, and StackSets will create those resources in every member account. The subscription filter streams selected log events from each account's log groups to the Lambda function, which then processes and forwards them to a centralized destination such as an S3 bucket, Kinesis, or a central CloudWatch account. This approach is fully automated, consistent, and the recommended pattern for centralized log forwarding.

Why this answer

AWS CloudFormation StackSets allows you to deploy infrastructure components across multiple accounts and regions in an AWS Organization. In this case, you can create a StackSet that includes a CloudWatch Logs subscription filter and a Lambda function to forward logs from each account to a central logging account. The subscription filter triggers the Lambda function to forward logs to a destination in the central account.

This ensures all accounts automatically forward their CloudWatch Logs. Option A is incorrect because AWS Config aggregator aggregates configuration and compliance data, not logs. Option B is incorrect because AWS Service Catalog is used to create and manage IT service catalogs for approved products, not for log forwarding.

Option D is incorrect because AWS Organizations does not have a native capability to forward logs; it manages policies and account structure.

440
Multi-Selectmedium

Which TWO of the following are benefits of using AWS Certificate Manager (ACM) to manage SSL/TLS certificates? (Choose two.)

Select 2 answers
A.Ability to use the same certificate on multiple EC2 instances.
B.Support for wildcard certificates only.
C.Automatic renewal of certificates.
D.Integration with Elastic Load Balancing and Amazon CloudFront.
E.Free certificates for use on any AWS service.
AnswersC, D

ACM automatically renews issued certificates when they are deployed on supported services, provided the domain validation remains valid, which spares engineers from manually tracking expiration dates. This automated lifecycle management is a key operational benefit because it prevents unexpected service interruptions caused by expired certificates. The renewal process works silently in the background, reissuing certificates before the current one expires, making it a primary reason organizations choose ACM.

Why this answer

Option C is correct because ACM automatically renews certificates that it issued and manages, as long as the certificate is in use and the domain validation remains valid, eliminating manual renewal overhead. Option D is correct because ACM certificates can be directly associated with Elastic Load Balancing (ALB/NLB) and Amazon CloudFront distributions, enabling seamless deployment of TLS termination without exporting private keys. Option A is not a listed benefit because ACM-issued public certificates cannot be exported or installed directly on EC2 instances; you must use services like ELB, CloudFront, or API Gateway.

Option B is incorrect because ACM supports both wildcard and non-wildcard (single-domain and multi-domain/SAN) certificates. Option E is incorrect because ACM public certificates are free only for integrated AWS services, not for arbitrary AWS services or exportable use, and ACM Private CA certificates incur cost.

441
MCQmedium

A company is using AWS CloudTrail to log API calls. The security team needs to ensure that log files are tamper-proof and can be used to verify integrity. Which feature should be enabled?

A.Server-side encryption (SSE-S3)
B.CloudTrail log file integrity validation
C.S3 Object Lock
D.MFA delete on the S3 bucket
AnswerB

CloudTrail log file integrity validation is purpose-built for exactly this need: it delivers signed digest files to your S3 bucket that cover the log files and link together in a hash chain. Each digest contains the SHA-256 hash of the prior digest and the current log files, and is signed with a private key by CloudTrail. You can use the corresponding public key to verify both the authenticity and the integrity of the logs, which lets you detect any modification, deletion, or forgery. Enabling this feature ensures that your audit trail itself is trustworthy, which is a key defense against attackers trying to cover their tracks by altering logs.

Why this answer

CloudTrail log file integrity validation uses SHA-256 hashing and digital signing to ensure logs have not been tampered with. S3 object lock prevents deletion but not modification. MFA delete protects deletion but not modification.

SSE encrypts data at rest but does not protect integrity.

442
Multi-Selectmedium

A DevOps engineer is designing a monitoring solution for a multi-tier web application hosted on AWS. The application consists of an Application Load Balancer (ALB), EC2 instances, and an RDS database. The engineer needs to capture and analyze HTTP request logs from the ALB to understand client behavior and troubleshoot errors. Which THREE steps are necessary to achieve this?

Select 3 answers
A.Install the CloudWatch Agent on the ALB
B.Enable AWS CloudTrail for the ALB
C.Use Amazon Athena to query the access logs in S3
D.Enable access logs on the ALB
E.Create an Amazon S3 bucket to store the access logs
AnswersC, D, E

Amazon Athena is the correct service for interactively querying ALB access logs directly from S3 without loading data into a database. The logs are stored as gzipped text files in a partitionable layout (AWSLogs/account-id/elasticloadbalancing/region/yyyy/mm/dd), which can be registered as a table in Athena using JSON or CSV SerDe. With Athena's SQL, you can analyze request patterns, error rates, latency, and client behavior by writing queries against the log fields, and partitioning or partition projection keeps query costs low. This makes Athena the natural final step after enabling ALB access logs and storing them in S3.

Why this answer

Option D is correct because ALB access logs must be explicitly enabled on the load balancer, which captures detailed information about every HTTP/HTTPS request including client IP, request path, response codes, and latency. Option E is correct because ALB access logs are delivered to an Amazon S3 bucket, so a target S3 bucket (with the proper bucket policy allowing the ALB to write) must exist before enabling logging. Option C is correct because Amazon Athena can query the ALB access logs stored in S3 directly using SQL, enabling analysis of client behavior and troubleshooting of errors without loading data into a database.

Option A is incorrect because the CloudWatch Agent runs on EC2 instances or on-premises servers, not on ALBs, which are managed services that cannot host agents. Option B is incorrect because AWS CloudTrail records API activity and management events, not HTTP request logs, so it does not capture ALB access log data.

Exam trap

DOP-C02 often tests the confusion between CloudTrail (API audit logs) and ALB access logs (HTTP request logs), and the misconception that the CloudWatch Agent can be installed on managed services like ALB.

443
MCQeasy

A DevOps engineer is investigating an incident where an EC2 instance became unreachable. The engineer checks the AWS Management Console and finds the instance is running, but the status check shows '2/2 checks passed' and the system log shows no errors. What should the engineer do NEXT to diagnose the connectivity issue?

A.Review the CloudWatch metrics for CPU utilization and network throughput.
B.Reboot the instance to reset the network interface.
C.Stop and start the instance to move it to new underlying hardware.
D.Check the security group and network ACL rules to ensure inbound traffic is allowed.
AnswerD

Security groups act as a stateful firewall at the instance level, while network ACLs are stateless at the subnet level, so both must allow the relevant inbound traffic and the NACL must also allow the corresponding outbound return traffic. An incorrect deny rule, an overly restrictive CIDR, or a missing allow for the source IP/port pairs will cause exactly this kind of unreachability even when the instance is running and healthy, making this the correct first troubleshooting step.

Why this answer

Since the instance is running, status checks pass, and the system log shows no errors, the issue is not with the operating system or underlying hardware. The most likely cause is a network-layer restriction, such as security group or network ACL rules blocking inbound traffic. Checking these rules is the correct next step because they control traffic at the instance and subnet levels, respectively, and misconfigurations here are a common cause of unreachability despite healthy instance status.

Exam trap

The trap here is that candidates assume a 'running' instance with passing status checks guarantees network reachability, overlooking that security groups and NACLs can silently drop traffic without any error in system logs or status checks.

How to eliminate wrong answers

Option A is wrong because CloudWatch metrics for CPU utilization and network throughput measure performance, not connectivity; they would not reveal whether inbound traffic is being blocked by security groups or NACLs. Option B is wrong because rebooting the instance resets the OS but does not change network configurations or underlying hardware; if the instance is running and status checks pass, a reboot is unlikely to resolve a network-level block. Option C is wrong because stopping and starting the instance moves it to new underlying hardware, which could help if the issue were hardware-related, but the status checks passing indicates the hardware is healthy; this action is more disruptive and unnecessary for a likely network configuration problem.

444
MCQhard

A company uses AWS CloudFormation StackSets to deploy a common security baseline across multiple AWS accounts. They have a new account that needs to be added to the StackSet. The StackSet is configured with self-service permissions and uses a service-managed IAM role. What must be done to include the new account?

A.Create an IAM role in the new account that trusts the StackSet.
B.Create a new StackSet that includes the new account.
C.Manually create a stack instance for the new account in the StackSet.
D.Add the new account to the AWS Organization.
AnswerD

Service-managed StackSets are tightly integrated with AWS Organizations, so adding the new account to the organization makes it a recognized target account for the StackSet. Once the account is part of the organization, CloudFormation automatically creates and manages stack instances in every configured region for that account, without requiring manual role setup or stack creation. This approach preserves the existing StackSet's configuration and operational tracking, and it is the intended mechanism for onboarding a new account to a service-managed StackSet.

Why this answer

When a StackSet uses service-managed permissions, it relies on AWS Organizations to manage accounts. To include a new account, you must add it to the AWS Organization; the StackSet will automatically deploy stack instances to that account based on the specified organizational units (OUs) or accounts. This is because service-managed permissions use a service-linked role created and managed by CloudFormation, not a manually created IAM role in the target account.

Exam trap

The trap here is that candidates often confuse self-service permissions (which require manual IAM role creation and stack instance management) with service-managed permissions (which rely on AWS Organizations for automatic account and stack instance management), leading them to choose Option A or C incorrectly.

How to eliminate wrong answers

Option A is wrong because with service-managed permissions, CloudFormation automatically creates and manages the necessary IAM role in the target account via a service-linked role; you do not need to manually create an IAM role. Option B is wrong because creating a new StackSet is unnecessary; you can add the new account to the existing StackSet by adding it to the Organization or updating the StackSet's target accounts/OUs. Option C is wrong because manually creating a stack instance is not possible with service-managed permissions; the StackSet automatically manages stack instances based on the Organization structure, and manual creation is only available with self-service permissions.

445
Multi-Selectmedium

A company uses AWS Organizations with SCPs to enforce security policies. The security team needs to ensure that no IAM user or role can disable AWS CloudTrail or delete CloudTrail logs. Which TWO approaches should be combined to achieve this? (Choose TWO.)

Select 2 answers
A.Use a service control policy to deny s3:DeleteObject on the CloudTrail S3 bucket.
B.Enable MFA Delete on the CloudTrail S3 bucket.
C.Apply an SCP that denies cloudtrail:StopLogging and cloudtrail:DeleteTrail for all accounts.
D.Enable CloudTrail log file validation.
E.Attach an IAM policy to all users denying cloudtrail:StopLogging.
AnswersA, C

An SCP denying s3:DeleteObject on the CloudTrail bucket prevents any account within the organization from deleting or overwriting the log objects that CloudTrail writes. This augments the previous SCP by protecting the evidence trail itself, not just the trail configuration. It is a preventive control at the organizational boundary, ensuring that even an administrator who manages to disable the trail cannot destroy historical log data that would reveal the actions.

Why this answer

To prevent IAM users and roles from disabling CloudTrail or deleting logs, a combination of two preventive controls is needed. Option A applies a service control policy (SCP) that denies s3:DeleteObject on the CloudTrail S3 bucket, preventing log deletion at the organizational level. Option C applies an SCP that denies cloudtrail:StopLogging and cloudtrail:DeleteTrail, preventing disabling of CloudTrail itself.

Both SCPs cannot be overridden by account administrators. Option D (log file validation) is a detective control—it detects tampering or deletion after it occurs, but does not prevent the action. Options B (MFA Delete) and E (IAM policy) are weaker or can be overridden, making them less reliable for this requirement.

Exam trap

Candidates often mistake detective controls (like log validation) for preventive controls. The question specifically asks for approaches that ensure no user or role can disable or delete—this requires preventive measures that block the action entirely.

446
MCQmedium

A DevOps engineer is troubleshooting an issue where an EC2 instance running a web application becomes unresponsive every few hours. CloudWatch logs show no application errors, but the instance's status checks are passing. The engineer suspects a memory leak. Which AWS service can be used to capture memory utilization metrics at a granular level to confirm the leak?

A.EC2 Status Checks
B.AWS Config
C.AWS CloudTrail
D.CloudWatch Agent
AnswerD

The unified CloudWatch Agent is the correct solution because it runs as a service on your EC2 instance and reads guest-OS metrics directly from /proc and the operating system, allowing it to report memory utilization as a custom metric to CloudWatch. You can install and configure it via the command line, SSM Run Command, or an AWS-supplied recipe, and attach an IAM role with `cloudwatch:PutMetricData` permissions. It also supports collecting disk usage, CPU per-core statistics, and logs, all of which are invisible to hypervisor-level default EC2 metrics.

Why this answer

The CloudWatch Agent can be installed on EC2 instances to collect custom metrics such as memory usage and send them to CloudWatch, allowing granular monitoring. Option A is incorrect because EC2 Status Checks only verify system status (e.g., network reachability). Option B is incorrect because AWS Config tracks configuration changes, not performance metrics.

Option C is incorrect because CloudTrail records API activity, not system metrics.

447
MCQmedium

A company runs a standard AWS CodePipeline with a CodeCommit source and a CodeBuild build stage. Compliance requires that every source revision be built with a buildspec that is immutable and shared across all pipelines in the organization; developers must not be able to change build instructions by editing files in the source repository. Which approach enforces this requirement with the least operational overhead?

A.Move the buildspec into an Amazon S3 object and use the S3 object URL in the CodeBuild project's source configuration.
B.Create a separate CodeCommit repository that holds only buildspec.yml and add it as a second source action with a lower runOrder than the application source.
C.Store the buildspec.yml in the CodeCommit repository and protect the branch with an approval rule template that requires senior approval.
D.Configure the CodeBuild project to use an inline buildspec defined in the project, and reference that project from every pipeline stage.
AnswerD

An inline buildspec is stored as part of the CodeBuild project configuration, not in the source repository, so source edits cannot change build instructions. Pipelines that reference this project inherit the same immutable definition, satisfying the shared and immutable requirement with no extra repository controls or per-pipeline duplication.

Why this answer

Storing the buildspec inline in the CodeBuild project makes build instructions part of the project definition rather than the source tree, so source repository changes cannot affect them. Because pipelines reference the same CodeBuild project, the definition is shared and consistent across the organization, which satisfies both immutability and low operational overhead without per-repository controls.

Exam trap

The trap here is assuming that protecting the source branch with approvals makes the buildspec immutable, when the buildspec is still source-controlled content that can change on any branch or via a merge.

448
MCQeasy

A company uses CloudWatch Synthetics canaries to monitor a critical API endpoint. Recently, a canary started failing with a '403 Forbidden' error. The DevOps engineer verifies that the canary's IAM role has the necessary permissions to invoke the API and that the API endpoint is publicly accessible. What should the engineer check NEXT?

A.Review the canary's CloudWatch Logs for any runtime errors.
B.Increase the canary's memory to 512 MB to prevent timeout-related issues.
C.Check if the API requires an API key or other authentication that the canary is not providing.
D.Verify that the canary is attached to the correct VPC and subnet.
AnswerC

A 403 Forbidden response from an API means the server understood the request but refused to authorize it; public APIs commonly require an API key, an Authorization header, or a signed payload, and the canary's HTTP request may be missing those credentials. Verify the canary script attaches the API key (e.g., x-api-key header) or uses the same authentication mechanism as your test client; also ensure the key is valid and not expired.

Why this answer

A 403 Forbidden error indicates that the request reached the API but was denied due to authentication or authorization issues. Since the canary's IAM role has the necessary permissions and the endpoint is publicly accessible, the most likely cause is that the API requires an API key or other authentication mechanism (e.g., a token) that the canary is not providing. Checking for missing authentication is the logical next step.

Exam trap

DOP-C02 often tests the distinction between IAM permissions and application-level authentication; candidates may assume that IAM roles are sufficient for API access, ignoring API keys or other auth mechanisms.

How to eliminate wrong answers

Option A is wrong because runtime errors would typically produce different error codes (e.g., 500 or timeout) and the 403 specifically points to an authentication/authorization issue, not a code error. Option B is wrong because increasing memory addresses timeout issues, which would manifest as timeouts, not 403 errors. Option D is wrong because VPC/subnet misconfiguration would result in network connectivity failures (e.g., timeouts or connection refused), not a 403 response from the API.

449
MCQmedium

A company runs a containerized application on Amazon EKS. They want to ensure that if a node fails, the pods are rescheduled on healthy nodes. Which configuration is necessary?

A.Configure a pod disruption budget to prevent too many pods from being terminated simultaneously.
B.Use a horizontal pod autoscaler to increase the number of pods during high load.
C.Configure the EKS managed node group with a health check and ensure that the Kubernetes control plane automatically reschedules pods from failed nodes.
D.Use a cluster autoscaler to automatically add new nodes when pods are pending.
AnswerC

An EKS managed node group is backed by an Auto Scaling group whose health checks include both Amazon EC2 status checks and Kubernetes node status (including the node-lost and NodeReady conditions). When a node fails or becomes unhealthy, the Auto Scaling group replaces the underlying instance, while Kubernetes' node controller marks the node as NotReady and evicts pods, leading them to be rescheduled onto healthy nodes. This combination of automatic instance replacement and control-plane-driven pod rescheduling directly addresses the failure of a node and is the expected solution for maintaining availability.

Why this answer

EKS managed node groups automatically register nodes with the Kubernetes control plane, and the Kubernetes node controller (part of the kube-controller-manager) monitors node health via the NodeLifecycleController. When a node fails (e.g., due to an EC2 instance termination or health check failure), the control plane marks the node as `NotReady` and, after the default pod eviction timeout (5 minutes), evicts pods from the failed node, rescheduling them on healthy nodes. This behavior is inherent to Kubernetes and does not require additional configuration beyond using a managed node group.

Exam trap

The trap here is that candidates confuse the Cluster Autoscaler (which adds nodes) with the node controller's pod rescheduling behavior, or they think a PodDisruptionBudget is needed for failure recovery when it only applies to voluntary disruptions.

How to eliminate wrong answers

Option A is wrong because a PodDisruptionBudget (PDB) controls voluntary disruptions (e.g., node drains during updates) and does not handle involuntary node failures; it would actually prevent pods from being rescheduled if the PDB's minAvailable or maxUnavailable constraints are violated. Option B is wrong because a HorizontalPodAutoscaler (HPA) scales the number of pod replicas based on CPU/memory metrics, not in response to node failures; it does not reschedule pods from failed nodes. Option D is wrong because the Cluster Autoscaler adds new nodes when pods are unschedulable due to resource constraints, but it does not reschedule pods from failed nodes—that is the responsibility of the Kubernetes node controller and kube-scheduler.

450
Multi-Selectmedium

A DevOps team is investigating a production incident where an Amazon RDS for MySQL database experienced a sudden spike in connections and CPU utilization. The team suspects a SQL injection attack. Which TWO actions should the team take to investigate and mitigate the incident?

Select 2 answers
A.Delete the error logs to free up storage space.
B.Enable automated backups and ensure point-in-time recovery is configured.
C.Enable RDS Enhanced Monitoring and audit logs to capture SQL queries.
D.Create a read replica to offload traffic from the primary instance.
E.Increase the DB instance size to handle the increased load.
AnswersB, C

Enabling automated backups with point-in-time recovery (PITR) is the correct immediate response because it lets you restore the database to any second within the retention window, such as a timestamp just before the compromise occurred. Automated backups take daily snapshots, and RDS continuously records transaction logs to enable PITR, minimizing data loss if tables were dropped or data was tampered with. This is the foundational recovery mechanism to return the system to a known-good state.

Why this answer

Enabling automated backups and point-in-time recovery ensures that the database can be restored to a state before the suspected SQL injection attack, preserving data integrity and enabling forensic analysis. Option C is correct because RDS Enhanced Monitoring provides OS-level metrics (CPU, memory, disk I/O) to correlate with the spike, while audit logs capture actual SQL queries, which are essential for identifying malicious patterns and confirming the attack vector.

Exam trap

The trap here is that candidates confuse reactive scaling (Option E) or read replicas (Option D) with proper incident response, failing to recognize that investigation and mitigation require enabling logging and backup capabilities, not just increasing capacity.

Page 5

Page 6 of 18

Page 7