Courseiva

AWS Certified DevOps Engineer Professional DOP-C02 (DOP-C02) — Questions 1–75

1298 questions total · 18pages · All types, answers revealed

Page 1 of 18

Page 2
1
MCQeasy

A team uses AWS CloudFormation to manage infrastructure. They want to deploy a stack that creates an S3 bucket and a DynamoDB table. The S3 bucket name must be unique across all AWS accounts. Which CloudFormation intrinsic function should be used to generate a unique bucket name?

A.!Ref 'AWS::StackName'
B.!GetAtt S3Bucket.Arn
C.!Sub 'mybucket-${AWS::AccountId}'
D.!Select [0, !Split ['-', !Ref 'AWS::Region']]
AnswerC

AWS::AccountId is a true pseudo parameter that is resolved during template evaluation, providing a unique, numeric identifier for the current AWS account. When combined with a fixed prefix like 'mybucket-', the resulting bucket name is globally unique across all accounts and regions because each account has a distinct ID, satisfying S3's global namespace requirement. This approach is deterministic, readable, and avoids the circular dependency issues of resource attributes, making it a widely adopted best practice for naming resources.

Why this answer

The `!Sub 'mybucket-${AWS::AccountId}'` intrinsic function substitutes the AWS::AccountId pseudo parameter, which is guaranteed to be unique per AWS account. Since S3 bucket names must be globally unique across all AWS accounts, appending the account ID ensures the generated name does not conflict with buckets in other accounts. This approach is a common pattern for creating unique resource names in CloudFormation.

Exam trap

The trap here is that candidates may think `!Ref 'AWS::StackName'` or `!Ref 'AWS::Region'` provide sufficient uniqueness, but they overlook the requirement for global uniqueness across all AWS accounts, which only `AWS::AccountId` guarantees.

How to eliminate wrong answers

Option A is wrong because `!Ref 'AWS::StackName'` returns the name of the CloudFormation stack, which is not guaranteed to be unique across AWS accounts—multiple accounts can have stacks with the same name. Option B is wrong because `!GetAtt S3Bucket.Arn` returns the Amazon Resource Name of the S3 bucket, which is only available after the bucket is created, and cannot be used to generate a name before creation. Option D is wrong because `!Select [0, !Split ['-', !Ref 'AWS::Region']]` extracts the first part of the region name (e.g., 'us' from 'us-east-1'), which is not unique across accounts or even across regions within the same account.

2
MCQmedium

A DevOps engineer is designing a CI/CD pipeline for an application that runs on Amazon ECS. The pipeline should automatically build a Docker image from source code, push it to Amazon ECR, and deploy it to the ECS service. The engineer wants to use AWS CodePipeline with Amazon ECR as a source. Which action provider should be used for the deploy stage?

A.AWS CloudFormation
B.AWS CodeDeploy
C.Amazon ECS
D.AWS Elastic Beanstalk
AnswerC

The Amazon ECS deploy action provider is the native integration in CodePipeline for pushing a new container image or task definition to an existing ECS service. When this action runs, it registers the revised task definition and calls UpdateService so the service starts a deployment with the new image. This makes it the correct, minimal-action choice for an ECS-based application pipeline, as it avoids abstracting the deployment behind another compute platform.

Why this answer

Amazon ECS is the correct action provider for the deploy stage when using AWS CodePipeline because it directly integrates with ECS to update an existing ECS service with a new task definition. When Amazon ECR is configured as a source, CodePipeline detects image pushes, and the ECS deploy action automatically triggers a rolling update of the ECS service using the new image, without requiring additional deployment tools.

Exam trap

The trap here is that candidates often confuse the Amazon ECS deploy action with AWS CodeDeploy, but CodeDeploy is only required for ECS blue/green deployments, while the Amazon ECS action directly handles rolling updates and is the simplest choice for standard ECS service deployments.

How to eliminate wrong answers

Option A is wrong because AWS CloudFormation is an infrastructure-as-code service used to provision and manage AWS resources, not a direct deploy action for updating an ECS service with a new Docker image; using it would require custom scripts and adds unnecessary complexity. Option B is wrong because AWS CodeDeploy can deploy to ECS only when using the ECS blue/green deployment type, which requires an additional CodeDeploy application and deployment group setup, but the question asks for the simplest direct deploy action provider for ECS, which is Amazon ECS itself. Option D is wrong because AWS Elastic Beanstalk is a PaaS service for deploying web applications, not designed for ECS service deployments; it abstracts underlying infrastructure and does not provide a native action to update an ECS service with a new image from ECR.

3
MCQhard

A DevOps engineer is troubleshooting a CloudFormation stack creation failure. The stack includes an AWS::EC2::Instance with a UserData script. The stack creation fails with the error: 'The following resource(s) failed to create: [EC2Instance]. The requested configuration is currently not supported. Please check the documentation for supported configurations.' The engineer suspects the instance type is not supported in the selected Availability Zone. Which action should the engineer take to resolve this issue and ensure successful stack creation?

A.Add an Availability Zone parameter and map it to an AZ that supports the instance type, or use the AWS::EC2::Instance AvailabilityZone property to specify an AZ that supports the instance type.
B.Use AWS OpsWorks to deploy the instance instead of CloudFormation.
C.Change the instance type to a previous generation that is supported in all AZs.
D.Modify the template to specify a different region that supports the instance type.
AnswerA

The error occurs because CloudFormation did not explicitly specify an Availability Zone, so AWS placed the instance in an AZ that does not offer the selected instance type. By adding an Availability Zone parameter that is constrained to zones where the instance type is supported—or by setting the AvailabilityZone property directly on the AWS::EC2::Instance resource—you force EC2 to launch in a compatible AZ. This addresses the root cause without changing the instance type, region, or underlying architecture.

Why this answer

The error indicates the instance type is not supported in the default Availability Zone (AZ) selected by CloudFormation. By explicitly specifying an AZ that supports the instance type via the `AvailabilityZone` property or by parameterizing the AZ, the engineer can override the default AZ selection and resolve the incompatibility. This approach directly addresses the root cause without changing the instance type or region.

Exam trap

The trap here is that candidates may assume the instance type itself is unsupported in the entire region and choose to change the region or instance type, rather than recognizing that the error is AZ-specific and can be resolved by explicitly specifying a supported Availability Zone.

How to eliminate wrong answers

Option B is wrong because AWS OpsWorks is a configuration management service that does not solve the underlying AZ compatibility issue; it would still rely on the same EC2 instance type and AZ constraints. Option C is wrong because changing to a previous generation instance type is unnecessary and may not be supported in all AZs either; the goal is to keep the desired instance type and find a compatible AZ. Option D is wrong because specifying a different region is an overreaction; the issue is AZ-specific within the current region, and changing the region may introduce other compatibility or latency issues.

4
MCQhard

A company runs a critical application on AWS Lambda that processes messages from an Amazon SQS queue. The application must be resilient to downstream service failures. The team notices that when the downstream service is unhealthy, messages are repeatedly retried and eventually sent to the dead-letter queue (DLQ) before the service recovers. What design change would improve resilience by allowing automatic retries after the downstream service recovers?

A.Configure the SQS queue with a large visibility timeout (e.g., 6 hours) and use a redrive policy only after a high number of receives. Keep the messages in the queue and retry when the downstream service becomes healthy.
B.Reduce the maxReceiveCount to 1 so that messages are sent to DLQ immediately, then reprocess them from DLQ later.
C.Increase the message retention period to 14 days and use a DLQ with high retention.
D.Use Amazon SNS to fan out messages to multiple SQS queues, each with different retry policies.
AnswerA

This approach leverages the SQS visibility timeout as a built-in retry buffer: setting it to 6 hours prevents messages from being redelivered to consumers while the downstream service is unhealthy, effectively holding them in the queue. Combined with a high maxReceiveCount (e.g., 100) in the redrive policy, messages are not prematurely moved to the DLQ, allowing the Lambda consumer to keep retrying for hours until the service recovers. This ensures no data loss and avoids manual intervention, making it the correct pattern for intermittent downstream outages.

Why this answer

Increasing the visibility timeout to a long duration (e.g., 6 hours) prevents messages from being repeatedly retried and sent to the DLQ while the downstream service is unhealthy. Instead, messages remain in the SQS queue and become visible again only after the visibility timeout expires, allowing automatic retries once the downstream service recovers. This approach avoids premature DLQ delivery and leverages SQS's built-in redrive policy based on maxReceiveCount.

Exam trap

The trap here is that candidates often think increasing the DLQ retention or reducing retries (maxReceiveCount) is the solution, but the real key is controlling the retry timing via the visibility timeout to allow the downstream service to recover before messages are exhausted.

How to eliminate wrong answers

Option B is wrong because reducing maxReceiveCount to 1 sends messages to the DLQ immediately after the first failure, which defeats resilience by not allowing any retries and requiring manual reprocessing from the DLQ. Option C is wrong because increasing the message retention period and using a DLQ with high retention does not prevent messages from being sent to the DLQ prematurely; it only keeps them in the DLQ longer, but the downstream service may recover before the messages are consumed from the DLQ. Option D is wrong because using SNS to fan out to multiple SQS queues with different retry policies adds complexity and does not address the core issue of preventing premature DLQ delivery; it still relies on the same visibility timeout and retry mechanism.

5
MCQeasy

A company wants to visualize the performance of their application running on EC2. They need to create a dashboard that shows CPU utilization, memory usage, and disk I/O. Which AWS service should they use?

A.Amazon CloudWatch Dashboards.
B.AWS CloudTrail.
C.AWS Systems Manager.
D.Amazon QuickSight.
AnswerA

Amazon CloudWatch Dashboards is the correct choice because it natively integrates with CloudWatch, which collects and stores EC2 standard metrics (CPU, network, disk) and, when the CloudWatch agent is installed on the instance, also collects custom operating-system-level metrics such as memory utilization and disk space. These dashboards can be customized to display multiple metrics in a single graph, aggregated across instances, and automatically refreshed for near-real-time monitoring. Additionally, CloudWatch Dashboards support alarm overlays and anomaly detection, making it purpose-built for visualizing application performance as a monitoring and observability tool.

Why this answer

Amazon CloudWatch Dashboards is the correct choice because it is the native AWS service for visualizing operational metrics such as CPU utilization, memory usage, and disk I/O. CloudWatch collects these metrics from EC2 instances (with the CloudWatch agent required for memory and disk metrics at the OS level) and allows you to create customizable, shareable dashboards. It supports cross-account, cross-region views and integrates with alarms and logs, making it the standard tool for performance monitoring.

The other services do not provide metric visualization for EC2 performance.

Exam trap

DOP-C02 often tests the distinction between monitoring (CloudWatch) and auditing (CloudTrail), or between operational dashboards and business intelligence (QuickSight), so candidates must remember that memory and disk metrics require the CloudWatch agent, not just default EC2 monitoring.

How to eliminate wrong answers

Option B is wrong because AWS CloudTrail records API activity and audit logs, not performance metrics like CPU or memory; it is used for governance, compliance, and operational auditing. Option C is wrong because AWS Systems Manager is a management suite for patching, automation, and inventory, not a dashboarding or metric visualization service, though it can collect some inventory data. Option D is wrong because Amazon QuickSight is a business intelligence service for analyzing data from various sources, not a real-time operational monitoring tool for EC2 metrics; it lacks native integration with CloudWatch metrics for live dashboards.

6
MCQhard

An e-commerce application runs on Amazon ECS with Fargate. The operations team notices that the application's latency increases during peak hours. The engineer needs to correlate high CPU usage with increased request latency to identify the root cause. Which approach should be used?

A.Use CloudWatch Logs Insights to query container logs
B.Enable Container Insights and ServiceLens to correlate metrics and traces
C.Configure CloudWatch Synthetics canaries to measure latency
D.Set up a Prometheus server on an EC2 instance to scrape container metrics
AnswerB

For Fargate, Container Insights must be enabled to collect infrastructure telemetry such as CPU and memory utilization at the cluster, service, task, and container levels; AWS then ships these metrics to CloudWatch automatically. ServiceLens builds on that by combining CloudWatch metrics, logs, and X-Ray traces into a single service map and console, allowing you to select a trace and see the corresponding CPU/memory metric trends for that underlying container. This native integration is exactly what is required to prove whether a CPU saturation event is causing an observed latency increase in a traced request.

Why this answer

Amazon CloudWatch Container Insights collects, aggregates, and summarizes metrics and logs from containerized applications on ECS/Fargate, including CPU utilization. ServiceLens integrates CloudWatch with AWS X-Ray traces, allowing you to correlate metrics with request traces and latency. Together they provide the end-to-end visibility needed to link high CPU usage with increased request latency.

Exam trap

DOP-C02 often tests whether candidates confuse monitoring tools that collect metrics (Container Insights) with those that trace requests (X-Ray/ServiceLens) or measure external latency (Synthetics), missing the need for correlation.

How to eliminate wrong answers

Option A is wrong because CloudWatch Logs Insights queries log data but does not natively correlate metrics with distributed traces or provide the integrated latency-to-CPU correlation. Option C is wrong because Synthetics canaries measure endpoint latency from outside but do not correlate that latency with internal CPU metrics or traces. Option D is wrong because a self-managed Prometheus server adds operational overhead and does not natively integrate with X-Ray traces or CloudWatch ServiceLens for correlation.

7
MCQmedium

A company is deploying a critical microservice on Amazon ECS with Fargate. They need to ensure that the service can tolerate an Availability Zone failure. What is the BEST approach?

A.Use a cluster placement constraint to spread tasks across instances
B.Use EC2 launch type and spread tasks across instance types
C.Define the service to spread tasks across multiple Availability Zones
D.Configure service auto scaling to add tasks when CPU is high
AnswerC

Defining the service to spread tasks across multiple Availability Zones ensures that the microservice remains available even if a single AZ fails. In Amazon ECS, you can use a placement strategy with 'spread' across 'attribute:ecs.availability-zone' for EC2, or for Fargate you can use the `AvailabilityZoneSpread` placement strategy (or specify AZs in the service definition) to distribute tasks evenly across AZs. This directly mitigates the risk of an AZ outage by maintaining running tasks in other zones.

Why this answer

Amazon ECS with Fargate allows you to define a service with a 'spread across Availability Zones' strategy. By setting the service's placement strategy to spread tasks across multiple Availability Zones, the service automatically distributes tasks evenly across the specified AZs. If one AZ fails, the remaining tasks in other AZs continue to serve traffic, ensuring high availability.

This is the most direct and effective method for tolerating an AZ failure in a Fargate-based deployment.

Exam trap

The trap here is that candidates often confuse 'spreading tasks across instances' (Option A) or 'across instance types' (Option B) with the correct concept of spreading across Availability Zones, or they mistakenly think auto scaling (Option D) alone provides AZ failure tolerance, when in fact it only scales based on load and does not guarantee multi-AZ distribution.

How to eliminate wrong answers

Option A is wrong because cluster placement constraints are used with the EC2 launch type to control task placement on specific instances, not to spread tasks across Availability Zones; Fargate manages the underlying infrastructure, so placement constraints are not applicable. Option B is wrong because the EC2 launch type and spreading tasks across instance types addresses instance-level diversity, not AZ-level resilience; it does not protect against an entire AZ failure. Option D is wrong because service auto scaling based on CPU handles performance scaling, not fault tolerance; it does not distribute tasks across AZs and cannot recover from an AZ failure without pre-existing multi-AZ distribution.

8
MCQeasy

A company uses AWS Systems Manager to manage configuration compliance. They want to ensure that all EC2 instances have a specific security patch installed. Which Systems Manager capability should they use?

A.Automation
B.Parameter Store
C.State Manager
D.Patch Manager
AnswerD

Patch Manager is the purpose-built Systems Manager capability that automates the process of patching managed nodes. It uses patch baselines to determine which patches are approved or rejected, and it can either scan for missing patches or install them directly on EC2 instances and on-premises servers. Patch Manager supports both Linux and Windows, integrates with maintenance windows and compliance reporting, and is the correct answer because it directly handles patch configuration and deployment.

Why this answer

Patch Manager is the correct Systems Manager capability because it is specifically designed to automate the process of patching managed instances with security updates and other types of patches. It can scan for missing patches and enforce compliance by applying patches on a schedule or on demand, directly addressing the requirement to ensure all EC2 instances have a specific security patch installed.

Exam trap

The trap here is that candidates often confuse State Manager (which manages configuration state) with Patch Manager (which specifically handles patch compliance), or they assume Automation can handle patching because it can run scripts, but Patch Manager provides the dedicated patching workflow and compliance reporting.

How to eliminate wrong answers

Option A is wrong because Automation is used to automate common maintenance and deployment tasks (e.g., creating AMIs, running scripts) but does not natively handle patch scanning or installation workflows. Option B is wrong because Parameter Store is a secure hierarchical store for configuration data and secrets, not a patching mechanism. Option C is wrong because State Manager is used to define and maintain consistent configuration of instances (e.g., ensuring a file exists or a service is running) but does not include built-in patch scanning or remediation capabilities.

9
MCQeasy

A DevOps engineer needs to manage configuration files for multiple applications across several EC2 instances. The configuration values are sensitive (e.g., database passwords) and must be encrypted at rest and in transit. Which AWS service should be used to store and retrieve these configuration values?

A.AWS Systems Manager Parameter Store (SecureString)
B.AWS CloudFormation template parameters
C.Amazon DynamoDB with encryption
D.Amazon S3 with server-side encryption
AnswerA

The AWS Systems Manager Parameter Store is purpose-built for configuration management, providing a secure, hierarchical namespace for keys and values. A SecureString parameter uses AWS KMS encryption at rest and in transit, and integrates natively with EC2 via the SSM agent or SDK without exposing plaintext. It supports versioning, tagging, and IAM policies for fine-grained access, making it the right choice for dynamic runtime retrieval.

Why this answer

AWS Systems Manager Parameter Store with the SecureString parameter type uses AWS KMS to encrypt values at rest and enforces TLS for values in transit, making it purpose-built for storing sensitive configuration such as database passwords. It integrates natively with EC2 instance profiles and IAM, so applications can retrieve secrets without embedding credentials in code or AMIs. This is the canonical DOP-C02 answer for encrypted configuration management.

Exam trap

DOP-C02 often tests the distinction between a configuration store (Parameter Store) and a general-purpose data store (S3, DynamoDB), so candidates pick S3 SSE or DynamoDB encryption thinking 'encrypted at rest' is the only requirement.

How to eliminate wrong answers

Option B is wrong because CloudFormation template parameters are only input values passed at stack creation/update time; they are not a persistent, encrypted-at-rest store and are visible in stack metadata and change sets. Option C is wrong because DynamoDB is a NoSQL database, not a configuration/secret store — encryption at rest exists but there is no SecureString semantics, no native parameter versioning, and it adds unnecessary cost and operational overhead. Option D is wrong because S3 server-side encryption protects objects at rest but does not provide a configuration-management API, parameter versioning, or the IAM-scoped retrieval model that Parameter Store offers.

10
MCQhard

A company runs a critical e-commerce application on AWS. They use AWS CodePipeline to manage deployments. The pipeline has a source stage (CodeCommit), a build stage (CodeBuild), and a deploy stage (CodeDeploy to an Auto Scaling group). Recently, a deployment caused a 5-minute outage because the new application version had a bug that caused the health checks to fail. The Auto Scaling group marked instances as unhealthy and replaced them, but during the replacement, traffic was routed to the remaining instances, which also failed health checks, causing a full outage. The company wants to implement a deployment strategy that prevents any traffic from being routed to unhealthy instances and automatically rolls back if the deployment fails. They also want to minimize deployment time and cost. Which solution should the DevOps team implement?

A.Add a manual approval step in CodePipeline before deploy
B.Use CodeDeploy in-place deployment with automatic rollback enabled
C.Use CodeDeploy blue/green deployment with automatic rollback enabled
D.Increase the health check grace period in the Auto Scaling group
AnswerC

Blue/green shifts traffic to a new fleet only after it passes health checks, so no traffic reaches unhealthy instances, and CodeDeploy automatic rollback reverts on failure. This satisfies the zero-traffic-to-unhealthy-instances and automatic rollback constraints while avoiding full in-place replacement cost.

Why this answer

The correct solution is to use a blue/green deployment with CodeDeploy and automatic rollback enabled. In a blue/green deployment, a new Auto Scaling group (green) is created alongside the existing one (blue). Traffic is shifted to the green group only after all health checks pass.

If health checks fail, the deployment is automatically rolled back by terminating the green group, ensuring no traffic is routed to unhealthy instances. This prevents any outage. Option B (in-place deployment with rollback) updates instances in place, which can cause downtime if instances fail health checks, as the Auto Scaling group replaces them sequentially, potentially routing traffic to unhealthy instances.

Option A (manual approval) slows down deployment and does not automate rollback based on health checks. Option D (increasing health check grace period) only delays detection of failures and does not prevent traffic from being routed to unhealthy instances.

11
MCQhard

A company uses AWS CodeBuild to build a Java application. The buildspec.yml file includes a post_build phase that runs a script to package the application and upload it to an Amazon S3 bucket. The build is failing with an error indicating that the S3 bucket is not accessible. The CodeBuild service role has the AmazonS3FullAccess managed policy attached. The S3 bucket is in a different AWS account. Which action should the DevOps engineer take to resolve the issue?

A.Add a bucket policy to the S3 bucket that grants the CodeBuild service role permissions to PutObject.
B.Configure the S3 bucket to allow public write access.
C.Enable cross-account access in the CodeBuild project settings.
D.Modify the CodeBuild service role to include the s3:PutObject permission for the specific bucket ARN.
AnswerA

When the S3 bucket is in a different account, the CodeBuild service role needs explicit permission from the bucket owner. The AmazonS3FullAccess policy grants permissions within the same account, but cross-account access requires a bucket policy that allows the role to perform actions. Adding a bucket policy that grants PutObject to the CodeBuild service role resolves the access issue.

Why this answer

Cross-account access to S3 requires a bucket policy that explicitly grants permissions to the external IAM role. The CodeBuild service role already has S3 permissions in its own account, but the bucket policy in the target account must allow the role to perform PutObject. Therefore, adding a bucket policy is the correct resolution.

Exam trap

The trap here is assuming that attaching AmazonS3FullAccess to the CodeBuild service role is sufficient for cross-account access, but it only grants permissions within the same account.

12
MCQmedium

A company uses Amazon S3 to store critical data. An incident occurs where an S3 bucket is accidentally deleted. The DevOps engineer needs to recover the bucket and its objects. What should the engineer do?

A.Restore the bucket from AWS CloudTrail event history
B.Contact AWS Support to restore the bucket from a backup if versioning was enabled
C.Recreate the bucket with the same name and restore objects from a previous backup
D.Use the AWS S3 console to undo the deletion
AnswerC

The only way to recover from a deleted bucket is to recreate it and restore objects from a backup. (Note: bucket name uniqueness may require waiting or using a different name.)

Why this answer

S3 bucket names are globally unique and, once deleted, cannot be restored by AWS — the bucket and its configuration are gone. The correct recovery approach is to recreate the bucket with the same name (if the name is still available) and restore the objects from a backup or from versioned copies if versioning was enabled. This is the only viable path to recovery.

Exam trap

DOP-C02 often tests whether candidates understand that S3 bucket deletion is irreversible and that versioning protects objects, not the bucket itself, leading candidates to incorrectly believe AWS Support or CloudTrail can restore a deleted bucket.

How to eliminate wrong answers

Option A is wrong because CloudTrail event history records API activity (who deleted the bucket) but does not store bucket contents or enable restoration. Option B is wrong because AWS Support cannot restore a deleted S3 bucket — AWS does not retain deleted buckets, and Support has no mechanism to recover them. Option D is wrong because there is no 'undo deletion' feature in the S3 console; deletion is permanent.

13
MCQmedium

A DevOps engineer notices that an EC2 instance running a critical web application has been terminated unexpectedly. The instance was part of an Auto Scaling group. Which step should the engineer take FIRST to investigate the root cause?

A.Review AWS CloudTrail logs for TerminateInstances API calls.
B.Look at the EC2 console's 'Termination Protection' setting.
C.Check the application logs on the instance's attached EBS volume (detached and attached to another instance).
D.Verify the Auto Scaling group's scaling policies and scheduled actions.
AnswerA

AWS CloudTrail is the authoritative source for who, when, and how an EC2 instance was terminated: every StopInstances/TerminateInstances API call is recorded as an event with the IAM user or role, source IP, user agent, and request parameters. Console clicks, CLI commands, SDK calls, and automated actions by services such as Auto Scaling or AWS Lambda all generate a TerminateInstances event. Reviewing CloudTrail event history or querying the CloudTrail S3 bucket with Athena reveals the entity that issued the termination and any accompanying error or access-denied information.

Why this answer

CloudTrail records all AWS API activity, including TerminateInstances calls, capturing the identity, source IP, timestamp, and whether the termination was user-initiated, by an Auto Scaling policy, or by another service. Reviewing CloudTrail first establishes who or what triggered the termination before examining instance-level artifacts.

Exam trap

The trap is jumping to Auto Scaling policies or instance settings as the cause — the exam expects you to know CloudTrail is the authoritative source for API-level 'who did what' before examining downstream artifacts.

How to eliminate wrong answers

Option B is wrong because Termination Protection is a setting that, if enabled, would have prevented the termination — checking it after the fact does not explain why the instance was terminated and is not the first investigative step. Option C is wrong because application logs on the EBS volume may show application behavior but not the cause of termination, and detaching/attaching volumes is a later forensic step, not the first. Option D is wrong because scaling policies and scheduled actions are a possible cause, but CloudTrail will reveal whether they were the actual trigger — checking policies without the API audit trail is speculative.

14
MCQhard

A company uses AWS Elastic Beanstalk to deploy a web application. The application requires environment-specific configuration values (database URL, API keys) that must be stored securely and rotated automatically. The team uses AWS Secrets Manager. Which configuration management strategy should the team implement to securely inject secrets into the Elastic Beanstalk environment?

A.Store the secrets in the Elastic Beanstalk environment configuration as plain text under 'aws:elasticbeanstalk:application:environment'.
B.Configure Secrets Manager to automatically push secrets to Elastic Beanstalk environment properties.
C.Use an Elastic Beanstalk platform hook script that retrieves secrets from Secrets Manager and sets them as environment variables.
D.Use AWS CloudFormation dynamic references to inject secrets into the Elastic Beanstalk environment.
AnswerC

An Elastic Beanstalk platform hook, such as a script in the .platform/hooks/postdeploy directory, runs on the instance during deployment and can call the AWS CLI or SDK to retrieve secrets from Secrets Manager using the instance role's IAM permissions. The script can then export the secret values as environment variables into the application runtime context (e.g., by writing to /etc/profile.d or a systemd environment file) before the app starts. This keeps secrets out of the environment configuration and source code, and it supports rotation by re-running the hook on subsequent deployments.

Why this answer

Elastic Beanstalk platform hooks allow custom scripts to run during deployment, enabling retrieval of secrets from AWS Secrets Manager and setting them as environment variables before the application starts. This approach keeps secrets out of the environment configuration and supports automatic rotation by having the script fetch the latest secret value on each deployment.

Exam trap

The trap here is that candidates often assume CloudFormation dynamic references (Option D) are the best fit for automatic rotation, but they only inject secrets at deployment time and do not handle in-place rotation without a stack update, whereas platform hooks can be used to fetch the latest secret on every instance start or deployment.

How to eliminate wrong answers

Option A is wrong because storing secrets as plain text in the Elastic Beanstalk environment configuration under 'aws:elasticbeanstalk:application:environment' exposes them in the environment properties, which can be viewed by anyone with access to the environment configuration and does not support automatic rotation. Option B is wrong because Secrets Manager does not have a native capability to automatically push secrets to Elastic Beanstalk environment properties; it requires an external mechanism (e.g., Lambda, custom script) to retrieve and set them. Option D is wrong because AWS CloudFormation dynamic references can inject secrets at stack creation or update time, but they do not handle automatic rotation of secrets within a running Elastic Beanstalk environment without additional custom logic.

15
MCQhard

An organization uses AWS OpsWorks for configuration management. They have a stack with multiple layers, including a PHP application layer and a MySQL database layer. The operations team needs to deploy a custom configuration file to all PHP application instances. How should this be accomplished using OpsWorks?

A.Use the OpsWorks agent to directly copy the file to each instance via SSH.
B.Add the configuration file as a stack-level custom cookbook and assign it to all layers.
C.Create a custom cookbook with a recipe that deploys the configuration file, and assign it to the Deploy lifecycle event of the PHP layer.
D.Use a custom JSON attribute in the stack settings to define the file content, and then use a built-in recipe to apply it.
AnswerC

A custom cookbook can contain a recipe with Chef resources like 'template' or 'file' that renders the configuration file content to the exact path expected by PHP on each instance. Assigning that recipe to the Deploy lifecycle event of the PHP layer ensures it runs only during a deployment operation on that layer, and only on instances that belong to the PHP layer. This approach is fully automated, idempotent, and version-controlled, and it aligns with OpsWorks best practices by using the layer-specific lifecycle hook to scope execution precisely.

Why this answer

OpsWorks uses Chef cookbooks to manage configuration. By creating a custom cookbook with a recipe that deploys the configuration file and assigning it to the Deploy lifecycle event of the PHP layer, the recipe runs on all PHP application instances during deployment, ensuring the file is placed correctly. This approach leverages OpsWorks' built-in lifecycle events and Chef's idempotent execution model.

Exam trap

The trap here is that candidates confuse stack-level custom cookbooks (which apply to all layers) with layer-specific lifecycle event assignments, leading them to choose Option B, which would incorrectly deploy the configuration file to the MySQL layer as well.

How to eliminate wrong answers

Option A is wrong because the OpsWorks agent does not support direct SSH file copying; OpsWorks uses Chef recipes to manage instances, and manual SSH operations bypass automation and are not scalable. Option B is wrong because stack-level custom cookbooks are assigned to layers, not to all layers automatically; assigning a cookbook to all layers would run its recipes on every instance, including the MySQL layer, which is not desired. Option D is wrong because custom JSON attributes can pass data to recipes but cannot directly deploy files; a custom recipe is required to interpret the JSON and write the file.

16
MCQhard

Refer to the exhibit. A security team reviews this CloudTrail log entry. Which finding is most concerning?

A.The event occurred in us-east-1.
B.The instance was terminated by an assumed role.
C.The source IP is from a public IP.
D.The user did not authenticate with MFA.
AnswerD

The absence of MFA in a CloudTrail session is a direct security-control failure because AWS relies on multi-factor authentication as a second factor to verify the caller's identity beyond long-term credentials. For sensitive actions like terminating an EC2 instance, missing MFA increases the risk that the request came from compromised keys or a stolen session. This makes the MFA status the key anomaly that warrants investigation, as it violates the security best practice of enforcing MFA for destructive operations.

Why this answer

The session was created without MFA (mfaAuthenticated: false). This is a security concern because the role allows console access and the user did not use MFA, increasing risk of unauthorized access. The termination is the action, but the lack of MFA is a security gap.

17
Multi-Selecthard

A company uses AWS CodeBuild to build and test their application. They want to integrate Infrastructure as Code (IaC) scanning into their build pipeline to detect security misconfigurations in CloudFormation templates before deployment. Which TWO tools or services can be used for this purpose? (Choose TWO.)

Select 2 answers
A.AWS CloudFormation Guard
B.AWS Security Hub
C.AWS Config
D.HashiCorp Terraform
E.cfn-nag
AnswersA, E

CloudFormation Guard is a policy-as-code engine from AWS that uses a Guard-specific DSL to define rules for template properties like encryption, tags, and allowed resource types. It reads CloudFormation JSON/YAML directly and can be invoked in CodeBuild via the `cfn-guard` CLI to fail the build when a violation is found. This enables automated, pre-deployment compliance checks, which is exactly what the scenario requires.

Why this answer

AWS CloudFormation Guard (cfn-guard) is a policy-as-code tool that allows you to define rules to enforce compliance and security best practices on CloudFormation templates. It can be integrated into a CodeBuild pipeline to scan templates for misconfigurations before deployment, making it a correct choice for IaC security scanning.

Exam trap

The trap here is that candidates may confuse AWS Config (which evaluates deployed resources) with a pre-deployment scanning tool, or assume Security Hub can scan templates directly, when in fact both operate on live infrastructure, not on template files.

18
MCQhard

A company uses AWS KMS to encrypt EBS volumes. The security team wants to ensure that EBS snapshots are shared with another account without exposing the underlying data. What is the correct approach?

A.Share the encrypted snapshot without modifying the KMS key policy.
B.Create an unencrypted copy of the snapshot and share it.
C.Share the encrypted snapshot and also share the KMS key with the target account.
D.Share the encrypted snapshot and update the KMS key policy to allow the target account to use the key.
AnswerD

To securely share an encrypted snapshot, the source account must share the snapshot and modify the KMS key policy to include a statement that grants the target account's root principal the kms:Decrypt and kms:CreateGrant permissions needed to use the customer managed key. After adding this cross-account authorization, the target account can create an encrypted EBS volume from the shared snapshot. The target account's IAM user or role must also have corresponding EBS and KMS permissions, but the key policy is the critical mechanism that enables cross-account decryption.

Why this answer

Sharing an encrypted EBS snapshot requires the KMS key policy to grant the target account permission to use the key (via kms:Decrypt and kms:CreateGrant). Without this, the target account cannot decrypt the snapshot to create volumes or copies. AWS KMS enforces that the key policy explicitly allows cross-account access, and the target account must have the corresponding IAM permissions.

Exam trap

The trap here is that candidates often confuse sharing the KMS key itself (which is impossible) with updating the key policy to grant cross-account usage, leading them to select Option C.

How to eliminate wrong answers

Option A is wrong because sharing an encrypted snapshot without modifying the KMS key policy denies the target account the ability to decrypt the snapshot, making it unusable. Option B is wrong because creating an unencrypted copy of an encrypted snapshot would expose the underlying data in plaintext, violating the security requirement. Option C is wrong because sharing the KMS key with the target account is not a supported operation; KMS keys cannot be shared or transferred; instead, you must update the key policy to grant cross-account usage permissions.

19
MCQmedium

A company uses AWS CodeDeploy to deploy applications to an Auto Scaling group. During a deployment, the deployment fails because the target instances are not passing the health checks. The DevOps engineer notices that the CodeDeploy agent logs show 'The overall deployment failed because too many individual instances failed deployment, too few healthy instances are available for deployment, or some instances in your deployment group are experiencing problems.' Which step should the engineer take to diagnose the issue?

A.Increase the size of the Auto Scaling group to ensure more instances are available.
B.Review the CodeDeploy deployment logs on a failed instance to identify script errors.
C.Create a CloudWatch alarm to monitor the deployment health.
D.Check AWS CloudTrail logs for the CodeDeploy API calls to see if permissions are missing.
AnswerB

The CodeDeploy agent on the failed instance writes execution logs that include the full stdout and stderr for every lifecycle event hook, such as BeforeInstall and ApplicationStop. These logs are available at /var/log/aws/codedeploy-agent/codedeploy-agent.log and under /opt/codedeploy-agent/deployment-root/deployment-logs, and they reveal the exact script command that failed and the error message. Reviewing these logs is the only way to directly observe the script's runtime behavior, making it the correct first step in troubleshooting.

Why this answer

The error message indicates that individual instances failed deployment, and the CodeDeploy agent logs on a failed instance contain detailed script output (e.g., AppSpec hooks like BeforeInstall, AfterInstall, ApplicationStart) that reveal the root cause, such as a missing dependency or incorrect file path. Reviewing these logs directly identifies script errors, which is the most targeted diagnostic step before considering infrastructure changes.

Exam trap

The trap here is that candidates often jump to infrastructure-level fixes (like scaling or monitoring) when the error message clearly points to per-instance script failures, which require examining the CodeDeploy agent logs on a failed instance.

How to eliminate wrong answers

Option A is wrong because increasing the Auto Scaling group size does not address the root cause of why instances are failing health checks; it only adds more instances that would likely fail with the same error. Option C is wrong because creating a CloudWatch alarm monitors deployment health over time but does not provide the granular, per-instance script execution details needed to diagnose why a specific deployment failed. Option D is wrong because CloudTrail logs record API calls (e.g., CreateDeployment, UpdateDeploymentGroup) and can show permission issues, but the error message explicitly states instances failed deployment, not an authorization failure; missing permissions would typically cause a different error (e.g., 'AccessDenied') during the deployment initiation, not during health checks.

20
Multi-Selecteasy

A company wants to ensure that its Amazon S3 bucket is resilient to accidental deletion of objects. Which TWO actions should be taken?

Select 2 answers
A.Enable MFA Delete on the bucket.
B.Enable S3 Object Lock.
C.Enable S3 Versioning.
D.Enable S3 Transfer Acceleration.
E.Configure a lifecycle policy to expire objects after 30 days.
AnswersA, C

MFA Delete requires an authenticated AWS user to provide both a valid AWS credential and a one-time code from an MFA device before permanently deleting an object version or toggling the versioning state. This effectively blocks accidental or malicious permanent deletion because an attacker who compromises credentials would still need physical access to the MFA token. It is a strong, targeted defense for critical S3 data, especially when combined with versioning.

Why this answer

Enabling MFA Delete on an S3 bucket requires multi-factor authentication for any delete operations, including object version deletion and bucket deletion. This adds a critical layer of protection against accidental or unauthorized deletions, as the user must present both their AWS credentials and a valid MFA code to perform these destructive actions.

Exam trap

The trap here is that candidates often confuse S3 Object Lock (which prevents overwrites and deletes during a retention period) with MFA Delete (which requires additional authentication for delete operations), but Object Lock does not protect against accidental bucket deletion or version deletion without MFA, and it is not a direct resilience mechanism for accidental deletion scenarios.

21
Multi-Selecteasy

A DevOps engineer is troubleshooting an Amazon RDS for PostgreSQL instance that is running out of storage. The engineer wants to resolve the issue without downtime. Which TWO actions can achieve this? (Choose two.)

Select 2 answers
A.Create a read replica and promote it to primary.
B.Enable storage auto scaling on the DB instance.
C.Delete old automated snapshots to free up storage.
D.Scale up the DB instance to a larger instance class.
E.Modify the DB instance to increase the allocated storage size.
AnswersB, E

Enabling storage auto scaling on the DB instance automatically increases the allocated storage when free space falls below the configured threshold, up to your specified maximum. Amazon RDS detects low storage conditions and modifies the storage volume dynamically, requiring no manual intervention or downtime. This is the most direct and proactive solution to prevent the instance from reaching a storage-full state.

Why this answer

Enabling storage auto scaling on an Amazon RDS for PostgreSQL instance allows the database to automatically increase its allocated storage when it detects that available storage is running low, preventing out-of-storage errors without requiring manual intervention or downtime. Option E is correct because modifying the DB instance to increase the allocated storage size is a dynamic operation that can be performed without downtime, as Amazon RDS supports online storage scaling for PostgreSQL instances, allowing the change to take effect while the database remains available.

Exam trap

The trap here is that candidates often confuse instance class scaling (compute/memory) with storage scaling, or mistakenly think that deleting snapshots (which are stored separately in S3) can free up space on the DB instance's attached storage volume.

22
MCQeasy

A company uses AWS CodePipeline to automate the deployment of a static website hosted on Amazon S3. The pipeline includes a source stage that pulls from a CodeCommit repository and a deploy stage that uses CodeBuild to sync the files to an S3 bucket. The team noticed that the website is not updating after a successful pipeline run. The CodeBuild logs show that the 'aws s3 sync' command completed successfully. However, the website still shows the old content. What is the MOST likely cause?

A.The CodeBuild project does not have permission to write to the S3 bucket.
B.The S3 bucket is not configured for static website hosting.
C.The website is fronted by Amazon CloudFront, which is caching the old content.
D.The S3 bucket policy is blocking public access to the updated objects.
AnswerC

When CloudFront fronts an S3 static website, edge locations serve cached objects until they expire or are invalidated. CodePipeline's S3 sync updates the origin, but CloudFront does not automatically know about those changes; it continues returning the old objects as long as they remain in the cache. To serve the new content, you must create a CloudFront invalidation for the changed paths (e.g., /*) or use versioned filenames to bypass the cache. This directly matches the symptom of a successful deployment that still shows stale content in the browser.

Why this answer

The most likely cause is that Amazon CloudFront is caching the old content at edge locations. Even though CodeBuild successfully synced new files to S3, CloudFront continues to serve cached objects until the TTL expires or an invalidation is created. The pipeline must include a CloudFront invalidation step (or use versioned object keys) to force edge locations to fetch the updated content.

Exam trap

DOP-C02 often tests the trap that a successful S3 sync means the website is updated, ignoring CloudFront caching, so candidates blame IAM or bucket policies instead of the CDN layer.

How to eliminate wrong answers

Option A is wrong because the CodeBuild logs show 'aws s3 sync' completed successfully, which means the IAM permissions allowed the write; a permission failure would have produced an AccessDenied error. Option B is wrong because if static website hosting were not configured, the website would not load at all, not show old content. Option D is wrong because a bucket policy blocking public access would cause 403 errors for new objects, not stale content; also, the old content is still being served, indicating the objects are reachable.

23
MCQhard

A DevOps engineer is responsible for a critical application that uses an Amazon Aurora MySQL cluster. The application experiences sudden spikes in read traffic. The engineer needs to ensure that read replicas automatically scale to handle the load and that the application can tolerate the failure of an Availability Zone without manual intervention. Which solution should the engineer implement?

A.Use Amazon RDS for MySQL with Multi-AZ and create read replicas in multiple Availability Zones, then configure an Auto Scaling group for the replicas.
B.Create an Aurora Auto Scaling policy that adds Aurora Replicas based on CPU utilization, and place the replicas in multiple Availability Zones.
C.Enable Aurora Global Database and promote a secondary Region during read spikes to distribute the load.
D.Configure an Application Load Balancer to distribute read traffic across multiple Aurora Replicas, and use AWS Lambda to add replicas when CPU exceeds a threshold.
AnswerB

Aurora Auto Scaling automatically adjusts the number of Aurora Replicas based on metrics like CPU utilization. Placing replicas in multiple Availability Zones ensures that if one AZ fails, replicas in other AZs continue to serve read traffic. This solution provides automatic scaling and AZ resilience with minimal manual intervention, meeting both requirements.

Why this answer

Aurora Auto Scaling automatically adds or removes Aurora Replicas based on performance metrics, such as CPU utilization, to handle read traffic spikes. Deploying replicas across multiple Availability Zones ensures that the cluster remains available if an AZ fails. This managed feature provides automatic scaling and high availability with minimal operational effort, directly satisfying the requirements.

Exam trap

The trap here is confusing Aurora Replicas with RDS read replicas or assuming that manual scaling via Lambda is needed, when Aurora Auto Scaling handles it natively.

24
MCQeasy

A DevOps engineer receives an alert that an Amazon ECS service is failing to start tasks. The service uses the Fargate launch type. The task definition includes a container that requires port 8080. The security group associated with the service allows inbound traffic on port 8080. What should the engineer check NEXT?

A.Verify that the VPC subnets have a route to a NAT Gateway or Internet Gateway.
B.Confirm that the task definition's container image exists in ECR.
C.Check if the task definition has sufficient CPU and memory allocated.
D.Review the security group rules for outbound traffic.
AnswerA

Fargate tasks must download their container images from Amazon ECR before they enter the RUNNING state. If the service is configured to use private subnets that lack a route to a NAT Gateway (or a subnet with an Internet Gateway for public IPs), the task cannot establish outbound connectivity to the image registry, causing the deployment to hang in PROVISIONING and eventually fail. Verifying the subnet route tables for a 0.0.0.0/0 route to a NAT Gateway or IGW directly addresses the most likely cause of image pull failures.

Why this answer

Fargate tasks require network connectivity to pull container images from ECR (or Docker Hub) and to send logs to CloudWatch. Without a route to a NAT Gateway (for private subnets) or an Internet Gateway (for public subnets), the task cannot pull the image and fails to start. Option B is incorrect: while the image must exist, the immediate symptom of tasks failing to start when the image is missing would be an 'image not found' error, not a generic failure; the security group already allows inbound traffic on port 8080, but outbound connectivity is the issue.

Option C is incorrect: insufficient CPU/memory would cause tasks to enter a 'CPU exhausted' or 'memory exhausted' state, not prevent them from starting entirely. Option D is incorrect: the security group allows inbound traffic, but the issue is about egress connectivity for the task to reach the image registry.

25
MCQeasy

A DevOps team is using AWS CloudFormation to manage infrastructure. They want to reuse the same template across multiple environments (dev, test, prod) with minor parameter variations. Which CloudFormation feature should they use to pass environment-specific values without modifying the template?

A.Conditions
B.Outputs
C.Mappings
D.Parameters
AnswerD

Parameters are the designated CloudFormation feature for passing environment-specific values, such as instance types, AMI IDs, or environment names, into a template during stack creation or update. They are declared in the template and can be referenced via Ref or Fn::Sub, allowing users to supply input without editing the template, with optional constraints like AllowedValues, type validation, and NoEcho for sensitive data. This makes Parameters exactly what the DevOps team needs for dynamic per-environment configuration.

Why this answer

CloudFormation Parameters let you pass environment-specific values (like instance sizes, CIDR ranges, or environment names) into a template at stack creation or update time without modifying the template itself. This is exactly the use case described: one template, multiple environments, different values supplied at deploy time via the console, CLI, or a parameters file.

Exam trap

DOP-C02 often tests the confusion between Parameters (inputs at deploy time) and Mappings (static in-template lookups) — candidates pick Mappings thinking they allow per-environment variation, but Mappings require editing the template, which the question explicitly forbids.

How to eliminate wrong answers

Option A is wrong because Conditions control whether resources are created based on a boolean expression — they do not pass environment-specific values into the template. Option B is wrong because Outputs export values from a stack (like resource ARNs) for cross-stack references; they do not feed values in. Option C is wrong because Mappings are static lookup tables hardcoded in the template (e.g., region-to-AMI-ID), which would require editing the template to change per-environment values — the opposite of what is needed.

26
MCQhard

A DevOps team is debugging a production incident where an Application Load Balancer (ALB) is returning 503 errors for some requests. The target group instances are healthy. What is the most likely cause?

A.The security group for the ALB does not allow inbound traffic on port 443
B.Health checks are misconfigured to use an incorrect path
C.The deregistration delay setting on the target group is too long
D.Cross-zone load balancing is disabled
AnswerC

An excessively long deregistration delay prolongs the draining period in which a target is excluded from new request routing but still waits for in-flight requests to complete. For example, if the delay is set to 3,600 seconds (the maximum) and a rolling deployment drains instances, replacement targets may not become healthy quickly enough, leaving zero targets in rotation. With no healthy targets, the ALB returns HTTP 503 to new requests even though the underlying instances might function correctly—the bottleneck is the delayed deregistration.

Why this answer

The deregistration delay setting controls how long the ALB continues to send requests to an instance that is being deregistered. If this delay is too long, the ALB may route traffic to an instance that has already stopped accepting connections, resulting in 503 errors even though the health checks pass. Option A is incorrect because a missing security group rule would prevent any traffic from reaching the ALB, causing connection timeouts rather than 503 errors.

Option B is incorrect because the instance health checks are passing (as stated), so the health check path must be correct. Option D is incorrect because disabling cross-zone load balancing affects traffic distribution but does not cause 503 errors.

27
MCQmedium

A team is using AWS CodeDeploy to deploy an application to EC2 instances. They want to ensure that if a deployment fails, the instances are automatically rolled back to the previous version. What should they configure?

A.Add a BeforeBlockTraffic hook to the AppSpec file to run a rollback script.
B.Set the deployment configuration to 'CodeDeployDefault.OneAtATime' for a gradual deployment.
C.Configure the deployment group to use the same Auto Scaling group for rollback.
D.Enable automatic rollback in the deployment group settings.
AnswerD

Enabling automatic rollback in the deployment group settings is the correct native mechanism for requiring CodeDeploy to redeploy the last successful revision when a deployment fails. This feature can also be configured to roll back when a deployment hits a limit or when a CloudWatch alarm triggers, making it flexible for various failure scenarios. It is precisely what the team needs to achieve automatic recovery without manual intervention.

Why this answer

CodeDeploy supports automatic rollback configuration. When enabled, if a deployment fails, CodeDeploy automatically redeploys the last successful revision. Option A is wrong because a hook is a lifecycle event, not a rollback mechanism.

Option B is wrong because a deployment group contains instances but does not have a rollback setting. Option C is wrong because a deployment configuration specifies traffic routing and failure thresholds, not automatic rollback.

28
MCQeasy

A company uses AWS CloudFormation to manage a stack that includes an Amazon SQS queue. The queue name must be unique. The developer wants to define the queue name in the CloudFormation template. Which intrinsic function should be used to generate a unique name?

A.Fn::Sub
B.Fn::Select
C.AWS::NoValue
D.Fn::GetAtt
AnswerA

Fn::Sub is the correct choice because it performs variable and pseudo-parameter substitution on a string literal, allowing you to embed ${AWS::StackName} and other pseudo parameters like ${AWS::AccountId} directly into a bucket name. This creates a predictable, globally unique S3 bucket name that is automatically tied to the stack, eliminating the need to hardcode a name and reducing the risk of collisions when the stack is deployed in multiple environments.

Why this answer

A is correct because `Fn::Sub` can embed a pseudo parameter like `AWS::StackName` or `AWS::AccountId` into a string to generate a unique queue name. By using `Fn::Sub` with a reference to the stack name or a random string, you can create a name that avoids collisions across accounts or regions, satisfying the uniqueness requirement for SQS queue names.

Exam trap

The trap here is that candidates confuse `Fn::GetAtt` (which retrieves attributes like ARN or URL) with the ability to generate a unique name, but `Fn::GetAtt` cannot create or modify a name string—it only reads existing resource properties.

How to eliminate wrong answers

Option B is wrong because `Fn::Select` returns a single element from a list based on an index, which does not generate or modify a string to ensure uniqueness. Option C is wrong because `AWS::NoValue` is used to conditionally omit a property from a template, not to produce a unique name. Option D is wrong because `Fn::GetAtt` retrieves an attribute value from a resource (e.g., the ARN of an SQS queue), but it cannot generate a unique name; it only returns existing attributes of already-created resources.

29
MCQmedium

A company experiences intermittent high latency for a web application running on EC2 behind an ALB. They want to monitor and automatically replace instances that have high CPU. Which solution meets this requirement?

A.Create a CloudWatch alarm on CPU utilization that triggers an Auto Scaling policy to replace the instance
B.Use Auto Scaling scheduled scaling actions to replace instances at peak times
C.Use AWS Lambda to periodically check CPU and terminate high-CPU instances
D.Configure the ALB health check to mark instances unhealthy when CPU is high
AnswerA

A CloudWatch alarm on CPUUtilization monitors the metric in near real-time and, when breached, triggers an Auto Scaling policy that terminates the affected instance, after which the Auto Scaling group automatically launches a fresh replacement. This directly addresses the intermittent high latency by reacting to the actual symptom (high CPU) as it occurs, rather than relying on predetermined schedules, and the managed replacement ensures capacity is maintained without manual intervention.

Why this answer

You can configure a CloudWatch alarm on the EC2 instance's CPU utilization metric, and then use that alarm to trigger an Auto Scaling lifecycle hook or a scaling policy that terminates the unhealthy instance and launches a replacement. This directly ties performance monitoring to automated instance replacement, meeting the requirement to replace instances with high CPU.

Exam trap

The trap here is that candidates often confuse ALB health checks with instance health monitoring, assuming ALB can react to CPU metrics, when in fact ALB health checks only verify application-level responsiveness (e.g., HTTP status codes) and cannot directly measure CPU utilization.

How to eliminate wrong answers

Option B is wrong because scheduled scaling actions replace instances at fixed times, not in response to real-time high CPU utilization, so they cannot address intermittent latency. Option C is wrong because while Lambda could terminate instances, it adds unnecessary complexity and latency, and Auto Scaling already provides native health-check-based replacement without custom code. Option D is wrong because ALB health checks are designed to detect application or network failures (e.g., HTTP 5xx, connection timeouts), not CPU utilization; they cannot be configured to mark instances unhealthy based on CPU metrics.

30
MCQhard

A company uses AWS CodeBuild to compile and test their Java application. The build takes about 20 minutes. They have enabled Amazon S3 cache to store the Maven repository to speed up subsequent builds. However, they notice that the build time has not improved significantly. The buildspec file includes the 'cache' section with 'paths' pointing to '/root/.m2'. The CodeBuild project has cache type set to 'S3' and a valid bucket. The build logs show that the cache is being downloaded and uploaded, but the Maven dependencies are still being downloaded from the internet each time. What is the most likely cause?

A.The cache is too large and takes as long to download as the build itself.
B.The S3 bucket is in a different region than the CodeBuild project.
C.The buildspec file does not include the 'cache' section correctly.
D.The Maven dependencies are not being stored in the local repository path specified in the cache.
AnswerD

If the Maven local repository is overridden (e.g., via a custom settings.xml, the -Dmaven.repo.local flag, or a different user's home directory), the path declared in CodeBuild's cache section will not contain the downloaded dependencies. The build will then re-download artifacts every time because the restore phase puts files into a location Maven never reads. The fix is to ensure the cache path exactly matches the effective local repository path used by the Maven build.

Why this answer

The cache is being downloaded and uploaded, but Maven still fetches dependencies from the internet, which means the dependencies are not actually landing in /root/.m2 during the build. This typically happens when the build runs as a non-root user (e.g., the CodeBuild default user), so Maven writes to a different local repository path such as /home/codebuild/.m2 or a path defined by settings.xml. The cache section is syntactically correct, so the fix is to align the cached path with the actual Maven local repository location.

Exam trap

The trap is assuming the cache section is misconfigured when the logs prove it is working — the real issue is a path mismatch between the cached directory and Maven's actual local repository.

How to eliminate wrong answers

Option A is wrong because the logs show the cache is being downloaded and uploaded successfully, and a 20-minute Java build's Maven cache is typically tens to low hundreds of MB — not large enough to negate the benefit. Option B is wrong because CodeBuild S3 cache works cross-region; a region mismatch would cause cache upload/download failures, not silent re-downloads of Maven dependencies. Option C is wrong because the question states the buildspec includes the cache section with paths pointing to /root/.m2, and the logs confirm cache activity — so the section is present and functional.

31
Multi-Selecthard

Which THREE factors should be considered when designing a deployment strategy using AWS CodeDeploy to minimize downtime during updates? (Choose three.)

Select 3 answers
A.Configure a load balancer to deregister instances before deployment.
B.Use a blue/green deployment to switch traffic instantly.
C.Use canary deployments to shift traffic gradually.
D.Deploy to all instances simultaneously to reduce total time.
E.Use the same instance type for all instances.
AnswersA, B, C

Deregistering instances from the load balancer before deployment engages connection draining, allowing in-flight requests to finish before the instance is stopped. Combined with health check grace periods, users never reach a terminating instance, preventing 5xx errors and dropped connections during rolling or batch replacements. This is the core of a zero-downtime rolling update.

Why this answer

Deregistering instances from a load balancer before deployment ensures that traffic is not routed to instances that are being updated, preventing downtime during the deployment process. AWS CodeDeploy integrates with Elastic Load Balancing to automatically deregister instances, wait for in-flight requests to complete, and then register them back after the deployment succeeds.

Exam trap

The trap here is that candidates may think deploying to all instances simultaneously (Option D) reduces total time and thus minimizes downtime, but they overlook that without traffic management, this approach can cause complete service interruption during the update window.

32
MCQmedium

A DevOps team is designing a disaster recovery solution for an Amazon RDS for MySQL database. The primary database is in us-east-1, and the recovery point objective (RPO) is 5 minutes, recovery time objective (RTO) is 1 hour. Which solution meets these requirements?

A.Enable Multi-AZ deployment for high availability.
B.Create a cross-Region read replica in the secondary Region.
C.Take manual snapshots and copy them to the secondary Region daily.
D.Configure automated backups with a retention period of 35 days.
AnswerB

A cross-Region read replica uses asynchronous replication to continuously copy changes from the source DB instance to a replica in the secondary Region, keeping data loss typically within 5 minutes. In a disaster, you can promote this replica to a standalone primary instance, which is a fast, reversible operation that meets the 1-hour RTO. This is the only option that both maintains an up-to-date copy in another Region and provides a ready-to-activate target for write traffic.

Why this answer

A cross-Region read replica in the secondary Region meets the RPO of 5 minutes because replication from the primary RDS instance to the read replica is asynchronous but typically completes within seconds to a few minutes, well under the 5-minute threshold. In a disaster, promoting the read replica to a standalone instance can be done manually or automated, and the RTO of 1 hour is achievable because promotion takes only a few minutes, leaving ample time for DNS and application failover. This solution provides a continuous replication stream without manual intervention, unlike snapshot-based approaches.

Exam trap

The trap here is that candidates confuse Multi-AZ (high availability within a Region) with cross-Region disaster recovery, assuming Multi-AZ protects against Regional failures, but it only protects against Availability Zone failures within the same Region.

How to eliminate wrong answers

Option A is wrong because Multi-AZ deployment provides high availability within a single Region (us-east-1) by synchronously replicating to a standby in a different Availability Zone, but it does not protect against a Regional disaster, so it cannot meet the cross-Region recovery requirement. Option C is wrong because taking manual snapshots daily and copying them to the secondary Region results in an RPO of up to 24 hours, far exceeding the required 5 minutes, and the copy operation adds additional latency. Option D is wrong because automated backups with a retention period of 35 days are stored within the same Region and cannot be used for cross-Region recovery; they also do not provide a mechanism to restore in a secondary Region within the required RPO/RTO.

33
MCQmedium

A company is running a production application on Amazon ECS with AWS Fargate. The application has unpredictable traffic patterns and occasionally experiences increased latency. The DevOps team needs to configure scaling based on a custom metric that tracks the number of active user sessions in real time. Which solution will allow the team to scale the ECS service based on this custom metric?

A.Use AWS Auto Scaling to scale the ECS service based on the custom metric.
B.Create a CloudWatch dashboard to visualize the metric and manually adjust the service count.
C.Publish the custom metric to Amazon CloudWatch, then create a target tracking scaling policy in Application Auto Scaling for the ECS service.
D.Use an AWS Lambda function to directly update the desired count of the ECS service based on the metric.
AnswerC

The recommended approach is to first publish the custom application metric to Amazon CloudWatch using the PutMetricData API, ensuring it has the appropriate ECS service dimensions. Then, in Application Auto Scaling, create a target tracking scaling policy that references that custom metric; the policy automatically calculates the required desired count to maintain the target value, and adjusts the ECS service accordingly. This provides native, automated scaling with built-in cooldowns and no additional compute, making it the most operationally efficient solution.

Why this answer

Application Auto Scaling is the service that scales ECS services, and it supports target tracking policies based on custom CloudWatch metrics. Publishing the custom session-count metric to CloudWatch and then creating a target tracking scaling policy for the ECS service is the correct, fully managed approach.

Exam trap

DOP-C02 often tests the confusion between 'AWS Auto Scaling' and 'Application Auto Scaling' — candidates who pick the generic AWS Auto Scaling option miss that ECS services are scaled specifically through Application Auto Scaling with target tracking policies.

How to eliminate wrong answers

Option A is wrong because 'AWS Auto Scaling' is a separate service for scaling groups of related resources (EC2, DynamoDB, etc.) and is not the mechanism used to scale ECS service desired count — that is Application Auto Scaling. Option B is wrong because a CloudWatch dashboard is a visualization tool and manual adjustment does not meet the requirement for automatic scaling based on a custom metric. Option D is wrong because a Lambda function directly updating the desired count is a custom, unmanaged workaround that bypasses Application Auto Scaling's built-in cooldowns, alarms, and policy evaluation — it is not the recommended solution and adds operational overhead.

34
MCQmedium

A company runs a stateless web application on Amazon ECS with Fargate. The application must be highly available across multiple Availability Zones. What is the BEST way to achieve this?

A.Create an ECS service with tasks in multiple AZs and place an ALB in front.
B.Use an Auto Scaling group of EC2 instances in a single AZ and run ECS tasks on them.
C.Deploy a CloudFront distribution with multiple origins in different AZs.
D.Deploy a single ECS service with tasks in one AZ and use an ALB.
AnswerA

Running ECS tasks in multiple Availability Zones ensures that an AZ failure does not cause total loss of capacity. The ALB performs health checks and routes traffic only to healthy targets, so if tasks in one AZ become unhealthy, traffic automatically shifts to tasks in the other AZ. Additionally, ECS service scheduler maintains desired task count across AZs, and with service auto scaling you can scale based on request count or CPU. This architecture provides both resiliency and elasticity.

Why this answer

Running an ECS service with tasks in multiple Availability Zones (AZs) and placing an Application Load Balancer (ALB) in front ensures that if one AZ becomes unavailable, the ALB can route traffic to healthy tasks in the remaining AZs. This architecture provides both high availability and fault tolerance for the stateless web application, as the ALB performs health checks and distributes requests across tasks in different AZs.

Exam trap

The trap here is that candidates may think CloudFront (Option C) provides high availability for dynamic web applications, but it is a CDN for static content caching and does not replace the need for an ALB with multi-AZ task placement for active traffic routing.

How to eliminate wrong answers

Option B is wrong because using an Auto Scaling group of EC2 instances in a single AZ creates a single point of failure; if that AZ goes down, all instances and tasks become unavailable, violating the high availability requirement. Option C is wrong because CloudFront is a content delivery network (CDN) that caches content at edge locations; it does not provide active-active load balancing across AZs for dynamic web application traffic, and its origins are not designed to replace an ALB for distributing requests to ECS tasks. Option D is wrong because deploying a single ECS service with tasks in one AZ means all tasks are in a single failure domain; even with an ALB, if that AZ fails, the application becomes unavailable.

35
MCQmedium

A company is using Amazon CloudWatch Synthetics to monitor the availability of a web application. The canary runs every 5 minutes from multiple locations. Recently, the canary has been failing intermittently with HTTP 503 errors, but the application team reports that the application is healthy. Which step should the DevOps engineer take to identify the cause of the false positives?

A.Increase the canary timeout setting to allow more time for the application to respond.
B.Add more canary locations to increase coverage.
C.Review the canary's CloudWatch Logs to check for network errors or timeouts.
D.Increase the canary run frequency to every 1 minute.
AnswerC

The CloudWatch Synthetics canary's logs contain detailed request/response records, error stack traces, and screenshots captured during the run, making them essential for pinpointing whether a failure is caused by network errors, timeouts, or the application itself. By reviewing these logs, you can check for DNS resolution failures, TCP handshake timeouts, TLS certificate errors, or 4xx/5xx responses that indicate the problem is not with the client-side script. This is the first diagnostic step because it separates infrastructure-level issues from application-level defects.

Why this answer

Reviewing the canary's CloudWatch Logs is the correct first step because Synthetics canaries write detailed logs and screenshots for each run, including network errors, timeouts, and HTTP response details. Since the application team reports the app is healthy, the 503 errors are likely caused by the canary's network path, DNS resolution, or a transient issue that the logs will reveal. This is the diagnostic step that identifies the root cause of the false positives.

Exam trap

DOP-C02 often tests the instinct to 'fix' a monitoring problem by changing the monitor's configuration (timeout, frequency, locations) rather than investigating the logs to find the actual cause of false positives.

How to eliminate wrong answers

Option A is wrong because increasing the timeout does not diagnose the cause of 503 errors; if the app is healthy and responding quickly, a longer timeout will not fix a false positive and may mask the real issue. Option B is wrong because adding more canary locations increases coverage but does not explain why existing runs fail intermittently — it may even add more false positives. Option D is wrong because increasing run frequency to every 1 minute increases noise and cost without diagnosing the root cause, and could exacerbate throttling or rate-limit issues that cause 503s.

36
MCQmedium

A company uses AWS CodePipeline with a GitHub source action. The pipeline is configured to trigger on changes to the main branch. After a recent commit, the pipeline did not trigger. The DevOps engineer verified that the webhook is configured correctly and the IAM role has the necessary permissions. What is the most likely cause?

A.The GitHub personal access token used for authentication has expired.
B.The pipeline is set to manual execution only.
C.The source action's branch filter is set to a different branch.
D.The GitHub webhook endpoint URL is incorrect.
AnswerA

The PAT is the credential CodePipeline uses to authenticate to GitHub for API calls, including the initial webhook registration and every source-revision lookup (e.g., GetCommit, ListCommits). When the token expires, CodePipeline receives 401/403 responses, so webhook events cannot be processed successfully and the Source stage fails or is skipped, which matches the observed post-push behavior.

Why this answer

The most likely cause is that the GitHub personal access token used for authentication has expired. CodePipeline uses this token to authenticate with GitHub and trigger the webhook; when the token expires, GitHub rejects the webhook call, preventing the pipeline from starting even though the webhook configuration and IAM permissions are correct.

Exam trap

The trap here is that candidates may focus on webhook configuration or IAM permissions, overlooking that the GitHub personal access token is a separate authentication mechanism that can expire independently of the webhook setup.

How to eliminate wrong answers

Option B is wrong because if the pipeline were set to manual execution only, it would never trigger automatically, but the question states the pipeline is configured to trigger on changes to the main branch, implying automatic execution is expected. Option C is wrong because the branch filter is explicitly set to the main branch, and the DevOps engineer verified the webhook is configured correctly; a different branch filter would prevent triggering only if the commit were to a different branch, but the commit was to main. Option D is wrong because the engineer verified the webhook is configured correctly, which includes the endpoint URL; an incorrect URL would be a configuration error that would have been caught during verification.

37
MCQeasy

A company runs a stateless web application on EC2 instances in an Auto Scaling group across three Availability Zones. The application uses an Application Load Balancer. The operations team needs to ensure that the application remains available if one AZ fails. Which solution is MOST resilient?

A.Configure the Auto Scaling group to launch instances in a single Availability Zone with a desired capacity of 6.
B.Configure the Auto Scaling group to launch instances in two Availability Zones with a desired capacity of 4.
C.Configure the Auto Scaling group to launch instances in three Availability Zones with a desired capacity of 3.
D.Configure the Auto Scaling group to launch instances in two Availability Zones with a desired capacity of 6, all in one AZ.
AnswerC

Configuring three Availability Zones with a desired capacity of three places one instance in each AZ, so if any single AZ fails, the remaining two instances continue serving traffic, preserving 66% of capacity. The Auto Scaling group will automatically detect the unhealthy instances and launch replacements in other healthy AZs, gradually restoring capacity to the desired level. Because the application is stateless, a load balancer can distribute traffic across the surviving instances, providing high availability with minimal disruption.

Why this answer

Distributing instances across three Availability Zones (AZs) with a desired capacity of 3 ensures that even if one AZ fails, the remaining two AZs still have at least 2 instances running, maintaining service capacity. The Application Load Balancer (ALB) automatically routes traffic away from the failed AZ, and the Auto Scaling group will replace lost instances in the healthy AZs, providing the highest resilience against a single-AZ failure.

Exam trap

The trap here is that candidates often think using two AZs is sufficient for high availability, but the question specifically asks for the 'MOST resilient' solution, and three AZs provide better fault isolation and recovery capacity than two, especially when the desired capacity is low.

How to eliminate wrong answers

Option A is wrong because launching all instances in a single AZ creates a single point of failure; if that AZ fails, all instances are lost and the application becomes unavailable. Option B is wrong because distributing instances across only two AZs with a desired capacity of 4 means that if one AZ fails, the remaining AZ may have only 2 instances (if evenly split), but the total capacity drops by 50%, and the Auto Scaling group cannot launch instances in the failed AZ, potentially leading to insufficient capacity. Option D is wrong because it configures instances in two AZs but places all 6 instances in one AZ, which is functionally identical to a single-AZ deployment and provides no resilience against an AZ failure.

38
MCQmedium

After deploying a new application version using AWS CodeDeploy, an EC2 instance fails the deployment. The deployment group is configured with an in-place deployment. The engineer sees the error 'ScriptMissing' in the CodeDeploy logs. What should the engineer check?

A.The deployment group's deployment configuration
B.The file path defined in the appspec.yml for the lifecycle hook
C.The security group attached to the instance
D.The AMI used for the EC2 instance
AnswerB

The appspec.yml file is the deployment specification that maps each lifecycle hook (e.g., BeforeInstall, AfterInstall, ApplicationStart) to a script path. The 'ScriptMissing' error occurs precisely when the CodeDeploy agent attempts to execute the script declared in a hook and finds no file at that path, either because the path is mistyped, uses an incorrect relative reference, or the script file was omitted from the deployment revision. Since the agent only knows script locations from appspec.yml, an incorrect path there is the direct and immediate cause of this error, making this option correct.

Why this answer

The 'ScriptMissing' error in CodeDeploy indicates that a lifecycle hook script referenced in appspec.yml cannot be found at the specified path on the instance. The engineer should verify the file path defined for that hook in appspec.yml.

Exam trap

DOP-C02 often tests whether candidates can distinguish configuration-level causes (deployment config, security groups, AMI) from the actual appspec-driven cause of a specific lifecycle error like ScriptMissing.

How to eliminate wrong answers

Option A is wrong because the deployment configuration controls how many instances are deployed to at a time (e.g., one-at-a-time, half-at-a-time), not script resolution. Option C is wrong because security groups affect network access, not the presence of a local script file. Option D is wrong because the AMI determines the base OS image; while a missing script could theoretically stem from a bad AMI, the direct cause of 'ScriptMissing' is the appspec path.

39
MCQmedium

A company is using AWS CodeDeploy to deploy a web application to an Auto Scaling group. The deployment fails with the error 'The overall deployment failed because too many individual instances failed deployment, too few healthy instances are available, or some instances in your deployment group are experiencing problems.' The deployment configuration uses a linear traffic shifting with a 10-minute interval. The application logs show that the new version of the application crashes on startup. What is the MOST effective way to handle this situation to ensure successful future deployments?

A.Increase the interval in the linear traffic shifting to 30 minutes to allow more time for instances to stabilize.
B.Configure the deployment to automatically roll back when a failure occurs and ignore the error.
C.Switch to a blue/green deployment strategy to minimize the impact on existing instances.
D.Add a script in the AppSpec file's 'Validate Service' lifecycle hook to check the application health and fail the deployment if the application does not start successfully.
AnswerD

Adding a script to the ValidateService lifecycle hook is the correct solution because this hook executes after ApplicationStart and is designed to verify application readiness. When the script detects that the application did not start successfully—for example, it curls a local endpoint or checks the listening port—it returns a nonzero exit code, causing CodeDeploy to fail the deployment immediately. This check ensures unhealthy instances are identified before any production traffic is shifted, preventing user-facing outages and allowing you to fix the root cause.

Why this answer

The 'Validate Service' lifecycle hook in the AppSpec file runs after the application is installed and started, allowing you to execute a custom script that verifies the application is healthy. If the script detects that the new version crashes on startup, it can return a non-zero exit code, which causes CodeDeploy to mark that instance as failed and trigger the deployment failure. This provides an early, automated validation that prevents the deployment from proceeding with a broken application, directly addressing the root cause of the crash.

Exam trap

The trap here is that candidates often confuse recovery mechanisms (like rollback or blue/green) with prevention mechanisms, failing to realize that the most effective solution is to catch the failure early using the ValidateService lifecycle hook, which directly validates application health before traffic is shifted.

How to eliminate wrong answers

Option A is wrong because increasing the linear traffic shifting interval to 30 minutes does not fix the underlying issue of the application crashing on startup; it only delays the inevitable failure and wastes time. Option B is wrong because configuring automatic rollback is a recovery mechanism, not a prevention strategy, and ignoring the error would mask the problem, leading to repeated failures without addressing the root cause. Option C is wrong because switching to a blue/green deployment strategy does not prevent the new application version from crashing; it only isolates the impact on existing instances, but the deployment would still fail if the new version is broken.

40
MCQeasy

A DevOps engineer needs to ensure that all API calls made to AWS are logged for compliance. The logs must be stored in S3 for at least 7 years. Which AWS service should they use?

A.VPC Flow Logs
B.AWS Config
C.Amazon CloudWatch Logs
D.AWS CloudTrail
AnswerD

CloudTrail records every API call across the account, including the identity, source IP and timestamp, and delivers these events to an S3 bucket. This satisfies the requirement to log all API calls, with S3 lifecycle policies retaining the logs for the mandated seven years.

Why this answer

AWS CloudTrail is the correct service because it records all API calls made to AWS, including the identity, source IP, and timestamp, and can deliver log files to an S3 bucket for long-term retention. The requirement to store logs for at least 7 years aligns with CloudTrail's ability to integrate with S3 lifecycle policies for archival or deletion after a specified period.

Exam trap

The trap here is that candidates often confuse CloudTrail with CloudWatch Logs or AWS Config, thinking that any logging service can capture API calls, but only CloudTrail is designed specifically for auditing AWS API activity.

How to eliminate wrong answers

Option A is wrong because VPC Flow Logs capture network traffic metadata (IP addresses, ports, protocols) for VPCs, not API calls to AWS services. Option B is wrong because AWS Config records resource configuration changes and evaluates compliance rules, but it does not log API calls. Option C is wrong because Amazon CloudWatch Logs is designed for real-time monitoring and log storage from applications and AWS services, but it is not the primary service for auditing AWS API calls; CloudTrail is the dedicated service for that purpose.

41
MCQhard

A company wants to enforce that S3 buckets are not publicly accessible. Which AWS service can continuously monitor and automatically remediate non-compliant buckets?

A.AWS Config
B.Amazon Macie
C.AWS Security Hub
D.AWS Trusted Advisor
AnswerA

AWS Config is the correct service because it offers managed rules specifically for detecting S3 public access, such as s3-bucket-public-read-prohibited and s3-bucket-public-write-prohibited. These rules can be evaluated continuously, and when a violation is found, AWS Config can automatically invoke remediation actions through SSM automation documents, such as removing the public policy. This provides a continuous, automated enforcement mechanism that goes beyond simple detection, making it the only listed service capable of enforcing the requirement.

Why this answer

AWS Config continuously records resource configuration changes and evaluates them against Config rules, including the managed rule `s3-bucket-public-read-prohibited` and `s3-bucket-public-write-prohibited`. When a bucket becomes non-compliant, Config can trigger an SSM Automation document (via a remediation action) to automatically block public access, making it the service that both monitors and remediates.

Exam trap

DOP-C02 often tests the difference between detection and remediation — candidates pick Security Hub or Macie because they 'monitor' security, but only Config has native auto-remediation via SSM Automation.

How to eliminate wrong answers

Option B is wrong because Amazon Macie is a data-security service that discovers and classifies sensitive data (PII, credentials) in S3 — it does not evaluate bucket ACLs or policies for public access and cannot remediate them. Option C is wrong because AWS Security Hub aggregates and normalizes findings from other services (including Config) but does not itself perform continuous configuration evaluation or automatic remediation. Option D is wrong because Trusted Advisor provides advisory checks and recommendations but has no continuous evaluation engine or auto-remediation capability — it is a point-in-time best-practice report.

42
MCQeasy

A company uses AWS Secrets Manager to store database credentials for a legacy application running on an on-premises server. The application retrieves the secret via the AWS SDK. Recently, the database password was rotated in Secrets Manager, but the application continued to use the old password and failed to connect. The application code is correct and uses the latest SDK. The IAM role attached to the server has the secretsmanager:GetSecretValue permission. What is the MOST likely cause?

A.The IAM role does not have permission to list secrets
B.The application is using the wrong secret ID
C.The secret rotation Lambda function is failing
D.The application is caching the secret and not refreshing it after rotation
AnswerD

The AWS SDK and AWS Secrets Manager client-side caching libraries cache secret values in memory with a default TTL (e.g., 1 hour in Java) to reduce API calls. After rotation, the cached entry still holds the old AWSCURRENT value, and the application will keep using it until the cache expires or an explicit cache refresh is triggered. To fix this, the application must implement rotation-aware behavior, such as forcing a cache reload on authentication failure, shortening the TTL, or using the cached secret's version ID and comparing it to the newly fetched AWSCURRENT version. Without refreshing, the application is pinned to the pre-rotation secret indefinitely within the cache window.

Why this answer

The AWS SDK does not automatically re-fetch a secret on every call unless the application explicitly calls GetSecretValue again. Many applications cache the secret in memory (or in a local config) at startup for performance, so after Secrets Manager rotates the credential, the app keeps using the stale value. Since the IAM permission and SDK are correct, the most likely cause is client-side caching without a refresh mechanism.

Exam trap

The trap here is assuming that because the SDK is 'latest' and IAM is correct, the problem must be server-side (rotation Lambda or secret ID) — DOP-C02 often tests the client-side caching behaviour of Secrets Manager consumers.

How to eliminate wrong answers

Option A is wrong because secretsmanager:ListSecrets is only needed to enumerate secret names, not to retrieve a specific secret value — GetSecretValue is sufficient for the app's retrieval. Option B is wrong because an incorrect secret ID would cause a ResourceNotFoundException immediately, not a failure that appears only after a rotation. Option C is wrong because a failing rotation Lambda would leave the old password valid in the database, so the app would still connect successfully — the symptom described is the opposite.

43
Multi-Selectmedium

A company is deploying a serverless application using AWS Lambda, Amazon API Gateway, and Amazon DynamoDB. The application must be resilient to regional outages. Which THREE steps should the company take to achieve multi-Region resilience? (Choose THREE.)

Select 3 answers
A.Use Amazon CloudFront with multiple origins pointing to each Region's API Gateway.
B.Configure Route 53 with a failover routing policy to direct traffic to the secondary Region if the primary fails.
C.Use DynamoDB global tables to replicate data across Regions.
D.Deploy Lambda@Edge functions to handle requests at edge locations.
E.Deploy a second API Gateway and Lambda function in another Region.
AnswersB, C, E

Route 53 failover routing enables traffic redirection.

Why this answer

Amazon Route 53 with a failover routing policy allows the company to route traffic to a secondary Region when health checks detect a failure in the primary Region. This provides DNS-level failover, which is a fundamental component of multi-Region resilience for HTTP-based applications.

Exam trap

The trap here is that candidates often confuse CloudFront's origin failover capability (which requires manual configuration of origin groups) with automatic multi-Region failover, or they mistakenly believe Lambda@Edge can serve as a full application backend across Regions, when in fact it is limited to edge processing and cannot replace regional Lambda deployments.

44
MCQmedium

An EC2 instance shows as 'running' in the AWS console, but the system status check is 'impaired'. What is the most likely cause?

A.The instance's security group rules are blocking traffic.
B.The EBS root volume is corrupted.
C.The instance's operating system is not responding.
D.The underlying physical host has experienced a failure.
AnswerD

System status checks specifically monitor the health of the EC2 physical host and the surrounding AWS infrastructure, including power, network connectivity, and host maintenance events. When the underlying physical host experiences a hardware failure, loss of power, or network degradation, the system status check will fail because the host can no longer provide the required virtualization and I/O capabilities. This type of issue is considered host-level and is not fixable by rebooting the instance or making guest OS changes; instead, AWS recommends stopping and starting the instance to move it to a new, healthy host. The symptom described—an instance that is running in the console but failing system status checks—is the classic indicator of a physical host impairment.

Why this answer

The system status check specifically monitors the underlying physical host and hypervisor-level networking. When it is impaired, it indicates a failure of the AWS infrastructure supporting the instance, such as a power loss, network connectivity issue, or hardware degradation on the physical server. Since the instance status check (which monitors the guest OS) is not mentioned, the most likely cause is a host-level failure.

Exam trap

DOP-C02 often tests the distinction between system status checks and instance status checks, and candidates frequently confuse which check corresponds to host-level versus OS-level failures.

How to eliminate wrong answers

Option A is wrong because security group rules only affect network traffic to and from the instance; they do not cause status check failures. Option B is wrong because a corrupted EBS root volume would typically cause the instance status check to fail (or the instance to not boot), not the system status check. Option C is wrong because an unresponsive operating system would cause the instance status check to fail, not the system status check.

45
Multi-Selectmedium

Which TWO actions can be taken to protect an S3 bucket from being publicly accessible? (Select TWO.)

Select 2 answers
A.Use an SCP to deny s3:PutBucketPolicy.
B.Enable default encryption on the bucket.
C.Enable S3 Block Public Access settings on the bucket.
D.Enable MFA Delete on the bucket.
E.Use CloudFront to serve the bucket content.
AnswersA, C

An SCP (Service Control Policy) is a preventive guardrail applied at the AWS Organizations level, and it can deny the s3:PutBucketPolicy action for all IAM principals within an account or OU. This ensures that even if a user has an IAM policy granting s3:PutBucketPolicy, they cannot attach a bucket policy that would grant public access, because the SCP takes precedence. It is an effective way to enforce a 'no public buckets' rule across the entire organization.

Why this answer

An SCP (Service Control Policy) can explicitly deny the s3:PutBucketPolicy action at the AWS Organizations level, which prevents any IAM principal in affected accounts from attaching a public bucket policy. This is a preventive guardrail that overrides any permissive IAM permissions, ensuring the bucket cannot be made publicly accessible via policy statements.

Exam trap

The trap here is that candidates often confuse security features like encryption or MFA Delete with access control mechanisms, failing to recognize that only explicit policy restrictions (SCP or Block Public Access) can prevent public accessibility.

46
MCQhard

Refer to the exhibit. A developer is troubleshooting a failed AWS CodeBuild build. The buildspec file contains the following build commands: 'pre_build' - run linting, 'build' - './gradlew build', 'post_build' - package artifact. The error occurs in the build phase. Which of the following is the MOST likely cause?

A.The Gradle build encountered compilation or test errors.
B.The build environment ran out of disk space.
C.The artifact packaging step failed.
D.The linting step failed.
AnswerA

The failing command `./gradlew build` invokes Gradle's primary lifecycle task, which compiles all production and test sources and then executes the test suite. Any compilation error or failed test causes Gradle to abort with a non-zero exit code (typically 1), producing a 'BUILD FAILED' status. Since this failure occurs in the `build` phase and the exit code originates from the Gradle process itself, this is the direct and correct explanation.

Why this answer

The error occurs in the build phase, which executes './gradlew build'. Gradle's build task compiles source code and runs tests by default. If compilation fails or tests fail, Gradle exits with a non-zero exit code, causing AWS CodeBuild to mark the build phase as failed.

The error message in the build logs would typically show compilation errors or test failures.

Exam trap

The trap here is that candidates may confuse the phase in which each command runs (pre_build, build, post_build) and attribute the error to a step that executes in a different phase, rather than recognizing that the error occurs specifically in the build phase where './gradlew build' runs.

How to eliminate wrong answers

Option B is wrong because disk space exhaustion would typically cause a different error (e.g., 'No space left on device') and would likely affect all phases, not just the build phase. Option C is wrong because the artifact packaging step runs in the post_build phase, which occurs after the build phase; if the build phase fails, post_build never executes. Option D is wrong because linting runs in the pre_build phase; if linting failed, the error would occur in pre_build, not build.

47
MCQeasy

A DevOps engineer receives a CloudWatch alarm for high CPU utilization on an EC2 instance. The engineer needs to investigate the cause. Which AWS service can provide a detailed analysis of the running processes and their resource consumption?

A.AWS Config
B.AWS CloudTrail
C.AWS Systems Manager Run Command
D.Amazon Inspector
AnswerC

AWS Systems Manager Run Command lets you securely execute commands on managed EC2 instances via the SSM Agent, without requiring SSH, RDP, or open inbound ports. To diagnose a high-CPU alarm, you can run an OS-level command such as `top -b -n1` or `ps aux --sort=-%cpu` to get a snapshot of per-process CPU and memory usage, and retrieve the output from the console, CLI, or an S3 bucket. This gives you process-level detail directly from the guest OS, making it the correct tool to determine the root cause of the CPU spike.

Why this answer

AWS Systems Manager Run Command lets you remotely execute commands on EC2 instances (via the SSM agent) without SSH or RDP, so you can run tools like top, ps, htop, or perf to inspect running processes and CPU/memory consumption in real time. This is the correct service for interactive, instance-level process investigation triggered by a CloudWatch CPU alarm.

Exam trap

DOP-C02 often tests whether candidates reach for monitoring/audit services (CloudWatch, CloudTrail, Config, Inspector) when the question actually asks for remote command execution — Run Command is the only option that gives OS-level process visibility.

How to eliminate wrong answers

Option A is wrong because AWS Config records resource configuration changes and evaluates compliance rules — it does not provide runtime process or CPU analysis on instances. Option B is wrong because CloudTrail logs API calls and account activity (who did what in AWS), not OS-level process metrics. Option D is wrong because Amazon Inspector is a vulnerability management service that scans instances and containers for CVEs and network exposure, not a live process/resource profiler.

48
MCQhard

Refer to the exhibit. The S3 bucket policy is applied to a bucket. An application attempts to upload an object to the bucket using HTTP (not HTTPS). What will happen?

A.The upload fails because the condition matches HTTP requests
B.The upload succeeds if the bucket also has an allow policy for the user
C.The upload succeeds because there is no explicit allow statement
D.The upload fails because the bucket policy does not allow any access
AnswerA

The upload fails because the S3 bucket policy contains an explicit Deny statement whose condition key `aws:SecureTransport` evaluates to `false` for any HTTP request. Since the request was made over HTTP rather than HTTPS, the condition is satisfied and S3 returns 403 Forbidden, regardless of any IAM permissions that would otherwise permit the PutObject action. An explicit deny always takes precedence over any allow.

Why this answer

The bucket policy includes an explicit deny for all S3 actions when aws:SecureTransport is false (i.e., HTTP). This deny overrides any allow policies, so the upload fails. Option A is correct because the condition matches HTTP requests.

Option B is incorrect because the explicit deny overrides any allow. Option C is incorrect because the deny is explicit, so the condition is evaluated. Option D is incorrect because the policy only denies HTTP, not HTTPS.

49
MCQhard

A company runs a stateless web application on AWS Lambda behind an Application Load Balancer (ALB). During a deployment, the team updates the Lambda function to a new version. Some users report seeing the old version of the application for several minutes after the deployment. What is the MOST likely cause?

A.The Lambda function versions are not immutable, causing a gradual rollout.
B.Lambda@Edge is overriding the function version at the edge locations.
C.Amazon CloudFront is caching the old response and has not been invalidated.
D.The ALB target group is still pointing to the old Lambda function version due to connection draining.
AnswerD

ALB invokes Lambda functions via a target group that references a specific function version or alias. When you publish a new version and update the target group, ALB's connection draining process allows existing in-flight connections to complete on the old version before deregistering it. As a result, the old Lambda version can continue serving requests for a short period, causing a temporary gradual rollout until draining finishes.

Why this answer

When an ALB is used with Lambda, the ALB invokes a specific Lambda function version or alias. If the deployment updates the Lambda function but the ALB target group alias is not updated atomically, or if connection draining keeps old connections active, some requests may still be routed to the old version. This can cause users to see the old application for several minutes.

Option A is wrong because Lambda versions are immutable, so gradual rollout is not related. Option B is wrong because Lambda@Edge is not used in this setup (the application runs behind an ALB, not CloudFront). Option C is wrong because CloudFront is not mentioned in the architecture—the traffic goes directly from ALB to Lambda.

50
MCQhard

A company uses AWS CodeBuild to run builds for a Java application. The buildspec includes a 'mvn test' command. The build succeeds but the tests fail. The team wants to fail the build if any test fails. What should they do?

A.Add a 'post_build' phase that fails the build if tests fail.
B.Add a 'test' phase in the buildspec before the 'build' phase.
C.Ensure the buildspec's 'build' phase includes the test command and that the command returns a non-zero exit code on failure.
D.Configure the build project to use batch builds.
AnswerC

The correct approach is to run the test command (for example, mvn test or gradle test) inside the build phase of the buildspec. CodeBuild executes each command in a shell where a non-zero exit code causes that phase to be marked failed and the overall build to be FAILED. Because Maven/Gradle test goals already return non-zero when tests fail, this single change ensures the build pipeline stops immediately. Any subsequent build phases, such as post_build artifact packaging, still run, but the final build status will reflect the failure.

Why this answer

CodeBuild determines build success or failure based on the exit code of commands in the buildspec. The 'mvn test' command returns a non-zero exit code when tests fail, which causes the build to fail if placed in the 'build' phase. By ensuring the test command is in the 'build' phase and not suppressed (e.g., with '|| true'), the build will fail on test failure.

Exam trap

The trap here is that candidates confuse the 'post_build' phase with a phase that can retroactively fail the build, or invent a 'test' phase that does not exist in CodeBuild's buildspec schema.

How to eliminate wrong answers

Option A is wrong because the 'post_build' phase runs after the 'build' phase and does not affect the build's final status; it is intended for cleanup or notifications, and commands there cannot retroactively fail a build that already succeeded. Option B is wrong because the 'test' phase does not exist in CodeBuild's buildspec schema; valid phases are 'install', 'pre_build', 'build', and 'post_build'. Option D is wrong because batch builds are used to run multiple builds concurrently or sequentially, not to change how individual build phases handle exit codes or test failures.

51
MCQhard

A company uses AWS OpsWorks for configuration management. The DevOps team wants to run a custom recipe on all instances in a layer during stack updates. Which OpsWorks lifecycle event should they hook the recipe into?

A.Deploy
B.Shutdown
C.Configure
D.Setup
AnswerD

The Setup lifecycle event is specifically designed to run on an instance after it has been booted and registered with the stack, but before it is available. This event is intended for initial configuration tasks that need to happen only once, such as installing packages, creating directories, or setting up system users. In AWS OpsWorks, the Setup event is part of the standard lifecycle (Setup -> Configure -> Deploy -> Undeploy -> Shutdown) and is the correct lifecycle event for boot-time provisioning tasks, making it the answer to this question.

Why this answer

The Setup lifecycle event runs on every instance in a layer when it first boots or during a stack update, making it the correct hook for running a custom recipe that must execute on all instances during updates. This event occurs after the instance is fully configured and before it enters service, ensuring the recipe runs consistently across the layer.

Exam trap

The trap here is that candidates confuse the Configure event (which runs on all instances during any state change) with the Setup event (which is specifically designed for initial and update-time configuration), leading them to pick Configure because it seems more broadly applicable.

How to eliminate wrong answers

Option A (Deploy) is wrong because the Deploy event is triggered only when you manually run a 'deploy' command or use the Deploy action, not automatically during stack updates. Option B (Shutdown) is wrong because the Shutdown event runs when an instance is being terminated, not during stack updates. Option C (Configure) is wrong because the Configure event runs on every instance in the stack whenever an instance enters or leaves the online state, but it is not specifically tied to stack updates and may run multiple times, whereas Setup is the intended lifecycle event for initial configuration and update scenarios.

52
MCQmedium

A DevOps engineer manages a CI/CD pipeline that builds and deploys a containerized application to an Amazon ECS cluster. The pipeline runs on an EC2 instance and needs to retrieve a secret from AWS Secrets Manager to pass to the ECS task definition. The secret must not be stored on the instance or in the pipeline's code. The engineer wants to grant the pipeline the least privilege necessary to retrieve only that specific secret. Which approach should be taken?

A.Store the secret in an environment variable in the EC2 instance's user data and reference it in the pipeline.
B.Use AWS Systems Manager Parameter Store to store the secret as a SecureString parameter, and grant the instance's IAM role access to that parameter.
C.Create an IAM role for the EC2 instance with a policy that allows secretsmanager:GetSecretValue on the specific secret's ARN, and attach it to the instance profile.
D.Create an IAM user with programmatic access and a policy that allows secretsmanager:GetSecretValue on all secrets, then store the access keys in AWS CodePipeline as a secret parameter.
AnswerC

This approach uses an IAM role attached to the instance profile, providing temporary credentials to the pipeline. The policy scopes permissions to only the specific secret's ARN, adhering to least privilege. The secret is retrieved at runtime and never stored on the instance or in code.

Why this answer

The requirement is to retrieve a secret from AWS Secrets Manager without storing it on the instance or in code, and with least privilege. Attaching an IAM role to the EC2 instance with a policy scoped to the specific secret's ARN provides temporary credentials and restricts access. This is the most secure and operationally sound solution.

Exam trap

The trap here is assuming that storing secrets in user data or environment variables is acceptable for production pipelines, when it exposes credentials to anyone with instance access.

53
MCQmedium

Refer to the exhibit. A DevOps engineer ran the above AWS CLI command after a CloudFormation stack update. What does the status 'ROLLBACK_COMPLETE' indicate?

A.The stack update is in progress.
B.The stack was deleted successfully.
C.The stack was created successfully.
D.The stack update failed and CloudFormation reverted to the previous stack.
AnswerD

When an update operation fails, CloudFormation automatically initiates a rollback to the stack's previous template and resources, and ROLLBACK_COMPLETE is the final state after that restoration finishes. The stack is still present and functional from its prior state, but the attempted changes were discarded. This is the only status in the output that matches both the 'update' command and the terminal rollback outcome, so it directly confirms a failed update followed by a rollback.

Why this answer

The 'ROLLBACK_COMPLETE' status indicates that the CloudFormation stack update operation failed, and CloudFormation automatically reverted the stack to its previous stable state. This is a built-in safety mechanism: if any resource fails to update, CloudFormation triggers a rollback to undo all changes made during the update, ensuring the stack returns to its last known good configuration.

Exam trap

The trap here is that candidates confuse 'ROLLBACK_COMPLETE' with a successful operation or a deletion, when in fact it specifically means the update failed and the stack was reverted to its prior state.

How to eliminate wrong answers

Option A is wrong because 'ROLLBACK_COMPLETE' is a terminal state, not an in-progress state; an update in progress would show 'UPDATE_IN_PROGRESS' or 'UPDATE_ROLLBACK_IN_PROGRESS'. Option B is wrong because a successful deletion would show 'DELETE_COMPLETE', not 'ROLLBACK_COMPLETE'. Option C is wrong because a successful creation would show 'CREATE_COMPLETE', not 'ROLLBACK_COMPLETE'.

54
Matchingmedium

Match each AWS security and identity service to its function.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Manages users, groups, roles, and permissions

Creates and manages encryption keys

Rotates and manages secrets like database credentials

DDoS protection service

Web application firewall

Why these pairings

IAM handles access control, AWS Organizations manages multi-account structure, and AWS Shield protects against DDoS attacks. The distractors incorrectly swap these definitions.

55
MCQhard

Refer to the exhibit. A CodePipeline deployment fails at the CloudFormation stage. The Lambda function creation is cancelled. What is the MOST likely cause?

A.The buildspec.yml file contains an invalid command.
B.The Lambda function is configured in a VPC without a NAT gateway or VPC endpoints, causing deployment timeout.
C.The Lambda function's execution role lacks permissions to create ENIs.
D.The CodeCommit branch is not configured correctly in the pipeline.
AnswerC

If the Lambda execution role were missing ec2:CreateNetworkInterface or related ENI permissions, the Lambda service would fail to create an elastic network interface during function initialization, producing a different error such as 'InvalidParameterValueException' or a resource creation failure. This would prevent the function from running at all, rather than causing a timeout on the deployment's wait condition. Since the observed failure is a timeout, the role likely has the required ENI permissions, but the network route is missing.

Why this answer

When a Lambda function is configured in a VPC, the Lambda service must create an Elastic Network Interface (ENI) in the VPC on behalf of the function. The function's execution role must have permissions for ec2:CreateNetworkInterface, ec2:DescribeNetworkInterfaces, and ec2:DeleteNetworkInterface. Without these permissions, the Lambda service cannot create the ENI, and the Lambda function creation fails.

A missing NAT gateway or VPC endpoints only affects internet access for the running function, not the creation process. Therefore, Option C is the correct answer.

Exam trap

The trap is that candidates often attribute the failure to a lack of internet access (NAT gateway/VPC endpoints) or to the buildspec, but the actual cause is that the Lambda execution role is missing the required EC2 permissions to create network interfaces inside the VPC.

How to eliminate wrong answers

Option A is wrong because the buildspec.yml file is used in the CodeBuild stage, not the CloudFormation stage; an invalid command would cause a CodeBuild failure, not a CloudFormation deployment failure. Option C is wrong because the Lambda function's execution role lacking permissions to create ENIs would cause a different error, such as 'EC2 access denied' or 'Failed to create ENI', not a generic timeout or cancellation during CloudFormation deployment. Option D is wrong because a misconfigured CodeCommit branch would cause the pipeline to fail at the source stage, not at the CloudFormation stage.

56
Multi-Selectmedium

A DevOps engineer is designing a CI/CD pipeline for a Python application using AWS CodeBuild and AWS CodeDeploy. The application is deployed to an Auto Scaling group of EC2 instances. The engineer wants to ensure that the deployment does not impact availability. Which TWO strategies can be used? (Choose 2.)

Select 2 answers
A.In-place deployment with a large batch size.
B.Rolling deployment with a small batch size.
C.Immutable deployment.
D.Blue/green deployment.
E.Canary deployment.
AnswersB, D

Rolling deployment with a small batch size updates only a handful of instances at a time while the rest of the fleet continues serving live traffic, so the application remains available throughout the pipeline. AWS CodeDeploy moves the entire ASG to the new revision in incremental batches, and it uses health checks to verify each batch before proceeding; if a batch fails, the deployment is halted and can be rolled back. This is the correct choice when you need a zero-downtime update on an Auto Scaling group.

Why this answer

Rolling deployment with a small batch size (Option B) updates a limited number of instances at a time, ensuring that the majority of the Auto Scaling group remains available throughout the deployment. Blue/green deployment (Option D) creates a separate, fully-provisioned environment (green) and switches traffic to it only after validation, which eliminates downtime during the cutover. Both strategies directly preserve application availability by avoiding full-scale disruption.

Exam trap

The trap here is that candidates confuse 'canary' (a traffic-shifting pattern for Lambda/ECS) with 'rolling' (an instance-by-instance update for EC2), or incorrectly assume immutable deployment is available in CodeDeploy for Auto Scaling groups, when it is not a supported deployment type in that service.

57
Matchingmedium

Match each AWS compute or container service with its description.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Container orchestration service supporting Docker

Managed Kubernetes service

Serverless compute engine for containers

Serverless, event-driven compute service

Automatically adjusts EC2 capacity based on demand

Why these pairings

Correct matches: EC2 for virtual servers, ECS for Docker containers, EKS for Kubernetes, Lambda for serverless computing. Common confusions include mistaking EC2 as serverless or mixing up ECS and EKS.

58
Multi-Selecthard

A company uses AWS CodePipeline with multiple stages: Source (CodeCommit), Build (CodeBuild), Test (CodeBuild), and Deploy (CodeDeploy to EC2). The Test stage runs integration tests that require network access to a private database in a VPC. The CodeBuild project is configured to use a VPC. However, the Test stage fails intermittently with timeout errors. Which TWO actions would MOST likely resolve the issue? (Choose 2)

Select 2 answers
A.Remove the VPC configuration from the CodeBuild project and use a public subnet instead.
B.Increase the timeout for the Test stage in the CodeBuild project to accommodate network delays.
C.Ensure the CodeBuild project's VPC configuration includes a NAT gateway for internet access.
D.Use a larger compute type for the CodeBuild project to improve network performance.
E.Configure the security group for the CodeBuild project to allow outbound traffic to the database security group on the required port.
AnswersB, C

Increasing the Test stage timeout in the CodeBuild project is a pragmatic mitigation when network congestion intermittently causes dependencies or database connections to be slower than the default limit. CodeBuild has a default timeout of 60 minutes for a build, but the pipeline's stage timeout may be shorter, and network retries during peak usage can push a stage over its configured limit without an underlying configuration error. This accommodates transient network delays rather than masking a permanent misconfiguration, making it a valid, if not root-cause, fix.

Why this answer

The intermittent timeout errors suggest that the Test stage's integration tests are occasionally taking longer than the configured timeout. Increasing the timeout accommodates these delays, preventing premature failures. Option C is correct because a NAT gateway is required for CodeBuild in a VPC to access resources outside the VPC, such as the private database, if the database is in a different subnet or requires internet-routable traffic.

Exam trap

The trap here is that candidates may focus on security group rules (Option E) as the primary fix, but the intermittent nature of the timeout points to network path delays or missing NAT gateway, not a static misconfiguration.

59
MCQmedium

A company runs a microservices architecture on Amazon ECS with Fargate. Services communicate via an internal Application Load Balancer. Recently, one service became unavailable due to a memory leak, causing cascading failures in downstream services. What design change would MOST effectively improve resilience and limit the blast radius?

A.Increase the memory limit for each ECS task to accommodate memory leaks.
B.Implement circuit breaker patterns in the service discovery and client libraries to stop calling unhealthy services.
C.Enable connection draining on the ALB to allow in-flight requests to complete.
D.Implement automatic scaling policies for ECS services based on memory utilization.
AnswerB

Circuit breakers are a client-side resilience pattern that monitor outgoing requests to a dependency and, after exceeding a failure threshold, automatically fail fast without attempting the network call. In an ECS microservices environment with service discovery, this stops unhealthy services from being flooded with retries, preventing latency spikes and thread exhaustion in healthy services. By quarantining the failing service from call traffic, circuit breakers effectively stop cascading failures and allow the unhealthy service time to recover.

Why this answer

Implementing a circuit breaker pattern in service discovery and client libraries stops requests to unhealthy services, preventing cascading failures and limiting blast radius. Option A is wrong because increasing memory limits only delays the inevitable failure and does not prevent downstream services from being affected. Option C (connection draining) only affects in-flight requests during deregistration, not active health issues.

Option D (auto scaling) helps but does not stop requests from being sent to a failing service; scaling cannot fix a memory leak.

60
Multi-Selectmedium

Which TWO actions should a DevOps engineer take to implement a CI/CD pipeline that automatically deploys a containerized application to Amazon ECS using AWS CodePipeline and AWS CodeBuild? (Choose TWO.)

Select 2 answers
A.Use AWS CloudFormation to deploy the ECS service
B.Use AWS CodeDeploy to deploy the container to ECS with a blue/green deployment
C.Store the Docker image in AWS CodeCommit
D.Use CodeBuild to build the Docker image and push it to Amazon ECR
E.Configure CodePipeline to use Amazon ECR as a source for the container image
AnswersD, E

CodeBuild should be used to build the Docker image because it can execute a buildspec that runs 'docker build' and then uses 'aws ecr put-image' or 'docker push' to store the image in ECR. This step is the pipeline's build stage, converting source code into an immutable container artifact tagged with a version or commit hash. Without this action, the pipeline would have no image to deploy, making it the essential first of the two required actions.

Why this answer

AWS CodeBuild can be configured to build the Docker image from source code and then push it to Amazon ECR using the `post_build` phase with commands like `docker push`. This is a standard pattern for containerized CI/CD pipelines, as ECR serves as the private registry for storing and versioning container images. Option E is correct because CodePipeline can use Amazon ECR as a source action, which triggers the pipeline automatically when a new image is pushed to the specified repository, enabling continuous deployment to ECS.

Exam trap

The trap here is that candidates often confuse the role of CodeDeploy (which is for EC2/on-premises deployments) with ECS deployment mechanisms, or mistakenly think CodeCommit can store Docker images instead of source code, leading them to select options B or C.

61
MCQmedium

A DevOps engineer needs to securely store and automatically rotate database credentials for a web application running on Amazon ECS. Which solution should be used?

A.Use AWS KMS to generate and rotate a data key for encrypting the credentials in a file on ECS.
B.Store the credentials in AWS Systems Manager Parameter Store as a SecureString. Use a Lambda function to rotate them.
C.Store the credentials in AWS Secrets Manager and configure rotation. Grant the ECS task IAM role permission to retrieve the secret.
D.Use AWS Certificate Manager to store the credentials as a certificate.
AnswerC

AWS Secrets Manager natively supports automatic rotation of secrets with a configurable rotation schedule, using a Lambda function that updates the secret in both Secrets Manager and the target database. The ECS task assumes an IAM role whose permissions include secretsmanager:GetSecretValue, allowing the container to fetch the current password at runtime without embedding it in the task definition. This approach centralizes secret storage, enables rotation without redeploying the ECS service, and follows AWS best practices for managing database credentials.

Why this answer

AWS Secrets Manager can store database credentials and automatically rotate them on a schedule. The ECS task can retrieve the credentials using the Secrets Manager secret. Option C is correct.

Option A (AWS KMS) is for encryption keys, not credential rotation. Option B (SSM Parameter Store) can store secrets but does not support automatic rotation. Option D (AWS Certificate Manager) is for SSL/TLS certificates.

62
MCQhard

A company runs an application on EC2 with a shared Elastic IP. The instance fails and an engineer manually attaches the Elastic IP to a standby instance. To automate this failover, which service should be used?

A.Use an Auto Scaling group with a lifecycle hook
B.AWS Elastic Beanstalk
C.CloudWatch Events with a Lambda target
D.Configure a second Elastic IP
AnswerC

CloudWatch Events (now EventBridge) can capture Simple Notification Service messages or CloudWatch alarm state changes triggered by a Route 53 health check or EC2 status check, and then invoke an AWS Lambda function. That function can programmatically call ec2-associate-address to detach the Elastic IP from the failed instance and attach it to a healthy one, fully automating active/passive failover. This approach is the standard serverless pattern for EIP failover because it reacts to health signals and directly manages the association.

Why this answer

C is correct because CloudWatch Events (now Amazon EventBridge) can detect the EC2 instance state change (e.g., 'stopped' or 'failed') and trigger a Lambda function. The Lambda function can then programmatically disassociate the Elastic IP from the failed instance and reassociate it to a standby instance using the AWS SDK (e.g., ec2.disassociate_address and ec2.associate_address). This provides a fully automated, event-driven failover without manual intervention.

Exam trap

The trap here is that candidates often confuse event-driven automation (CloudWatch Events + Lambda) with scaling or deployment services (Auto Scaling, Elastic Beanstalk), or they assume adding a second Elastic IP solves failover without considering the need for automated reassignment.

How to eliminate wrong answers

Option A is wrong because an Auto Scaling group with a lifecycle hook is designed to manage instance launch/termination events (e.g., running custom scripts during scale-out/scale-in), not to reassociate an Elastic IP to a standby instance; it does not handle Elastic IP failover. Option B is wrong because AWS Elastic Beanstalk is a PaaS service for deploying and scaling web applications, not a mechanism for automating Elastic IP reassignment; it manages environments but does not provide direct control over Elastic IP failover. Option D is wrong because configuring a second Elastic IP does not automate failover; it simply adds another static IP, but the engineer would still need to manually update DNS or routing to switch traffic, which does not solve the automation requirement.

63
Multi-Selecthard

A company runs a critical application on AWS that uses an Auto Scaling group of EC2 instances. The application must remain available even if an entire Availability Zone fails. Which THREE actions should the company take?

Select 3 answers
A.Configure an ALB health check to automatically replace unhealthy instances.
B.Use a single instance in each Availability Zone to minimize cost.
C.Use multiple subnets in each Availability Zone for the instances.
D.Configure the Auto Scaling group to launch instances in at least two Availability Zones.
E.Use an Elastic Load Balancer (ELB) to distribute traffic across the instances in different AZs.
AnswersA, D, E

An ALB's health check marks an instance unhealthy when it repeatedly fails the configured protocol and path check, and the ALB then stops sending traffic to that instance. When the Auto Scaling group is configured to use ELB health checks, the 'ReplaceUnhealthy' process automatically terminates the unhealthy instance and launches a new one to maintain desired capacity. Thus, the health check is the trigger that enables the self-healing replacement of failed instances, ensuring application availability.

Why this answer

Configuring an ALB health check allows the Auto Scaling group to automatically detect and replace unhealthy instances. The ALB health check pings the instances at a specified interval (e.g., every 30 seconds) and marks them as unhealthy if they fail to respond. This triggers the Auto Scaling group to terminate the unhealthy instance and launch a new one, maintaining application availability even if an instance fails within a single AZ.

Exam trap

The trap here is that candidates might think using multiple subnets within a single AZ (Option C) provides redundancy, but it does not protect against an AZ failure, which requires instances to be spread across at least two distinct AZs.

64
Multi-Selecteasy

A company uses AWS CodeCommit for source control. Developers need to automatically run tests on every push to a feature branch, but only if the push includes changes to the 'src/' directory. Which TWO AWS services can be used together to achieve this?

Select 2 answers
A.Amazon CloudWatch Logs
B.AWS CodePipeline
C.AWS CodeDeploy
D.AWS Lambda
E.AWS CodeBuild
AnswersB, D

AWS CodePipeline is a fully managed continuous delivery service that can be configured with CodeCommit as a source action, automatically detecting new commits and triggering the pipeline execution. It supports path filters in the source action (e.g., requiring changes only in a specific folder) to avoid unnecessary runs. This makes it the native, event-driven orchestration service to start automated tests (via a CodeBuild build stage) exactly when developers push code.

Why this answer

AWS CodePipeline can be configured to trigger a pipeline execution when a push occurs to a specific branch in CodeCommit, and it supports path-based filtering using the 'PathFilter' condition in the source action. By combining CodePipeline with AWS Lambda, you can run custom validation logic (e.g., checking if changes are within the 'src/' directory) before invoking further actions like tests. This pair allows you to conditionally run tests only when the push includes changes to 'src/'.

Exam trap

The trap here is that candidates often assume AWS CodeBuild alone can handle event-driven triggers with path filtering, but it requires an orchestrator like CodePipeline or a Lambda-based custom trigger to implement conditional execution based on directory changes.

65
Multi-Selecthard

Which THREE are features of AWS Key Management Service (KMS) that help with compliance requirements? (Choose 3)

Select 3 answers
A.Automatic password generation for databases.
B.Automatic key rotation every year (optional).
C.Key policies to control access to keys.
D.Integration with AWS CloudTrail for auditing key usage.
E.Automatic deletion of keys after a specified period.
AnswersB, C, D

KMS offers optional automatic rotation of customer-managed keys on an annual basis. When enabled, KMS generates new backing key material each year while preserving the old material so that previously encrypted data can still be decrypted. This feature helps meet compliance requirements for cryptographic key lifecycle management.

Why this answer

AWS KMS supports optional automatic annual key rotation for customer managed keys. This helps meet compliance frameworks (e.g., PCI DSS, SOC, HIPAA) that require periodic cryptographic key rotation to limit the amount of data encrypted under a single key. When enabled, KMS automatically rotates the key material once per year, creating a new backing key while retaining the old one for decryption of previously encrypted data.

Exam trap

The trap here is that candidates confuse KMS key rotation with the automatic deletion or expiry of keys, or they mistakenly associate KMS with password generation features that belong to other AWS services like Secrets Manager.

66
MCQhard

A company uses AWS CodeDeploy to deploy a web application to an Auto Scaling group. The deployment fails because the new instances cannot connect to the database. The previous deployment succeeded. The DevOps engineer checks the CodeDeploy deployment configuration and finds that the deployment uses the 'CodeDeployDefault.AllAtOnce' configuration. What is the MOST likely cause of the failure?

A.The load balancer health check is misconfigured, causing new instances to be deregistered.
B.The deployment group is not configured to handle traffic routing, causing a routing loop.
C.The security group for the Auto Scaling group was updated during deployment, blocking database access.
D.The deployment replaced all instances at once, causing a temporary loss of connectivity if the new application version has incompatible changes.
AnswerD

With an AllAtOnce deployment strategy, CodeDeploy stops every current instance at the same time, deploys the new application revision to each, and then restarts them; this creates a zero-capacity window where no instance is serving traffic. If the new revision contains incompatible changes—such as expecting a new database schema, missing environment variables, or a different API version—the restarted instances may fail to connect to the database, leaving the entire application unavailable until the issues are resolved. This is the classic failure mode of AllAtOnce deployments: every instance fails together, so a single backward-incompatibility causes a full, fleet-wide outage rather than a partial degradation.

Why this answer

The 'CodeDeployDefault.AllAtOnce' deployment configuration causes CodeDeploy to attempt to deploy the new application revision to all instances in the Auto Scaling group simultaneously. If the new application version contains incompatible changes—such as a database schema mismatch, altered connection strings, or missing environment variables—all instances will fail at the same time, resulting in a complete loss of connectivity to the database. This contrasts with a rolling or canary deployment, which would limit the blast radius by updating only a subset of instances at a time.

Exam trap

The trap here is that candidates may assume the failure is due to a misconfiguration (like security groups or health checks) rather than recognizing that the deployment strategy itself—replacing all instances at once—amplifies the impact of any application-level incompatibility.

How to eliminate wrong answers

Option A is wrong because a misconfigured load balancer health check would cause instances to be deregistered from the target group, but the failure described is that new instances cannot connect to the database—a connectivity issue, not a registration issue. Option B is wrong because CodeDeploy deployment groups do not handle traffic routing in a way that creates routing loops; traffic routing is managed by the load balancer, and a routing loop would affect all traffic, not just database connections. Option C is wrong because security groups are not updated automatically during a CodeDeploy deployment; the engineer would have to manually modify them, and the question states the previous deployment succeeded, implying no security group changes occurred.

67
MCQmedium

A company runs an internal REST API on Amazon API Gateway backed by AWS Lambda. The operations team wants a single CloudWatch alarm that fires when the API's overall error rate exceeds 5% during any 5-minute period, without creating one alarm per resource. The API is deployed as a REST API (not HTTP API). Which approach should a DevOps engineer take?

A.Enable AWS X-Ray tracing on the stage and create a CloudWatch alarm on the X-Ray ErrorRate metric for the API.
B.Create a CloudWatch metric math alarm using the expression m1/m2 where m1 is the sum of the 4XXError and 5XXError metrics and m2 is the Count metric for the stage, all with period 300 and statistic Sum.
C.Create one CloudWatch alarm per method with a 5% threshold and combine them into a composite alarm using an OR rule.
D.Create a CloudWatch Logs metric filter on the API Gateway access logs and alarm on the resulting custom metric with a 5% threshold.
AnswerB

Metric math lets a single alarm combine several API Gateway metrics. Summing 4XXError and 5XXError and dividing by Count at a 300-second period yields the stage-level error rate, so one alarm covers the whole API rather than one alarm per method or resource. This directly satisfies the 5% threshold requirement.

Why this answer

API Gateway publishes 4XXError, 5XXError, and Count per stage. A metric math alarm can sum the error metrics and divide by Count at a five-minute period, producing a single stage-level error-rate alarm. Per-method alarms, X-Ray metrics, and log metric filters either fragment the view or do not directly yield a percentage.

Exam trap

The trap here is assuming API Gateway exposes a ready-made percentage error metric, when in fact you must compute the ratio yourself with metric math.

68
MCQmedium

A company uses AWS Lambda functions behind an Amazon API Gateway REST API. The DevOps team wants to monitor the end-to-end latency of API requests, including the time spent in API Gateway and Lambda. Which approach provides the most granular breakdown?

A.Enable Lambda Insights to get per-request latency breakdown.
B.Enable AWS X-Ray tracing on API Gateway and Lambda.
C.Enable VPC Flow Logs to capture network round-trip times.
D.Use CloudWatch metrics for API Gateway and Lambda, then add them together.
AnswerB

X-Ray propagates a trace ID through API Gateway and the downstream Lambda invocation, producing segment-level timings for each component. This granular breakdown isolates API Gateway overhead from Lambda init and execution duration, satisfying the requirement to measure end-to-end latency across both services.

Why this answer

AWS X-Ray provides end-to-end tracing with detailed segments for API Gateway and Lambda, allowing per-request breakdown of latency across each component. Option A (Lambda Insights) offers OS-level metrics like CPU and memory, not request-level latency breakdown. Option C (VPC Flow Logs) captures network traffic metadata but not application layer latency.

Option D (combining CloudWatch metrics) only gives aggregate statistics and cannot break down latency per request or within a single request's path.

69
MCQmedium

A company wants to centralize IAM user management across multiple AWS accounts. The company currently uses individual IAM users in each account. What is the BEST practice for centralized access control?

A.Use AWS Organizations and AWS IAM Identity Center (AWS SSO) to manage users centrally.
B.Create the same IAM users in each account with identical permissions.
C.Create IAM roles in each account and allow cross-account access from a central account.
D.Use IAM federation with an external identity provider and assign permissions based on SAML attributes.
AnswerA

AWS IAM Identity Center (formerly AWS SSO) integrates natively with AWS Organizations, providing a single place to manage users and groups and then assign them access across multiple AWS accounts. It uses permission sets to define IAM policies that are applied consistently to accounts, and it issues short-term AWS credentials, eliminating the need to create and rotate IAM users. This is the only option that genuinely centralizes user management while preserving fine-grained, auditable access control.

Why this answer

AWS IAM Identity Center (successor to AWS SSO) combined with AWS Organizations is the AWS-recommended best practice for centralizing workforce identity and access across multiple accounts. It provides a single place to manage users and groups, assign permission sets to accounts/OUs, and federate to external IdPs if desired — eliminating per-account IAM user sprawl. This is the canonical 'centralized access management' answer for multi-account environments.

Exam trap

The trap is confusing 'centralized access' with 'cross-account roles' — candidates pick option C because cross-account roles sound centralized, but the exam expects IAM Identity Center as the AWS-recommended best practice for multi-account user management.

How to eliminate wrong answers

Option B is wrong because duplicating IAM users across accounts creates credential sprawl, inconsistent permissions, and no single source of truth — it is explicitly an anti-pattern AWS advises against. Option C is wrong because while cross-account IAM roles are a valid pattern, they still require managing identities and trust relationships per account and do not provide centralized user lifecycle management the way Identity Center does; it is a partial solution, not the best practice. Option D is wrong because IAM federation with an external IdP is a component of a solution but by itself does not centralize account assignments across an AWS Organization — Identity Center is the AWS-native service that wraps federation plus centralized permission assignment.

70
MCQhard

A company uses AWS Config to track resource changes. They want to automatically remediate non-compliant security group rules that allow public SSH access. What is the MOST effective approach?

A.Set up an AWS Config rule that triggers a Lambda function to remove the SSH rule.
B.Use Amazon CloudWatch Events to detect the change and invoke a Lambda function.
C.Use AWS Service Catalog to enforce security group templates.
D.Create an AWS Config rule with an automatic remediation action using AWS Systems Manager Automation.
AnswerD

This is the correct approach because AWS Config rules continually evaluate resources against a desired policy, and when they detect non-compliance they can trigger an automatic remediation action—a Systems Manager Automation document—to fix the resource. In this case the rule (such as the managed RESTRICTED_SSH rule) would flag any security group with port 22 open to 0.0.0.0/0, and the associated SSM Automation document (for example, AWS-RevokeSecurityGroupIngress) would revoke the offending rule automatically. AWS Config tracks the remediation status and retries until the resource becomes compliant, providing a closed-loop, auditable remediation process without manual involvement.

Why this answer

AWS Config can directly associate an AWS Systems Manager Automation document as a remediation action for a non-compliant rule. This approach provides a fully managed, idempotent, and auditable remediation workflow without requiring custom Lambda code or external event orchestration. The automation document can be configured to automatically remove the SSH ingress rule (port 22) from the security group when the Config rule detects non-compliance.

Exam trap

The trap here is that candidates often assume a custom Lambda function (Option A) is the most flexible or effective approach, but AWS Config's native remediation with Systems Manager Automation is the recommended, fully managed, and less error-prone solution for automatic compliance enforcement.

How to eliminate wrong answers

Option A is wrong because while a Lambda function can remove the SSH rule, this approach requires you to write, deploy, and maintain custom code, and it does not natively integrate with AWS Config's remediation lifecycle (e.g., automatic retries, resource exclusion, or rollback). Option B is wrong because Amazon CloudWatch Events (now Amazon EventBridge) can detect security group changes, but it only provides an event notification; it does not include built-in remediation orchestration, compliance evaluation, or the ability to automatically trigger a remediation action directly from a Config rule evaluation. Option C is wrong because AWS Service Catalog is used to provision and govern pre-defined product templates, not to automatically remediate existing non-compliant resources; it cannot react to a Config compliance change or modify an already deployed security group.

71
Multi-Selectmedium

A company is designing a secure CI/CD pipeline. Which TWO actions should be taken to protect secrets (e.g., API keys) used in the pipeline? (Choose TWO.)

Select 2 answers
A.Encrypt secrets with AWS KMS and store the encrypted value in the source code
B.Store secrets in AWS Secrets Manager
C.Use IAM roles to grant the CI/CD service access to secrets
D.Store secrets in plaintext in the buildspec file
E.Pass secrets as environment variables in the build
AnswersB, C

AWS Secrets Manager is the correct choice because it natively stores secrets as encrypted objects with fine-grained IAM policies, automatic rotation for both AWS and custom secrets, and direct integration with services like CodeBuild, RDS, and Lambda. The pipeline retrieves a reference to the secret at build time, so the material value never appears in source control, buildspec files, or logs. Secrets Manager also logs every retrieval via CloudTrail, enabling security auditing and immediate revocation if a secret is compromised.

Why this answer

Option B is correct because AWS Secrets Manager is purpose-built to store, encrypt (using KMS), and rotate sensitive values such as API keys, so the pipeline retrieves them at runtime rather than embedding them in code or config. Option C is correct because granting the CI/CD service an IAM role with least-privilege permissions to read specific secrets enables secure, credential-free access via temporary STS credentials instead of long-lived static keys. Option A is wrong because even KMS-encrypted secrets committed to source code expose the ciphertext to anyone with repo access and risk decryption-key misuse.

Option D is wrong because plaintext secrets in a buildspec file are directly readable in the repository and build logs. Option E is wrong because environment variables can leak through logs, process listings, and child processes, and are not a secure secret-management mechanism on their own.

Exam trap

DOP-C02 often tests the misconception that encrypting secrets and committing them to source control is acceptable, when the correct pattern is to keep secrets entirely out of code and retrieve them at runtime via IAM-authorized services.

72
MCQeasy

A company uses AWS CodeBuild to build and test code. They need to securely store sensitive parameters, such as database passwords, and inject them into the build process. Which AWS service should they use?

A.Storing them in the CodeBuild project environment variables
B.AWS Secrets Manager
C.AWS Key Management Service (KMS)
D.AWS Systems Manager Parameter Store
AnswerD

AWS Systems Manager Parameter Store is the correct choice because it securely stores configuration data and secrets as parameters, supports typed values including SecureString using KMS encryption, and CodeBuild has native integration by allowing environment variables to reference parameter names and fetch values at build time. It is ideal for static parameters because it imposes no rotation overhead, and you only need to grant the build project's IAM role ssm:GetParameter(s) permission.

Why this answer

AWS Systems Manager Parameter Store is the correct choice because it provides secure, hierarchical storage for configuration data and secrets, such as database passwords, and integrates natively with AWS CodeBuild via the `parameter-store` environment variable type. This allows you to reference parameters without hardcoding sensitive values, and you can optionally use secure string parameters encrypted with KMS for additional protection.

Exam trap

The trap here is that candidates often confuse AWS Secrets Manager with Systems Manager Parameter Store, but the exam expects you to know that Parameter Store is the simpler, more cost-effective choice for injecting static or semi-static configuration values into CodeBuild, while Secrets Manager is intended for secrets requiring automatic rotation.

How to eliminate wrong answers

Option A is wrong because storing sensitive parameters directly in CodeBuild project environment variables exposes them in plaintext in the build configuration and logs, violating security best practices. Option B is wrong because while AWS Secrets Manager can store secrets, it is designed for automatic rotation of credentials and is more complex and costly than needed for simple parameter injection into CodeBuild; the question specifically asks for a service to 'store and inject' parameters, which Parameter Store handles with lower overhead. Option C is wrong because AWS Key Management Service (KMS) is a key management service for creating and controlling encryption keys, not a storage service for secrets or parameters; it can be used to encrypt parameters in Parameter Store or Secrets Manager, but it does not itself store or inject values into CodeBuild.

73
Multi-Selecteasy

A startup runs a stateless web application on AWS Elastic Beanstalk with a single environment. The application uses an Amazon RDS for MySQL database instance. The startup is preparing for a marketing campaign that is expected to increase traffic by 10x. The CTO is concerned about the application's ability to handle the load and wants to ensure high availability and resilience. The current architecture has a single RDS instance (db.t3.medium) and a single Elastic Beanstalk environment with one EC2 instance (t3.medium). The startup has a limited budget but wants to improve resilience without over-provisioning. Which combination of actions should the DevOps engineer recommend? (Choose THREE.)

Select 3 answers
A.Add an Amazon ElastiCache cluster to cache frequent database queries.
B.Use dedicated instances for the EC2 instances to ensure consistent performance.
C.Switch the Elastic Beanstalk environment to a load-balanced, auto-scaled environment with a minimum of 2 instances across 2 Availability Zones.
D.Enable Multi-AZ deployment for the RDS instance to provide a standby in another AZ.
E.Add Amazon RDS Proxy in front of the RDS instance to handle connection pooling.
AnswersC, D, E

Deploying the Elastic Beanstalk environment as a load-balanced, auto-scaling configuration with a minimum of two instances in separate Availability Zones eliminates the web tier as a single point of failure. The load balancer distributes traffic across instances and health-checks them, while Auto Scaling replaces unhealthy instances and can scale out during load spikes. This is the foundational action for high availability of stateless applications because it provides both redundancy and elasticity.

Why this answer

Option C is correct because converting the single-instance Elastic Beanstalk environment to a load-balanced, auto-scaled environment with a minimum of two instances across two Availability Zones removes the single point of failure at the web tier and lets the environment scale horizontally to absorb the 10x traffic spike. Option D is correct because enabling Multi-AZ on the RDS for MySQL instance creates a synchronous standby replica in a second AZ with automatic failover, improving database resilience without requiring application changes. Option E is correct because RDS Proxy pools and shares database connections, which prevents the connection exhaustion and overhead that occur when many new EC2 instances and users open connections directly to MySQL during a traffic surge.

Option A is not among the marked answers, and while caching could reduce read load, it is not required to achieve the stated high availability and resilience goals. Option B is not marked correct because dedicated instances raise cost and do not provide the multi-AZ redundancy or elasticity that the scenario demands.

Exam trap

DOP-C02 often tests the misconception that adding caching or dedicated instances is necessary for resilience, but the core actions are auto-scaling, Multi-AZ, and connection pooling; candidates may overlook RDS Proxy or choose cost-ineffective options.

74
MCQhard

A company runs a microservices architecture on Amazon ECS with Fargate launch type. Each microservice is deployed using AWS CodePipeline with a source stage from CodeCommit, a build stage in CodeBuild, and a deploy stage that updates the ECS service. The team wants to implement a blue/green deployment strategy to reduce downtime and enable quick rollbacks. Which combination of AWS services and configurations should be used?

A.Use AWS CloudFormation with a 'DeploymentPreference' set to 'BlueGreen' for the ECS service.
B.Use AWS CodeDeploy with a deployment group configured for blue/green deployment, and an Application Load Balancer (ALB) to shift traffic between the blue and green target groups.
C.Use the ECS service's built-in rolling update with a 'minimumHealthyPercent' of 100 and 'maximumPercent' of 200.
D.Configure the ECS service with an 'AutoScaling' policy that replaces instances gradually.
AnswerB

AWS CodeDeploy natively supports blue/green deployments for ECS services when integrated with an Application Load Balancer. During deployment, CodeDeploy registers the new task set with a 'green' target group while the existing 'blue' task set remains registered with the original target group; the ALB listener's traffic is then shifted from blue to green using either a canary, linear, or all-at-once traffic-shifting configuration. CodeDeploy also runs lifecycle hooks (BeforeInstall, AfterInstall, etc.) and preserves the original task set for instant rollback if the deployment fails. This is the only option that provides immutable infrastructure and controlled traffic shifting, making it the correct answer.

Why this answer

AWS CodeDeploy natively supports blue/green deployments for Amazon ECS (Fargate) by orchestrating traffic shifting between two target groups behind an Application Load Balancer (ALB). This allows the new task set (green) to be validated before production traffic is fully shifted, and enables instant rollback by reverting traffic to the original (blue) target group. CodePipeline can integrate CodeDeploy as a deploy action to automate this workflow.

Exam trap

The trap here is that candidates confuse the ECS rolling update configuration (minimumHealthyPercent/maximumPercent) with a blue/green strategy, but rolling updates do not provide separate target groups or instant rollback, which are key requirements for the scenario.

How to eliminate wrong answers

Option A is wrong because AWS CloudFormation does not support a 'DeploymentPreference' property set to 'BlueGreen' for ECS services; CloudFormation's 'DeploymentController' only supports 'ECS' (rolling update) or 'EXTERNAL' (external deployment), and blue/green for ECS requires CodeDeploy. Option C is wrong because setting 'minimumHealthyPercent' to 100 and 'maximumPercent' to 200 still performs a rolling update, not a blue/green deployment; it replaces tasks incrementally without maintaining two separate environments for traffic shifting, and does not support instant rollback. Option D is wrong because ECS Auto Scaling policies adjust the desired count of tasks based on metrics (e.g., CPU/memory) and do not implement deployment strategies; they cannot shift traffic between distinct task sets or provide blue/green rollback capabilities.

75
MCQeasy

A DevOps team is deploying a new web application on AWS Elastic Beanstalk. They want to monitor the application's health and receive notifications when the environment's health status changes to 'Degraded' or 'Severe'. What is the simplest way to achieve this?

A.Use the Elastic Beanstalk management console to manually check the health status twice a day.
B.Create a CloudWatch alarm on the 'EnvironmentHealth' metric published by the Elastic Beanstalk environment.
C.Write a custom script that polls the Elastic Beanstalk DescribeEnvironmentHealth API and sends an email using Amazon SES.
D.Configure an AWS CloudTrail trail to monitor Elastic Beanstalk API calls and create a CloudWatch alarm on the trail.
AnswerB

Elastic Beanstalk with enhanced health reporting publishes the 'EnvironmentHealth' metric to CloudWatch, and creating an alarm on this metric allows you to react automatically when the environment transitions to states such as severe or degraded. The alarm can publish to an SNS topic to email or page the team, enabling real-time incident response without any custom code. This approach is the native, supported integration between Elastic Beanstalk health and CloudWatch observability.

Why this answer

Elastic Beanstalk automatically publishes environment health metrics to Amazon CloudWatch, including the 'EnvironmentHealth' metric which reports values like 0 (Ok), 1 (Info), 5 (Unknown), 10 (NoData), 15 (Warning), 20 (Degraded), and 25 (Severe). Creating a CloudWatch alarm on this metric with a threshold of >= 20 (or specifically for Degraded/Severe) and configuring an SNS notification is the simplest, most direct, and fully managed solution. This requires no custom code, no polling infrastructure, and leverages native AWS integration.

Exam trap

DOP-C02 often tests the difference between CloudWatch (metrics and alarms) and CloudTrail (API activity logging), so candidates may incorrectly choose CloudTrail for monitoring health status changes.

How to eliminate wrong answers

Option A is wrong because manual console checks are not automated, are error-prone, and cannot provide real-time notifications, violating the requirement for automatic alerts on health changes. Option C is wrong because writing a custom polling script with SES introduces unnecessary complexity, requires managing credentials, scheduling, and error handling, and is not the simplest approach when CloudWatch alarms already exist. Option D is wrong because CloudTrail records API calls for auditing, not environment health metrics; it does not capture the 'EnvironmentHealth' metric and cannot be used to alarm on health status changes.

Page 1 of 18

Page 2