Courseiva

AWS Certified DevOps Engineer Professional DOP-C02 (DOP-C02) — Questions 451–525

1298 questions total · 18pages · All types, answers revealed

Page 6

Page 7 of 18

Page 8
451
MCQhard

A company is migrating its on-premises applications to AWS and wants to maintain the same level of monitoring for its Linux-based EC2 instances. They currently use Nagios for monitoring. They want a managed AWS service that can monitor instance health, system metrics, and application logs. Which solution should they use?

A.Install the Amazon CloudWatch agent on each EC2 instance to collect system metrics and logs, and send them to CloudWatch.
B.Use AWS CloudTrail to monitor instance activity and capture log files.
C.Use AWS Systems Manager Inventory to collect system configuration and log files.
D.Use AWS Config to track instance configuration changes and trigger alerts.
AnswerA

The CloudWatch agent runs on each instance to collect system metrics and application logs, then ships them to CloudWatch. This delivers the managed monitoring service the company requires, replacing Nagios without self-managed infrastructure, and satisfies the instance health, metrics and logs monitoring requirement.

Why this answer

The CloudWatch agent is the managed AWS solution for collecting guest-level system metrics (memory, disk, swap) and application/system logs from EC2 instances and delivering them to CloudWatch. It replaces the need for a self-managed Nagios server by providing native metric and log ingestion, alarms, and dashboards. Installing it on each instance satisfies the requirement for instance health, system metrics, and application logs in a managed service.

Exam trap

The trap is confusing AWS monitoring services: CloudTrail (API audit), Config (configuration compliance), and Systems Manager Inventory (metadata) are often mistaken for a metrics/logs monitoring solution, but only the CloudWatch agent collects guest OS metrics and application logs.

How to eliminate wrong answers

Option B is wrong because CloudTrail records API activity and management events, not guest OS metrics or application log files, so it cannot replace Nagios-style monitoring. Option C is wrong because Systems Manager Inventory collects configuration metadata (installed applications, OS details) rather than real-time system metrics and application logs. Option D is wrong because AWS Config tracks resource configuration changes and compliance, not runtime health metrics or log content.

452
Multi-Selecthard

Which THREE components are required to set up a fully automated CI/CD pipeline for a static website hosted on Amazon S3 using AWS CodePipeline? (Choose THREE.)

Select 3 answers
A.An AWS CodeBuild project to run build commands (e.g., minification)
B.An Amazon CloudFront distribution for content delivery
C.An AWS CodeCommit repository to store the website source code
D.An S3 bucket configured for static website hosting as the deployment target
E.An AWS Lambda function to invalidate CloudFront cache
AnswersA, C, D

CodeBuild is the required compute stage in a static website pipeline: it pulls the source from CodeCommit, executes build commands such as minification, Sass compilation, or image optimization, and outputs the production-ready artifacts to a staging location. Without a build project, the pipeline would simply copy raw source files, missing any transformation needed for the final site. CodeBuild integrates directly with CodePipeline and writes artifacts to S3, making it the only component explicitly running build commands.

Why this answer

AWS CodeBuild is required to execute build commands such as minification, bundling, or transpilation of static website assets before deployment. In a fully automated CI/CD pipeline, CodeBuild processes the source code from the repository and produces the deployable artifacts that are then uploaded to the S3 bucket.

Exam trap

The trap here is that candidates often confuse optional performance enhancements (CloudFront) or cache invalidation mechanisms (Lambda) as mandatory pipeline components, when the question specifically asks for the three required components to set up a fully automated CI/CD pipeline for a static website hosted on S3.

453
MCQmedium

Refer to the exhibit. An IAM policy is attached to a CodePipeline service role. When the pipeline tries to start a CodeBuild project, it fails with an 'AccessDenied' error. The CodeBuild project uses a different service role (arn:aws:iam::123456789012:role/CodeBuildServiceRole2). What is the MOST likely cause?

A.The policy has a condition on the s3 actions that is not satisfied.
B.The policy does not allow codebuild:StartBuild for the specific project.
C.The policy does not allow s3:GetObject on the artifact bucket.
D.The policy only allows iam:PassRole for a specific role ARN, but the CodeBuild project uses a different role.
AnswerD

The pipeline role's policy scopes `iam:PassRole` to a single resource ARN, so CodePipeline cannot hand CodeBuildServiceRole2 to CodeBuild. Starting a build requires passing the project's service role; the mismatched ARN triggers AccessDenied. Broadening the resource to the intended role ARN resolves it.

Why this answer

When CodePipeline starts a CodeBuild project, it must pass the CodeBuild service role to CodeBuild using iam:PassRole. If the pipeline's service role policy only allows iam:PassRole for a specific role ARN (e.g., CodeBuildServiceRole1) but the CodeBuild project is configured with a different role (CodeBuildServiceRole2), the PassRole action will be denied, causing an AccessDenied error. This is the most likely cause because the error occurs at the start of the build, and the policy explicitly restricts the role that can be passed.

Exam trap

DOP-C02 often tests the iam:PassRole permission and its role in service-to-service delegation; candidates may overlook that the pipeline role needs explicit PassRole permission for the specific CodeBuild service role, leading to AccessDenied errors.

How to eliminate wrong answers

Option A is wrong because the error is about starting the CodeBuild project, not about S3 actions; S3 conditions would affect artifact retrieval or storage, not the StartBuild call. Option B is wrong because if the policy lacked codebuild:StartBuild permission, the error would occur, but the question implies the policy is attached and the failure is due to role passing; also, the policy might allow StartBuild but not PassRole. Option C is wrong because missing s3:GetObject would cause failures when downloading artifacts, not when starting the build.

454
MCQeasy

A company uses AWS Organizations with multiple accounts. The security team needs to enforce that all new member accounts automatically receive a specific AWS Config rule to require encryption on Amazon EBS volumes. Which solution meets this requirement with the least operational overhead?

A.Use an SCP to deny the creation of unencrypted EBS volumes and use AWS Config to detect noncompliant volumes.
B.Use a service control policy (SCP) to deny the ability to disable the AWS Config rule and use a custom AWS Config rule that evaluates EBS encryption.
C.Use an AWS Config aggregator in the management account to monitor compliance across accounts.
D.Use AWS CloudFormation StackSets to deploy a stack with the Config rule to all existing and new accounts.
AnswerD

AWS CloudFormation StackSets let you define a stack containing the AWS Config rule once and deploy it to specified accounts across the organization. With service-managed StackSets, you can enable automatic deployment so that any new account added to the organization is automatically provisioned with the stack, ensuring the Config rule exists in every account. This centrally manages the rule's lifecycle and avoids needing per-account manual setup, making it the appropriate solution.

Why this answer

AWS CloudFormation StackSets can deploy a stack containing the AWS Config rule across all accounts in the organization, and with automatic deployment enabled, new accounts automatically receive the stack. This centrally manages the rule with minimal operational overhead. Option B is incorrect because SCPs only deny API actions and cannot deploy a Config rule; they would need a separate mechanism to deploy the rule initially.

Exam trap

The trap is that candidates may believe an SCP can be used to enforce a Config rule, but SCPs only deny or allow actions; they don't deploy configurations. The correct approach uses StackSets or organization-level Config rules for automatic deployment.

How to eliminate wrong answers

Option A is wrong because an SCP that denies the creation of unencrypted EBS volumes does not enforce an AWS Config rule; it only prevents creation but does not detect or remediate existing noncompliant volumes, and it does not automatically deploy the Config rule to new accounts. Option C is wrong because an AWS Config aggregator only provides a centralized view of compliance across accounts but does not enforce or deploy the Config rule to new accounts. Option D is wrong because CloudFormation StackSets require manual setup and ongoing management to deploy to new accounts as they are added, which introduces higher operational overhead compared to using organization-level AWS Config rules or SCPs.

455
MCQhard

A company runs a containerized application on Amazon ECS with Fargate launch type. The application consists of three microservices: frontend, backend, and database. The ECS cluster is in a VPC with public and private subnets. The frontend service is publicly accessible via an Application Load Balancer (ALB) in public subnets. The backend service communicates with the database service, which runs as a stateful service with persistent storage using Amazon EFS. The DevOps team is using CloudWatch Container Insights and has enabled Prometheus metrics for the ECS cluster. Recently, the team observed that the frontend service's response time has increased significantly, and some requests are timing out. The team checked the ALB metrics and saw an increase in 5xx errors. They also noticed that the backend service's CPU utilization is high, and the database service's disk I/O is high. The team suspects a bottleneck in the backend service. Which course of action should the team take FIRST to identify the root cause?

A.Disable the health check for the backend service in the ALB target group.
B.Migrate the database service to Amazon RDS for better performance.
C.Check the backend service's application logs in CloudWatch Logs to identify errors or slow database queries.
D.Increase the desired count of the backend service to reduce load per task.
AnswerC

Checking the backend service's application logs in CloudWatch Logs is the correct initial action because it provides direct visibility into application errors, database query execution times, and slow transactional paths. These logs, combined with ECS task metrics and ALB access logs, help isolate whether the high latency is due to application code, database contention, or an upstream dependency. Logs are the least intrusive and most informative diagnostic step, enabling an evidence-based decision before changing infrastructure.

Why this answer

The first step is to analyze the backend service's application logs to identify any errors or slow operations. High CPU and disk I/O may be caused by inefficient queries or code issues. Option A is incorrect because disabling health checks would hide the problem and could route traffic to unhealthy tasks.

Option B is incorrect because migrating to RDS does not address the immediate issue and is a significant change without root cause analysis. Option D is incorrect because increasing the desired count without understanding the root cause may temporarily alleviate load but does not fix underlying performance issues and can increase costs.

456
Multi-Selectmedium

A DevOps engineer is setting up centralized logging for a multi-account environment using AWS Organizations. The engineer needs to aggregate logs from all accounts into a single Amazon S3 bucket. Which TWO steps are necessary?

Select 2 answers
A.Create IAM roles in each account to allow the central bucket to read logs.
B.Create a bucket policy on the central S3 bucket that grants permissions to the source accounts.
C.Enable CloudTrail organization trail in the management account to deliver logs to the central bucket.
D.Set up a cross-account subscription in CloudWatch Logs to forward logs to the central account.
E.Configure each account’s services (e.g., CloudTrail, VPC Flow Logs) to deliver logs to the central S3 bucket.
AnswersB, E

The central S3 bucket needs a resource-based policy that explicitly grants the source accounts' log-delivery services (e.g., CloudTrail, VPC Flow Logs) permission to write objects and read the bucket ACL. Without this bucket policy, cross-account writes from other accounts will be denied by default. The policy must reference the source account IDs or the organization ID and include conditions like aws:SourceAccount or aws:SourceArn to prevent confused deputy attacks. This is a mandatory step to enable centralized log collection.

Why this answer

A bucket policy on the central S3 bucket can grant cross-account permissions to source accounts to write logs. This allows services like CloudTrail and VPC Flow Logs from member accounts to deliver logs directly to the central bucket without requiring IAM roles in each account for reading logs.

Exam trap

The trap here is that candidates often confuse the need for IAM roles in each account (Option A) with the correct bucket policy approach, or they assume that enabling an organization trail (Option C) is mandatory when the question allows for individual account configuration.

457
MCQhard

A company uses Amazon S3 to store sensitive data. The security team wants to be notified when an S3 bucket policy is modified. Which approach is most efficient?

A.Create an Amazon EventBridge rule that matches the 'PutBucketPolicy' API call and sends a notification to an SNS topic.
B.Set up an AWS Config rule to detect changes to the bucket policy.
C.Configure S3 event notifications for 's3:PutBucketPolicy' on the bucket.
D.Enable S3 server access logs and use CloudWatch Logs Insights to run queries periodically.
AnswerA

EventBridge matches the PutBucketPolicy API call from CloudTrail and routes it to an SNS topic, giving near-real-time notification without polling. This is more efficient than scheduled log analysis or S3 event notifications, which do not cover bucket policy changes.

Why this answer

Amazon EventBridge can match AWS API calls such as PutBucketPolicy via CloudTrail management events and route them to an SNS topic for notification. This is the most efficient, near-real-time, event-driven approach for alerting on bucket policy changes.

Exam trap

DOP-C02 often tests the difference between event-driven detection (EventBridge) and periodic/compliance-based detection (Config, access logs) — candidates may pick S3 event notifications, which do not cover bucket policy API calls.

How to eliminate wrong answers

Option B is wrong because AWS Config rules detect configuration changes but are not designed for immediate event-driven notification and can be slower and less direct for this use case. Option C is wrong because S3 event notifications only fire for object-level events (like PutObject), not for bucket policy API calls such as PutBucketPolicy. Option D is wrong because S3 server access logs record data-plane requests and querying them periodically is not real-time or efficient for policy-change alerting.

458
MCQhard

Refer to the exhibit. An IAM policy is attached to a role used by a CI/CD system. The policy is intended to allow starting the pipeline 'MyPipeline' from the same account. However, the CI/CD system receives an 'AccessDenied' error when trying to start the pipeline. What is the problem?

A.The Allow statement does not specify the correct pipeline ARN.
B.The policy needs an additional Allow for 'codepipeline:GetPipeline' to start the pipeline.
C.The role does not have permission to pass the policy to the CI/CD system.
D.The Deny statement with the 'aws:SourceAccount' condition denies access if the condition key is not present in the request.
AnswerD

This Deny statement uses a condition key such as aws:SourceAccount with an operator like StringNotEquals, which evaluates as true when the key is missing from the request context. Because an explicit Deny overrides all Allow statements, the request is denied whenever the expected source account is not present in the request. This is the precise cause of the AccessDenied error.

Why this answer

The Deny statement with the `aws:SourceAccount` condition key denies access unless the request includes that condition key. When the CI/CD system assumes the role and makes a `StartPipelineExecution` API call, the request context does not automatically include the `aws:SourceAccount` key unless explicitly added by the caller. Since the condition is not satisfied, the Deny statement applies, resulting in an 'AccessDenied' error even though the Allow statement grants the necessary action.

Exam trap

The trap here is that candidates overlook the explicit Deny statement with a condition key, assuming the Allow statement alone is sufficient, and instead focus on missing permissions or incorrect ARNs, not realizing that an explicit Deny with an unsatisfied condition will block access regardless of any Allow.

How to eliminate wrong answers

Option A is wrong because the pipeline ARN in the Allow statement is specified as 'arn:aws:codepipeline:us-east-1:123456789012:MyPipeline', which is the correct format and matches the pipeline name 'MyPipeline'. Option B is wrong because `codepipeline:GetPipeline` is a read-only action and is not required to start a pipeline; the required action is `codepipeline:StartPipelineExecution`, which is already allowed. Option C is wrong because the policy is attached to the role, not passed to the CI/CD system; the role itself is assumed by the CI/CD system, and there is no 'pass policy' permission issue here.

459
MCQhard

A company is running a critical application on Amazon ECS with Fargate launch type. The application writes logs to Amazon CloudWatch Logs. The DevOps team needs to set up an alert when the application generates more than 100 error logs in any 5-minute window. Which configuration should be used?

A.Create a CloudWatch Logs Insights query that runs every 5 minutes and triggers an SNS notification
B.Create an Amazon EventBridge rule that matches CloudWatch Logs events for the word 'ERROR' and triggers an alarm
C.Create a CloudWatch Logs metric filter for 'ERROR' and a CloudWatch alarm on the resulting metric with a period of 5 minutes
D.Enable AWS CloudTrail logging for the ECS task and create a metric filter on CloudTrail logs
AnswerC

A CloudWatch Logs metric filter is applied to a log group in near-real time as log events are ingested, and it uses pattern syntax to count occurrences of the string 'ERROR' and emit a custom metric (e.g., ErrorCount) for that log group. The CloudWatch alarm is then configured on that custom metric with an evaluation period of 5 minutes, so if the number of ERROR log lines within a 5-minute interval exceeds the configured threshold, the alarm state changes to ALARM and can trigger an SNS notification. This is the standard, fully managed mechanism for log-pattern-based alerting because it does not require any custom code or additional infrastructure, and it integrates directly with CloudWatch alarm actions.

Why this answer

A CloudWatch Logs metric filter can be configured to count occurrences of the word 'ERROR' in log streams. This filter creates a custom metric that can be monitored by a CloudWatch alarm with a period of 5 minutes. When the metric exceeds the threshold of 100, the alarm triggers an action such as an SNS notification.

Option A is incorrect because CloudWatch Logs Insights is a query tool for interactive analysis, not for continuous real-time alerting. Option B is incorrect because EventBridge events are not generated from log content directly; you would need a metric filter or subscription filter to turn log data into events. Option D is incorrect because AWS CloudTrail logs API activities, not application error logs written to CloudWatch Logs.

460
MCQmedium

A company uses AWS CodeBuild to compile Java applications. The builds often fail due to insufficient memory. The buildspec currently specifies 'compute-type: BUILD_GENERAL1_SMALL'. What is the most cost-effective solution to resolve the memory issues without changing the build logic?

A.Change the compute-type to 'BUILD_GENERAL1_MEDIUM' or 'BUILD_GENERAL1_LARGE' in the buildspec.
B.Enable Amazon S3 caching for the build artifacts to reduce memory usage.
C.Split the build into multiple parallel CodeBuild actions in the pipeline, each compiling a subset of the code.
D.Set the environment variable 'MEMORY_OVERPROVISION=2' in the buildspec.
AnswerA

Upgrading the compute type in the buildspec—from the default BUILD_GENERAL1_SMALL (3 GB memory) to BUILD_GENERAL1_MEDIUM (7 GB) or BUILD_GENERAL1_LARGE (15 GB)—is the correct fix because CodeBuild's compute-type property directly controls the memory and vCPU allocated to the build container. Java compilation through Maven or Gradle is memory-intensive, and a small instance can easily cause the JVM or compiler daemon to run out of heap space. Selecting a larger general-purpose tier gives the build process a proportionally larger heap and OS memory, resolving the OOM without changing application code.

Why this answer

The most direct and cost-effective way to resolve insufficient memory in AWS CodeBuild is to increase the compute type to a larger size (e.g., BUILD_GENERAL1_MEDIUM or BUILD_GENERAL1_LARGE), which provides more memory without altering the build logic. This change incurs additional cost only when builds run, and it avoids unnecessary complexity or invalid approaches.

Exam trap

The trap here is that candidates may think caching or parallelism can solve memory issues, but neither addresses the fundamental lack of memory per build instance, and the fabricated environment variable 'MEMORY_OVERPROVISION' is designed to lure those who guess at undocumented features.

How to eliminate wrong answers

Option B is wrong because enabling Amazon S3 caching for build artifacts reduces build time by reusing cached dependencies, but it does not increase the available memory during the build process, so it cannot resolve out-of-memory errors. Option C is wrong because splitting the build into multiple parallel CodeBuild actions does not increase the memory per build instance; each action still runs on the same compute type with the same memory limit, and the compilation of a single subset may still fail if it requires more memory than the small instance provides. Option D is wrong because there is no such environment variable as 'MEMORY_OVERPROVISION' in AWS CodeBuild; this is a fabricated option that does not exist in the CodeBuild documentation.

461
MCQhard

A DevOps engineer receives the error shown in the exhibit when attempting to update an existing CloudFormation stack that deploys a VPC with subnets. The stack was created successfully earlier using the same template. What is the most likely cause of this error?

A.The subnet ID in the template is already used by another stack in the same account.
B.The IAM role used for the stack update lacks the 'ec2:DescribeSubnets' permission.
C.The subnet specified in the template does not exist in the selected AWS region.
D.The CloudFormation template has a syntax error in the subnet definition.
AnswerB

CloudFormation uses the IAM service role's permissions when performing stack updates. If the role lacks ec2:DescribeSubnets, the update fails with an access denied error, exactly as shown in the exhibit. This permission is required for CloudFormation to validate the subnet and retrieve its attributes during the update. Adding ec2:DescribeSubnets to the role's policy resolves the issue.

Why this answer

When updating a CloudFormation stack that deploys a VPC with subnets, the update operation must be able to read the current state of the subnet resources to determine if changes are needed. The IAM role used for the stack update must have the 'ec2:DescribeSubnets' permission to query the existing subnet configuration. Without this permission, CloudFormation cannot verify the subnet's current properties, leading to the error shown in the exhibit.

Exam trap

The trap here is that candidates often assume the error is due to a template syntax issue or resource conflict, but the real cause is insufficient IAM permissions for the update operation, which is a subtle but critical distinction in CloudFormation stack management.

How to eliminate wrong answers

Option A is wrong because subnet IDs are unique within an AWS account per region, and CloudFormation does not reuse subnet IDs across stacks; the error is not about ID conflicts. Option C is wrong because the stack was created successfully earlier using the same template, so the subnet does exist in the region; the error occurs during the update, not the initial creation. Option D is wrong because a syntax error in the template would have been caught during the initial stack creation, not during an update of an existing stack that was previously created successfully.

462
MCQeasy

A development team wants to automate infrastructure provisioning using AWS CloudFormation. Which tool is specifically designed to manage CloudFormation templates as part of a deployment pipeline?

A.AWS CodePipeline
B.AWS CodeCommit
C.AWS CodeBuild
D.AWS CodeDeploy
AnswerA

AWS CodePipeline is a fully managed continuous delivery service that orchestrates the entire release workflow, from source through build to deployment. As an orchestration engine, it can directly invoke AWS CloudFormation actions to provision or update infrastructure stacks at any stage. It supports custom actions, manual approval gates, and seamless integration with other AWS services, making it the correct choice for automating infrastructure provisioning as part of a CI/CD pipeline.

Why this answer

AWS CodePipeline is the correct answer because it is a fully managed continuous delivery service that allows you to model, visualize, and automate the steps required to release your infrastructure changes. You can integrate CloudFormation as a deployment action within a pipeline stage, enabling the automated creation, update, or deletion of stacks based on template changes committed to a source repository.

Exam trap

The trap here is that candidates often confuse AWS CodeDeploy (which deploys application code) with CloudFormation (which provisions infrastructure), leading them to select CodeDeploy instead of recognizing that CodePipeline is the service that orchestrates the entire deployment pipeline including CloudFormation actions.

How to eliminate wrong answers

Option B (AWS CodeCommit) is wrong because it is a source control service for storing Git repositories, not a pipeline orchestrator; it cannot execute CloudFormation templates or manage deployment stages. Option C (AWS CodeBuild) is wrong because it is a fully managed build service that compiles source code, runs tests, and produces artifacts, but it does not orchestrate multi-stage pipelines or directly deploy CloudFormation stacks. Option D (AWS CodeDeploy) is wrong because it automates code deployments to compute services like EC2 or Lambda, but it is not designed to manage infrastructure provisioning via CloudFormation templates.

463
MCQhard

A company runs a production e-commerce platform on AWS. The architecture includes an Application Load Balancer (ALB) that distributes traffic to a fleet of Amazon EC2 instances running in an Auto Scaling group across three Availability Zones (AZs). The application stores session state in Amazon ElastiCache for Redis (cluster mode disabled) with a single node. The database is an Amazon Aurora MySQL DB cluster with one writer and two reader instances in different AZs. The platform experiences intermittent slowdowns and occasional timeouts during peak traffic hours. The CloudWatch metrics show that the ALB's TargetResponseTime is elevated, and the Redis CPU utilization is consistently above 80% during these periods. The Auto Scaling group is scaling out, but new instances take several minutes to become healthy. The DevOps team has been asked to improve the resilience and performance of the application with minimal changes to the application code. Which solution should the team implement?

A.Replace the ALB with a Network Load Balancer (NLB) to reduce latency, and use an Auto Scaling group with a step scaling policy based on Redis CPU utilization.
B.Increase the instance size of the ElastiCache for Redis node and the size of the Aurora writer instance. Also, increase the cooldown period for the Auto Scaling group to allow new instances to warm up.
C.Implement Amazon RDS Proxy in front of the Aurora cluster to reduce database connection overhead, and increase the size of the Redis instance to handle more connections.
D.Migrate ElastiCache for Redis to a cluster mode enabled configuration with multiple shards and enable Multi-AZ with automatic failover. Also, use an ElastiCache replication group with read replicas in different AZs.
AnswerD

Migrating to cluster mode enabled with multiple shards horizontally partitions the Redis keyspace across nodes, which directly lowers per-shard CPU utilization and enables linear scaling as traffic grows. Enabling Multi-AZ with automatic failover and placing read replicas in different AZs provides high availability and lets reads be served by replicas, reducing primary node load and cutting failover time from minutes to seconds—this tackles both the immediate CPU bottleneck and the resilience requirement for a production e-commerce platform.

Why this answer

The primary bottleneck is the single-node Redis instance (CPU > 80%), which cannot scale horizontally and lacks high availability. Migrating to cluster mode enabled with multiple shards distributes CPU load across shards, while Multi-AZ with automatic failover and read replicas in different AZs provides high availability and read scaling. This directly addresses the elevated ALB TargetResponseTime caused by Redis latency.

Note that the application will need to use a Redis Cluster-compatible client; since the stem allows minimal code changes, this solution is still appropriate.

Exam trap

The trap here is that candidates focus on scaling the database or load balancer (options A, B, C) instead of recognizing that the single-node Redis cache is the bottleneck and requires horizontal scaling and high availability to resolve both performance and resilience issues.

How to eliminate wrong answers

Option A is wrong because replacing the ALB with an NLB does not reduce application-layer latency (NLB operates at Layer 4, not Layer 7, and cannot offload TLS or inspect HTTP sessions), and a step scaling policy based on Redis CPU utilization does not fix the single-node Redis bottleneck or the slow instance warm-up. Option B is wrong because increasing the instance size of the single Redis node and the Aurora writer instance only vertically scales the existing bottlenecks, and increasing the Auto Scaling group cooldown period would delay scaling further, worsening the timeouts. Option C is wrong because RDS Proxy reduces database connection overhead but does not address the Redis CPU bottleneck (the primary cause of elevated response times), and increasing the Redis instance size alone does not provide the read scaling or high availability needed.

464
MCQmedium

A DevOps engineer is setting up centralized logging for multiple AWS accounts. They need to collect VPC Flow Logs, CloudTrail logs, and application logs into a single Amazon S3 bucket. What is the most efficient approach?

A.Configure a Lambda function in each account to copy logs to a central S3 bucket.
B.Create an S3 bucket in each account and use S3 replication.
C.Use Amazon Kinesis Data Firehose to stream logs from all accounts to a central S3 bucket.
D.Use an S3 bucket in a centralized logging account with a bucket policy that grants write access from all other accounts.
AnswerD

Placing a single S3 bucket in a centralized logging account with a bucket policy that grants s3:PutObject to principals in all source accounts is the most efficient pattern because CloudTrail, VPC Flow Logs, and similar services can deliver logs cross-account natively without additional moving parts. The bucket policy can restrict writes using aws:SourceArn or aws:SourceAccount conditions to prevent the confused-deputy problem, and the central account owns objects for unified lifecycle and access control. This direct-write model achieves lower latency and zero maintenance compared to middleware or copying mechanisms.

Why this answer

It uses a centralized logging account with a single S3 bucket configured with a bucket policy that grants write access (s3:PutObject) to all other accounts. This approach avoids data duplication, eliminates the need for replication or intermediate compute resources, and is the most efficient and cost-effective method for aggregating logs from multiple accounts into a single destination.

Exam trap

The trap here is that candidates often overcomplicate the solution by choosing managed services like Kinesis or Lambda, when a simple S3 bucket policy with cross-account write access is the most efficient and AWS-recommended approach for centralized log aggregation.

How to eliminate wrong answers

Option A is wrong because using a Lambda function in each account to copy logs to a central S3 bucket introduces unnecessary complexity, potential single points of failure, and additional cost from Lambda invocations and data transfer, making it less efficient than a direct write approach. Option B is wrong because creating an S3 bucket in each account and using S3 replication results in data duplication, increased storage costs, and replication latency, and it requires managing multiple buckets and replication rules, which is less efficient than a single bucket with a cross-account policy. Option C is wrong because Amazon Kinesis Data Firehose is designed for streaming data ingestion and transformation, but it adds unnecessary complexity and cost for log aggregation when a simpler S3 bucket policy can achieve the same result; Firehose is better suited for real-time processing needs, not for batch log collection from multiple accounts.

465
Multi-Selecteasy

Which TWO actions should a DevOps engineer take to ensure that an AWS CodeBuild project can access a private Amazon S3 bucket to download build artifacts? (Choose two.)

Select 2 answers
A.Generate an access key and secret key for the CodeBuild project to use in the buildspec
B.Add a bucket policy that explicitly allows access from the CodeBuild project's IAM role ARN
C.Configure the CodeBuild project to run in a VPC with an S3 VPC endpoint
D.Attach an IAM role to the CodeBuild project with s3:GetObject permissions for the bucket
E.Make the S3 bucket publicly readable
AnswersB, D

Adding a bucket policy that explicitly grants the CodeBuild project's IAM role ARN the s3:GetObject action is a resource-based policy approach. It complements the identity-based policy attached to the role and ensures S3 evaluates the permission grant directly, which is required when the bucket account enforces resource policies or when the bucket resides in a different account. This explicit allow on the bucket side removes any dependency on implicit same-account authorization and is a secure, scoped way to permit object retrieval.

Why this answer

A bucket policy can explicitly grant access to the CodeBuild project's IAM role ARN, allowing the CodeBuild service to download objects from the private S3 bucket. This is a common method to cross-account or service-specific access without making the bucket public.

Exam trap

The trap here is that candidates often confuse network-level controls (VPC endpoints) with IAM permissions, thinking that a VPC endpoint alone is sufficient to grant access to S3, when in fact IAM policies or bucket policies are still required.

466
Multi-Selecthard

A company runs a web application on Amazon EC2 instances behind an Application Load Balancer. The DevOps team has enabled detailed CloudWatch metrics for the ALB and is using CloudWatch Logs for the EC2 instances. Recently, users report intermittent 503 errors. The team notices that the ALB's 'RequestCount' metric shows a sudden drop during error periods, while the 'ActiveConnectionCount' remains steady. Which TWO steps should the team take to diagnose the issue? (Choose two.)

Select 2 answers
A.Enable and analyze the ALB access logs to see the HTTP response codes and target processing time.
B.Check the Amazon Route 53 health checks for the ALB DNS name.
C.Review the EC2 instances' CloudWatch metrics for CPU utilization and network traffic.
D.Observe the ALB's 'UnhealthyHostCount' metric and check target group health checks.
E.Inspect AWS CloudTrail logs for the ALB to see if there are any configuration changes.
AnswersA, D

ALB access logs capture full request-level detail for every HTTP request the load balancer receives, including the exact HTTP response code returned to the client (e.g., 503) and the target processing time, which is the time the EC2 instance took to respond. By filtering for 503 responses and correlating them with target IPs and timestamps, you can confirm whether errors originate from unhealthy targets, overloaded instances, or misconfigured routing. This is the most direct diagnostic because access logs are designed to expose the request/response path, rather than merely indicating that a threshold was crossed.

Why this answer

ALB access logs contain detailed per-request data, including HTTP response codes (e.g., 503) and target processing time. Analyzing these logs will reveal whether the 503 errors are coming from the ALB itself (e.g., due to request queue overflow) or from the targets, and whether the sudden drop in RequestCount is due to clients aborting or the ALB throttling requests.

Exam trap

The trap here is that candidates often focus on EC2-level metrics (CPU, network) or CloudTrail config changes, missing that the ALB's own health check and access log data are the direct sources for diagnosing 503 errors tied to target unavailability or request queue limits.

467
MCQmedium

A DevOps engineer receives an alarm that an EC2 instance's StatusCheckFailed metric has been in ALARM state for 10 minutes. Which action should the engineer take first to investigate?

A.Review the instance's system log and application logs in CloudWatch Logs
B.Use AWS Config to check the instance's configuration compliance
C.Check AWS CloudTrail for any API calls that modified the instance
D.Restart the EC2 instance to clear the alarm
AnswerA

Instance status checks detect OS-level problems such as failed system boot, kernel panic, or filesystem corruption. The instance's system log (console output) shows the boot sequence and kernel messages, while application logs streamed to CloudWatch Logs reveal software-level errors that may have caused the failure. Reviewing both gives the evidence needed to identify the root cause before taking any corrective action.

Why this answer

When StatusCheckFailed is in ALARM, the first investigative step is to examine the instance's system log and application logs in CloudWatch Logs, because these reveal whether the failure is an OS-level crash, kernel panic, or application error causing the instance to fail its status checks. Reviewing logs is non-destructive and provides the diagnostic evidence needed before taking any remediation action.

Exam trap

The trap is treating a remediation action (restart the instance) as the first step — the exam tests whether you know to investigate and gather diagnostic evidence from logs before taking any corrective action that could destroy the evidence.

How to eliminate wrong answers

Option B is wrong because AWS Config tracks resource configuration changes and compliance, not runtime health or the cause of a status check failure — it would not explain why the instance is failing. Option C is wrong because CloudTrail records API activity (who changed what), which is useful if you suspect a recent configuration change, but it does not diagnose the instance's runtime failure and is not the first step for a health alarm. Option D is wrong because restarting the instance is a remediation action, not an investigation step, and it destroys the running state that could reveal the root cause — rebooting before gathering logs is a classic anti-pattern.

468
Multi-Selecteasy

A company runs a critical application on EC2 instances in an Auto Scaling group. The application must be highly available across multiple Availability Zones. Which TWO configurations are necessary to achieve this? (Choose TWO.)

Select 2 answers
A.Use a single Availability Zone to reduce latency.
B.Use Spot Instances to reduce costs.
C.Use a Classic Load Balancer to distribute traffic.
D.Place an Application Load Balancer in front of the Auto Scaling group.
E.Configure the Auto Scaling group to launch instances in multiple Availability Zones.
AnswersD, E

An Application Load Balancer operates at Layer 7 and can distribute HTTP/HTTPS traffic across instances in multiple Availability Zones, providing health checks and automatic registration of instances in the Auto Scaling group. This integration allows the ALB to route traffic only to healthy instances and to scale with the ASG's changes, enhancing availability and fault tolerance. It is the recommended pattern for web applications requiring high availability and advanced routing.

Why this answer

To achieve high availability across multiple Availability Zones (AZs), two key configurations are required. First, the Auto Scaling group must be configured to launch instances in multiple AZs (Option E). This ensures that if one AZ fails, the application can continue serving traffic from instances in other AZs.

Second, an Application Load Balancer (ALB) should be placed in front of the Auto Scaling group (Option D). The ALB distributes incoming traffic across instances in all AZs where the Auto Scaling group has launched instances. It also performs health checks and automatically routes traffic away from unhealthy instances.

Option A is incorrect because using a single AZ would create a single point of failure, undermining high availability. Option B is incorrect because Spot Instances can be terminated with little notice, which is unsuitable for a critical application requiring consistent availability. Option C is incorrect because a Classic Load Balancer lacks advanced features like path-based routing and native cross-zone load balancing, making the ALB a better choice for modern applications.

469
MCQhard

An organization uses AWS Elastic Beanstalk for application deployments. They want to implement immutable updates to minimize downtime and ensure that if the new environment fails health checks, the old environment remains intact. Which deployment policy should they choose?

A.Traffic splitting.
B.Immutable update.
C.All at once.
D.Rolling update based on health.
AnswerB

Immutable update deploys the new application version to a completely new set of instances (or a fully separate environment) that runs alongside the old fleet. It waits for all new instances to pass health checks before shifting traffic to them, and if any fail, the old environment remains untouched, enabling an immediate, zero-impact rollback. This provides the strongest availability and isolation, making it the safest deployment policy.

Why this answer

Immutable updates in AWS Elastic Beanstalk launch a completely new environment with the new application version. If the new environment fails health checks, Elastic Beanstalk automatically terminates it, leaving the original environment untouched. This ensures zero downtime and a safe rollback, which matches the requirement to keep the old environment intact if health checks fail.

Exam trap

The trap here is that candidates confuse 'immutable update' with 'traffic splitting' because both involve a new environment, but traffic splitting does not automatically terminate the new environment on health check failure—it requires manual intervention or additional automation to roll back.

How to eliminate wrong answers

Option A is wrong because traffic splitting gradually shifts a percentage of traffic to a new environment, but if health checks fail, the old environment is not guaranteed to remain intact—the new environment may still be partially serving traffic and the rollback is not fully automated. Option C is wrong because all-at-once deploys replace all instances simultaneously, causing downtime and leaving no fallback environment if health checks fail. Option D is wrong because rolling update based on health replaces instances in batches and can terminate unhealthy instances in the old environment, potentially disrupting the original environment before the new one is fully verified.

470
MCQmedium

A company uses AWS CloudTrail to log API activity across multiple accounts. The security team needs to ensure that all CloudTrail logs are delivered to a centralized S3 bucket in the audit account, and that any log file validation failures trigger an immediate notification. What should the engineer do to meet this requirement?

A.Enable CloudTrail log file validation and create a CloudWatch alarm on the DigestDeliveryFailed metric
B.Create a Lambda function that checks the integrity of logs and publishes to SNS
C.Configure CloudTrail to deliver logs to the S3 bucket and enable SNS notifications for all events
D.Send CloudTrail logs to CloudWatch Logs and create a metric filter for validation errors
AnswerA

CloudTrail's built-in log file validation creates hash-chained digest files that are digitally signed with a private key, and you can verify them with the public key published by AWS. The CloudTrail service emits the DigestDeliveryFailed metric to CloudWatch when it cannot deliver a digest file, which is often the first sign of tampering or a delivery problem. Creating a CloudWatch alarm on this metric gives you immediate, proactive notification without building custom code.

Why this answer

Enabling CloudTrail log file validation triggers the generation of digest files that contain hash values for verifying log file integrity. CloudTrail also emits the DigestDeliveryFailed metric to CloudWatch when a digest file delivery fails. Creating a CloudWatch alarm on this metric allows you to send immediate notifications via SNS when a validation failure occurs.

Option B (Lambda function) is unnecessary because CloudTrail already provides the necessary metrics for this alerting. Option C (SNS notifications for all events) would generate excessive notifications and does not directly address log file validation failures. Option D (CloudWatch Logs and metric filter) is not the standard approach; CloudTrail directly emits the DigestDeliveryFailed metric, which is simpler and more reliable.

471
MCQmedium

A company has a multi-account AWS environment using AWS Organizations. They want to centrally manage user access to all accounts using single sign-on (SSO) and enforce multi-factor authentication (MFA). Which service should they use?

A.Use AWS Secrets Manager to store and rotate IAM user credentials.
B.Create IAM users in each account and share the credentials securely.
C.Use Amazon Cognito user pools with an identity broker.
D.Use AWS IAM Identity Center (AWS SSO) to manage access and enforce MFA.
AnswerD

AWS IAM Identity Center centralizes workforce identity and builds on AWS Organizations to give users SSO access to all accounts via permission sets, which define granular IAM role permissions. It enforces MFA with policy settings such as requiring MFA for all users or context-dependent MFA, and supports both its built-in directory and external identity providers. Users sign in once at the portal or via the CLI, and IAM Identity Center automatically creates temporary credentials for each account.

Why this answer

AWS IAM Identity Center (formerly AWS SSO) is the correct service because it provides a centralized place to manage user access and permissions across all AWS accounts in an AWS Organization. It natively supports enforcing multi-factor authentication (MFA) through an identity source (e.g., the built-in identity store or an external IdP) and integrates directly with AWS Organizations to grant single sign-on access without needing to create IAM users in each account.

Exam trap

The trap here is that candidates often confuse Amazon Cognito (a customer identity service) with workforce identity management, or assume that storing credentials in Secrets Manager or creating per-account IAM users is a viable centralized solution, when in fact AWS IAM Identity Center is the only service designed for multi-account SSO with MFA enforcement in an AWS Organizations context.

How to eliminate wrong answers

Option A is wrong because AWS Secrets Manager is designed to securely store and rotate secrets (like database credentials or API keys), not to manage user identities or enforce MFA for SSO access. Option B is wrong because creating IAM users in each account and sharing credentials manually violates the principle of least privilege, creates a massive administrative overhead, and does not provide centralized SSO or consistent MFA enforcement across accounts. Option C is wrong because Amazon Cognito user pools are intended for customer-facing identity and access management for web and mobile applications, not for managing workforce access to AWS accounts via SSO with MFA enforcement across an AWS Organization.

472
MCQmedium

A DevOps team needs to monitor failed API calls in their AWS account. They want to receive notifications when specific IAM actions, such as DeleteBucket, fail. Which service should they use?

A.AWS CloudTrail and Amazon EventBridge.
B.AWS Config rules.
C.Amazon S3 server access logs.
D.CloudWatch Logs and metric filters.
AnswerA

AWS CloudTrail records all management API calls in the account, including failed attempts, with metadata such as the IAM principal, source IP, event name, and error codes. Amazon EventBridge can be configured with a rule whose event pattern matches CloudTrail's api_call events and a filter condition on errorCode, routing the matched events to an SNS topic for real-time alerting. This combination is purpose-built for monitoring failed API calls.

Why this answer

AWS CloudTrail captures API calls, and Amazon EventBridge (formerly CloudWatch Events) can be used to create rules that match specific failed API calls (e.g., DeleteBucket) and trigger notifications. Option B is incorrect because AWS Config rules monitor resource configuration compliance, not API call failures. Option C is incorrect because S3 server access logs log requests made to an S3 bucket, not IAM API calls.

Option D is incorrect because CloudWatch Logs and metric filters are used to monitor log data, but they are not the primary service for capturing API calls; CloudTrail is needed for that.

473
MCQeasy

A DevOps engineer is troubleshooting an Auto Scaling group (ASG) that is not launching instances as expected. The ASG is configured with a launch template that uses an Amazon Linux 2 AMI. The engineer checks the EC2 Auto Scaling console and sees that the group's desired capacity is set to 2, but only 1 instance is running. The last scaling activity shows 'Failed to launch instance. Error: Your quota allows for 0 more running instance(s).' What is the most likely cause?

A.The launch template has insufficient IAM permissions to create instances.
B.The account has reached the EC2 instance limit for the selected instance type in the region.
C.The VPC subnet does not have enough available IP addresses.
D.There is an instance that is not passing health checks, preventing new instances.
AnswerB

The account has reached the EC2 instance limit for the selected instance type in the region, which is the only option that matches the error 'You have requested more instances than your current EC2 instance limit allows'. Auto Scaling attempts to call RunInstances, but EC2 rejects it with an InstanceLimitExceeded error, causing the scaling activity to fail. This is a service quota that applies per instance type per region, and the fix is to request a quota increase via the Service Quotas console.

Why this answer

The error message 'Your quota allows for 0 more running instance(s)' directly indicates that the AWS account has reached its EC2 instance limit for the specific instance type in the region. Auto Scaling groups cannot launch instances beyond the service quota, regardless of the desired capacity. This is a common issue when the default limit (e.g., 5 or 20 instances per instance family) has been exhausted.

Exam trap

The trap here is that candidates confuse IAM permissions with service quotas, or assume a VPC subnet IP shortage is the cause, when the explicit quota error message is the definitive clue.

How to eliminate wrong answers

Option A is wrong because IAM permissions affect the ability to call EC2 APIs (e.g., RunInstances), but the error message explicitly cites a quota issue, not an authorization failure (which would return 'UnauthorizedOperation' or 'AccessDenied'). Option C is wrong because insufficient IP addresses in the subnet would produce an error like 'InsufficientFreeAddressesInSubnet' or 'NoFreeAddressesInSubnet', not a quota limit message. Option D is wrong because an instance failing health checks does not prevent new instances from launching; the ASG would still attempt to launch replacements, and the error would relate to health check failures, not a quota limit.

474
MCQmedium

A company uses AWS Elastic Beanstalk to deploy a web application. The environment is running behind an Application Load Balancer. The DevOps team notices that during deployments, the new application version fails health checks and the deployment rolls back. The team wants to reduce deployment time while maintaining safety. Which configuration change should the engineer recommend?

A.Increase the number of EC2 instances in the environment.
B.Use the all-at-once deployment policy.
C.Change the deployment policy to immutable.
D.Increase the rolling update batch size to 100%.
AnswerC

Changing the deployment policy to immutable launches a fully new Auto Scaling group with the new application version while the old group continues serving traffic. Once the new instances pass health checks, Elastic Beanstalk swaps the environment's capacity to the new group, providing both safety and minimal downtime. This is the correct choice because it verifies the deployment before cutting over, unlike rolling or all-at-once strategies that update existing instances in place.

Why this answer

The immutable deployment policy in AWS Elastic Beanstalk creates a new Auto Scaling group with fresh instances running the new application version, performs health checks, and only then swaps the environment's target group to route traffic to the new fleet. This avoids the rolling update's problem of health check failures on existing instances (which trigger a rollback) and ensures the entire new environment is validated before any traffic is shifted. While immutable deployments can take longer than a successful rolling update because they provision a full new fleet, they reduce overall deployment time in this scenario by preventing repeated failed deployments and rollbacks, and they maintain safety by not exposing partially updated instances.

Exam trap

The trap is assuming that increasing batch size or instance count accelerates a rolling update, but the real issue is failed health checks during the rolling update causing rollbacks. Immutable deployments bypass this by launching and validating a completely new set of instances before switching traffic, rather than relying on modifications to the existing fleet.

How to eliminate wrong answers

Option A is wrong because increasing the number of EC2 instances does not change the deployment mechanism; it only adds capacity, which does not prevent health check failures during a rolling update or reduce deployment time. Option B is wrong because the all-at-once deployment policy deploys the new version to all instances simultaneously, which would cause immediate health check failures on all instances and a full rollback, increasing downtime and not reducing deployment time safely. Option D is wrong because increasing the rolling update batch size to 100% is effectively the same as all-at-once, which would cause all instances to be updated at once, leading to health check failures and a full rollback, not a safe or faster deployment.

475
MCQhard

A team manages a large fleet of EC2 instances using AWS Systems Manager. They want to enforce a consistent configuration across all instances, including installed software packages, firewall rules, and user accounts. The team also needs to audit configuration changes and remediate drift automatically. Which AWS service should the team use?

A.AWS OpsWorks for Chef Automate
B.AWS Systems Manager State Manager
C.AWS Systems Manager Run Command
D.AWS Config
AnswerB

AWS Systems Manager State Manager is the correct choice because it lets you define a desired configuration state (such as specific software packages, user accounts, or agent settings) and automatically apply and maintain that state on your EC2 fleet. It uses associations that run on a schedule, detect drift from the defined state, and reapply the configuration whenever needed. Unlike ad-hoc tools, State Manager continuously enforces the desired state across instances with built-in rate controls and error handling, making it ideal for managing large fleets.

Why this answer

AWS Systems Manager State Manager is the correct choice because it is designed to enforce a consistent configuration across EC2 instances by defining and applying desired state configurations (DSCs). It can manage software packages, firewall rules, and user accounts, and it automatically remediates drift by re-applying the desired state on a schedule. This directly meets the requirement for configuration enforcement, auditing, and automated drift remediation.

Exam trap

The trap here is confusing AWS Config (which only audits and detects drift) with State Manager (which enforces and remediates drift), leading candidates to choose Config because they focus on the auditing requirement without realizing it lacks enforcement capabilities.

How to eliminate wrong answers

Option A is wrong because AWS OpsWorks for Chef Automate is a configuration management service that uses Chef cookbooks, but it requires managing a Chef server and does not natively integrate with Systems Manager for drift remediation or auditing without additional setup. Option C is wrong because AWS Systems Manager Run Command is designed for ad-hoc, one-time command execution across instances, not for enforcing ongoing desired state configurations or automatically remediating drift. Option D is wrong because AWS Config is a service for auditing resource configurations and tracking changes, but it does not enforce configurations or remediate drift; it only detects non-compliance and can trigger remediation actions via other services like Systems Manager Automation.

476
MCQmedium

A company is using Amazon CloudWatch Logs Insights to analyze application logs. The DevOps team needs to create a metric filter that counts occurrences of the word 'ERROR' in the log events. Which CloudWatch Logs Insights query should be used to test the metric filter?

A.fields @timestamp, @message | stats count() by bin(5m)
B.fields @timestamp, @message | filter @message like /ERROR/
C.fields @timestamp, @message | parse @message '[*] *' as @severity, @log
D.fields @timestamp, @message | sort @timestamp desc
AnswerB

fields @timestamp, @message | filter @message like /ERROR/ applies a regular expression filter to the raw message text, returning only the individual log events that contain the substring 'ERROR'. This directly mirrors the behavior of the CloudWatch Logs metric filter pattern "ERROR", allowing you to see the exact events that would increment the metric. It is the appropriate query for testing whether the pattern matches the intended production logs before creating or updating the metric filter.

Why this answer

To test a metric filter that counts occurrences of the word 'ERROR' in log events, you need a CloudWatch Logs Insights query that filters log events containing 'ERROR'. The query `fields @timestamp, @message | filter @message like /ERROR/` does exactly that: it selects the timestamp and message fields and filters for messages that match the regular expression /ERROR/. This allows you to verify that the filter pattern will correctly identify the relevant log events.

Exam trap

The trap is choosing a query that counts or sorts logs without filtering for the specific pattern. Candidates might think that any query that includes @message is sufficient, but the key is to filter for 'ERROR'. Also, some might forget that the filter must match the exact pattern used in the metric filter, which often includes a regex.

How to eliminate wrong answers

Option A is wrong because it uses `stats count() by bin(5m)`, which counts all log events in 5-minute bins without filtering for 'ERROR'; it does not test the metric filter's pattern. Option C is wrong because it parses the message into severity and log fields but does not filter for 'ERROR'; it would return all logs, not just those containing 'ERROR'. Option D is wrong because it sorts logs by timestamp descending but does not filter for 'ERROR'; it would return all logs in reverse chronological order.

477
MCQmedium

A company uses Amazon CloudFront to distribute content globally. Users in some regions report slow load times. The DevOps team wants to identify the geographic regions where performance is worst. Which tool should they use?

A.Amazon CloudWatch Metrics for CloudFront
B.CloudFront access logs in S3
C.Amazon Route 53 latency records
D.CloudFront reports in the AWS Management Console
AnswerD

CloudFront reports in the AWS Management Console include a Geo Distribution report and viewer reports that break down requests, bytes served, and other metrics by country and by edge location. These built-in reports are pre-aggregated by AWS and displayed directly in the console, giving a quick geographic view of traffic patterns and performance without needing additional infrastructure or custom queries. They are the intended way to answer global distribution questions out of the box.

Why this answer

CloudFront reports in the AWS Management Console provide performance metrics (e.g., total requests, error rates, latency) broken down by geographic region, enabling the team to identify regions with worst performance. Option A is wrong: CloudWatch metrics for CloudFront are aggregated per distribution, not per region. Option B is wrong: CloudFront access logs in S3 record individual requests but do not aggregate performance by region.

Option C is wrong: Amazon Route 53 latency records are used for DNS-based routing decisions, not for analyzing CloudFront performance.

478
MCQhard

A company runs a high-traffic web application on a fleet of EC2 instances behind an Application Load Balancer (ALB) with Auto Scaling. The application uses an Amazon RDS for PostgreSQL database. Recently, during a traffic spike, the application became unresponsive. Investigation revealed that the database CPU utilization reached 100%, causing queries to timeout. The Auto Scaling group added more EC2 instances, which only increased the load on the database. The DevOps team needs to implement a solution that prevents the database from being overwhelmed during traffic spikes while maintaining application availability. The solution must be cost-effective and require minimal changes to the application code. Which solution should the DevOps team implement?

A.Implement read replicas for the RDS database and modify the application to use read replicas for read queries.
B.Increase the instance size of the RDS database to a larger instance type to handle more connections.
C.Use Amazon RDS Proxy between the application and the database to pool and reuse connections.
D.Configure Auto Scaling to launch EC2 instances based on a custom metric that tracks database CPU utilization, and throttle the number of instances.
AnswerC

Amazon RDS Proxy presents a single endpoint to the application while pooling and reusing database connections on the backend, dramatically cutting the CPU and memory load caused by connection handling. It is fully managed and transparent, and during Multi-AZ failovers it can keep connections warm for faster recovery. For a high-traffic web fleet, this directly mitigates CPU exhaustion due to connection volume, making it the most efficient and cost-effective option.

Why this answer

RDS Proxy manages database connections efficiently, reducing the number of connections and CPU overhead. It also provides connection pooling, which helps handle spikes without overwhelming the database.

479
Multi-Selecteasy

A company is designing a CI/CD pipeline using AWS CodePipeline, CodeBuild, and CodeDeploy. They need to ensure that the pipeline can deploy to multiple environments (dev, test, prod) with manual approval gates. Which TWO actions should they take? (Choose TWO.)

Select 2 answers
A.Create separate stages in the pipeline for dev, test, and prod
B.Configure a single pipeline with multiple branches in the source stage
C.Use CodeDeploy deployment groups to represent each environment
D.Add a manual approval stage before each environment deployment
E.Use CodeBuild batch builds to manage environment promotion
AnswersA, D

Separate stages in CodePipeline provide a sequential execution model where each stage contains a set of actions that deploy or test an environment. By defining dev, test, and prod as distinct stages, you enforce an ordered promotion flow, allow artifacts to progress only after each stage succeeds, and can insert approval actions between them. This is the fundamental way to model environment promotion in CodePipeline.

Why this answer

AWS CodePipeline allows you to define separate stages for each environment (dev, test, prod) within a single pipeline. This enables sequential or parallel deployments with clear separation of concerns, and each stage can have its own actions, such as deployment to a specific CodeDeploy application or environment. Option D is correct because you can add a manual approval action as a stage gate before each environment deployment, ensuring that a human reviewer must explicitly approve the promotion before the pipeline proceeds to the next environment.

Exam trap

The trap here is that candidates often confuse CodeDeploy deployment groups with environment stages, thinking that a single deployment group can represent an entire environment, when in fact deployment groups are compute targets within an environment and do not provide the stage-level orchestration or manual approval gates that CodePipeline stages offer.

480
MCQhard

A DevOps engineer is designing a CI/CD pipeline for a microservices architecture. The pipeline must deploy to Amazon ECS using blue/green deployments. The team wants to automatically roll back if the new deployment fails health checks. Which combination of AWS services and configurations should the engineer use?

A.Use AWS CodeDeploy with an ECS compute platform, configure a CloudWatch alarm for health checks, and enable automatic rollback.
B.Use AWS CloudFormation with a custom resource to perform blue/green deployment.
C.Use AWS Elastic Beanstalk with blue/green environment swapping.
D.Use AWS CodePipeline with ECS deployment action and manual approval for rollback.
AnswerA

CodeDeploy with the ECS compute platform natively orchestrates blue/green deployments by shifting traffic between the original and replacement ECS task sets using a load balancer target group. You can associate CloudWatch alarms with the deployment and enable automatic rollback so that if health checks or custom metrics fail, CodeDeploy reverts traffic to the original task set without manual intervention. This is the only option that combines native ECS support, health monitoring, and automated rollback.

Why this answer

AWS CodeDeploy with an ECS compute platform natively supports blue/green deployments for Amazon ECS, including automatic rollback triggered by CloudWatch alarms. By configuring a CloudWatch alarm based on ECS service health checks (e.g., ELB target group health), CodeDeploy can automatically revert to the original blue task set if the new green deployment fails, meeting the requirement for zero-touch rollback.

Exam trap

The trap here is that candidates may assume CodePipeline's ECS action supports automatic rollback, but it only provides a deployment action without native health check monitoring or rollback logic, requiring additional custom steps or manual intervention.

How to eliminate wrong answers

Option B is wrong because AWS CloudFormation does not natively support blue/green deployments for ECS; custom resources would require significant custom code and lack built-in health check rollback integration. Option C is wrong because AWS Elastic Beanstalk is a PaaS service for web applications, not designed for microservices on ECS, and its blue/green environment swapping does not integrate with ECS task definitions or service health checks. Option D is wrong because AWS CodePipeline with an ECS deployment action does not support automatic rollback; manual approval for rollback contradicts the requirement for automatic rollback on health check failure.

481
MCQhard

A DevOps team is deploying a multi-tier application on AWS. The application must comply with PCI DSS. Which combination of services should be used to encrypt data in transit between the web tier and the application tier?

A.AWS Certificate Manager (ACM) and Application Load Balancer (ALB)
B.AWS CloudHSM and Classic Load Balancer
C.AWS KMS and VPC Peering
D.AWS WAF and Amazon CloudFront
AnswerA

ACM issues and automatically renews public or private TLS certificates that integrate natively with an ALB's HTTPS listener, allowing the ALB to terminate TLS and encrypt traffic between the client and load balancer. For multi-tier architectures, the ALB can front each layer (e.g., web and application), providing encrypted inter-tier communication without manual certificate deployment or key management. This approach leverages AWS-managed distribution, per-listener policies, and SNI support, directly addressing the requirement for in-transit encryption.

Why this answer

AWS Certificate Manager (ACM) provisions and manages the TLS certificates, and an Application Load Balancer (ALB) terminates TLS and re-encrypts traffic to backend targets, providing encryption in transit between the web tier and the application tier. This combination is the standard AWS pattern for PCI DSS-compliant in-transit encryption on a multi-tier application.

Exam trap

The trap is picking KMS or CloudHSM for 'encryption' — candidates forget that KMS encrypts data at rest and CloudHSM stores keys, while encryption in transit requires TLS, which on AWS means ACM plus a load balancer that terminates and re-encrypts.

How to eliminate wrong answers

Option B is wrong because CloudHSM is a hardware security module for key storage and cryptographic operations, not a load-balancing or TLS-termination service, and Classic Load Balancer is a legacy offering that lacks the modern TLS policy and target-group features needed for tier-to-tier encryption. Option C is wrong because KMS is a key management service and VPC Peering is a network connectivity feature — neither encrypts traffic in transit between tiers. Option D is wrong because AWS WAF is a web application firewall that filters HTTP requests and CloudFront is a CDN — neither provides the TLS termination and re-encryption needed between the web and application tiers.

482
Multi-Selectmedium

A company uses AWS CloudTrail to log API calls in a multi-account environment. The security team wants to be alerted immediately when an IAM user or role performs a specific sensitive action (e.g., DeleteTrail, DeleteDBInstance). Which TWO services can be used together to achieve near real-time alerting? (Choose TWO.)

Select 2 answers
A.CloudWatch Logs metric filters and alarms
B.CloudTrail with CloudWatch Logs integration
C.CloudTrail with Amazon S3 event notifications
D.Amazon Athena and CloudWatch dashboards
E.AWS Config and AWS Lambda
AnswersA, B

CloudWatch Logs metric filters are the correct mechanism for real-time alerting on CloudTrail API activity. A metric filter defines a pattern that matches specific CloudTrail log events, such as `errorCode = "UnauthorizedOperation"`, and continuously increments a custom CloudWatch metric. A CloudWatch alarm can then evaluate that metric over a fixed period (e.g., 5 minutes) and trigger an SNS notification when the threshold is breached, enabling fast, automated responses to suspicious API calls.

Why this answer

CloudTrail logs can be delivered to CloudWatch Logs, where metric filters can be created to match specific API actions (e.g., DeleteTrail, DeleteDBInstance) and trigger CloudWatch alarms for near real-time notification. Option B is correct because CloudTrail integration with CloudWatch Logs is the prerequisite step that enables the log delivery required for metric filters and alarms. Option C is incorrect: CloudTrail with Amazon S3 event notifications is not near real-time; S3 event notifications can have delays and are not designed for immediate alerting on specific API calls.

Option D is incorrect: Amazon Athena is an interactive query service for ad-hoc analysis, not for real-time alerting. Option E is incorrect: AWS Config is used for resource configuration compliance and change tracking, not for real-time alerting on API calls.

483
MCQhard

A DevOps engineer executed the CLI command shown in the exhibit. After creation, the security team requires that the log files be encrypted with a KMS key that is rotated every 90 days. The current key is a customer managed key with automatic rotation enabled set to 365 days. What should the engineer do to meet the requirement?

A.Use the existing key and change the rotation period in KMS
B.Disable automatic rotation and manually rotate the key every 90 days
C.Modify the KMS key to set the rotation period to 90 days
D.Create a new KMS key with automatic rotation set to 90 days and update the trail with the new key
AnswerD

To achieve a 90-day automatic rotation, you must create a new symmetric customer managed key with `--rotation-period-in-days 90` in the `create-key` CLI call and then associate it with the trail using `update-trail --kms-key-id <new-key-arn>`. After updating, CloudTrail will use the new key to encrypt future log files, and the new key's key policy must include CloudTrail's account and the required `kms:GenerateDataKey` and `kms:Decrypt` permissions. Existing log files remain encrypted under the old key, so that key should still be available for decryption.

Why this answer

The requirement is to encrypt log files with a KMS key that rotates every 90 days. The current key rotates every 365 days, and you cannot change the rotation period of an existing customer managed key; you must create a new key with the desired rotation period. Once created, you update the CloudTrail trail to use the new key by specifying the --kms-key-id parameter.

Option D correctly describes this process. Option A is wrong because you cannot change the rotation period of an existing key. Option B is wrong because manually rotating a key does not meet the automatic rotation requirement and is not recommended.

Option C is wrong because you cannot modify the rotation period of an existing key.

484
MCQhard

A company runs a stateful web application on EC2 instances behind a Network Load Balancer (NLB) in a single Availability Zone. The application stores session state locally on the instance. The company wants to achieve high availability across multiple AZs with minimal application changes. What should the DevOps engineer do?

A.Add more AZs and configure the NLB with cross-zone load balancing.
B.Replace the NLB with an ALB and use ElastiCache for session storage.
C.Use a Multi-AZ RDS instance to store session state.
D.Replace the NLB with an ALB and enable sticky sessions (session affinity) using the ALB's cookie.
AnswerD

Replacing the NLB with an ALB and enabling sticky sessions via the ALB's load balancer-generated cookie is the correct minimal-change solution. The ALB inserts a stickiness cookie on the first response, and all subsequent requests from that client are routed to the same EC2 instance, preserving the locally stored session state without any application modifications. This leverages the ALB's Layer 7 capabilities to achieve session affinity while keeping the existing web application code unchanged.

Why this answer

Replacing the NLB with an ALB and enabling sticky sessions (session affinity) using the ALB's cookie allows the stateful web application to maintain session state across multiple AZs without modifying the application code. The ALB generates a cookie (AWSALB) that binds a client's session to a specific target instance, ensuring subsequent requests from the same client are routed to the same EC2 instance. This achieves high availability across AZs with minimal changes, as the application continues to store session state locally on the instance.

Exam trap

The trap here is that candidates often assume cross-zone load balancing or adding more AZs inherently solves high availability for stateful applications, but they overlook that session affinity is required to keep a client's requests directed to the same instance when session state is stored locally.

How to eliminate wrong answers

Option A is wrong because adding more AZs and configuring cross-zone load balancing with an NLB does not solve the session state problem; the NLB distributes traffic across instances without session affinity, so a client's requests may be routed to different instances in different AZs, breaking the locally stored session. Option B is wrong because replacing the NLB with an ALB and using ElastiCache for session storage requires application code changes to read/write session data to ElastiCache, which contradicts the requirement for minimal application changes. Option C is wrong because using a Multi-AZ RDS instance for session storage also requires significant application code changes to store and retrieve session data from the database, and it introduces unnecessary complexity and latency for session management.

485
Multi-Selecthard

A company's application uses Amazon DynamoDB as its primary data store. The application experiences occasional throttling errors during traffic spikes. The DevOps team needs to implement a solution that ensures consistent performance without manual intervention. Which TWO actions should the team take? (Choose TWO.)

Select 2 answers
A.Use eventually consistent reads for all queries.
B.Move the data to Amazon RDS with read replicas.
C.Implement DynamoDB Accelerator (DAX) to cache read requests.
D.Enable DynamoDB Auto Scaling for read and write capacity.
E.Switch DynamoDB to On-Demand capacity mode.
AnswersC, D

DynamoDB Accelerator (DAX) is an in-memory cache that sits in front of a DynamoDB table, serving read-heavy workloads with microsecond latency while absorbing a large fraction of read requests. By caching frequently accessed items (including strongly consistent reads when DAX is enabled), DAX reduces the number of read requests that actually reach DynamoDB, thereby reducing the table's consumed read capacity and preventing read throttling. It is a native, fully managed solution specifically designed for this scenario, preserving the DynamoDB API and requiring no application rewrite beyond adding a DAX client endpoint.

Why this answer

To handle occasional throttling during traffic spikes without manual intervention, the team should use DynamoDB Accelerator (DAX) to cache read requests, reducing read load on the table, and enable DynamoDB Auto Scaling to automatically adjust read and write capacity based on traffic patterns. DAX absorbs spikey read traffic, while Auto Scaling ensures sufficient capacity for writes and uncached reads, together providing consistent performance without manual scaling. Option E (On-Demand) also handles spikes automatically but can be costlier for predictable workloads; the combination of DAX and Auto Scaling is often more cost-effective for read-heavy applications.

Exam trap

The trap here is that candidates may think On-Demand capacity mode (Option E) is the only way to handle spikes without manual intervention, but it ignores the cost implications and the fact that DAX plus Auto Scaling provides a more balanced and cost-effective solution for read-heavy workloads.

486
MCQmedium

A DevOps engineer manages a DynamoDB table that serves a high-traffic ordering API. During a flash sale, ProvisionedThroughputExceededException errors spike and the on-call engineer must reduce customer impact quickly. The table currently uses provisioned capacity and the workload is expected to surge unpredictably for the next several hours. Which action should the engineer take FIRST to mitigate the incident?

A.Enable DynamoDB Accelerator (DAX) in front of the table to absorb the surge.
B.Enable DynamoDB on-demand capacity mode for the table to absorb the unpredictable traffic without throughput errors.
C.Increase the table's provisioned read and write capacity units manually to a value matching the observed peak.
D.Create a global secondary index with higher capacity and redirect all API reads to it.
AnswerB

Switching the table to on-demand capacity mode removes the provisioned throughput ceiling and instantly accommodates unpredictable spikes, which is exactly the flash-sale pattern here. It is the fastest mitigation because the table continues serving reads and writes during the switch, so the ProvisionedThroughputExceededException errors stop without capacity math or client changes.

Why this answer

The scenario describes unpredictable burst traffic against a provisioned DynamoDB table producing throttling errors. Switching to on-demand capacity mode immediately removes the provisioned throughput limit and lets the table scale with the surge while still serving requests, making it the fastest and most reliable mitigation. Approaches that tune provisioned values or cache reads do not remove the write bottleneck during an unpredictable spike.

Exam trap

The trap here is assuming that caching with DAX or adding an index solves throughput throttling, when throttling in a write-heavy burst is caused by the base table's provisioned capacity ceiling.

487
MCQmedium

An organization uses OpsWorks to manage application stacks. They notice that custom cookbooks are not being executed during the lifecycle events. What is the most likely cause?

A.The layer's IAM role does not have permissions to execute the cookbook
B.The custom cookbook repository URL is misconfigured or inaccessible
C.The cookbook is not configured with CodeDeploy
D.The cookbook uses a Chef version that is not supported by OpsWorks
AnswerB

OpsWorks Stacks downloads custom cookbooks from the source repository during the setup phase and then runs the recipes defined for each lifecycle event. If the repository URL is malformed, points to a private repo with invalid credentials, or the endpoint is unreachable, the agent cannot retrieve the cookbook and therefore cannot execute any recipes. The failure may appear as 'non-execution' because the instance comes online but nothing from the cookbook is applied, and the error is only visible if you inspect the OpsWorks agent logs or stack activity.

Why this answer

Custom cookbooks in AWS OpsWorks are fetched from a repository (e.g., Git, S3, HTTP) during lifecycle events. If the repository URL is misconfigured (e.g., wrong branch, invalid path) or inaccessible (e.g., private repo without proper SSH keys or S3 bucket permissions), OpsWorks cannot retrieve the cookbooks, causing them to not execute. This is the most common cause of cookbook execution failures.

Exam trap

The trap here is that candidates often confuse IAM permissions with repository access, assuming the layer's IAM role controls cookbook retrieval, when in fact OpsWorks uses separate SSH keys or S3 bucket policies for repository access.

How to eliminate wrong answers

Option A is wrong because the layer's IAM role is used for AWS API calls (e.g., EC2, CloudWatch), not for executing Chef cookbooks; cookbook execution is handled by the Chef client locally. Option C is wrong because CodeDeploy is a separate AWS service for application deployments and is not involved in OpsWorks Chef cookbook execution; OpsWorks uses its own lifecycle event system. Option D is wrong because OpsWorks supports Chef 11.10, 12, and 12.2; if a cookbook uses an unsupported version, the error would occur during Chef run, not silently skip execution, and OpsWorks would log a version mismatch error.

488
MCQeasy

A company wants to securely store database credentials used by an application running on Amazon EC2. The credentials should be automatically rotated every 90 days. Which AWS service should be used?

A.AWS IAM
B.AWS KMS
C.AWS Systems Manager Parameter Store
D.AWS Secrets Manager
AnswerD

AWS Secrets Manager is purpose-built for storing secrets like database credentials and natively supports automatic rotation through an integrated Lambda rotation function. It manages secret versions with AWSCURRENT and AWSPREVIOUS labels, allowing applications to reliably fetch rotated credentials without downtime. It also provides fine-grained access via IAM and resource policies, making it the correct service when the requirement is both secure storage and automatic rotation of database credentials.

Why this answer

AWS Secrets Manager is designed to securely store and manage secrets such as database credentials, and it provides built-in automatic rotation every 90 days (or custom intervals) using Lambda functions. It integrates natively with Amazon RDS, Redshift, and DocumentDB for rotation. IAM and KMS do not store secrets, and Parameter Store does not offer automatic rotation natively.

Exam trap

The trap is confusing Parameter Store with Secrets Manager; candidates may think Parameter Store can rotate secrets automatically, but it does not—rotation is a key differentiator of Secrets Manager.

How to eliminate wrong answers

Option A is wrong because AWS IAM is for identity and access management, not for storing database credentials. Option B is wrong because AWS KMS is a key management service for encryption keys, not for storing secrets. Option C is wrong because AWS Systems Manager Parameter Store can store secrets, but it does not provide automatic rotation; rotation must be implemented manually or via custom automation.

489
MCQhard

A company is implementing a disaster recovery strategy for its Amazon Aurora MySQL database. The primary database is in us-west-2. The company requires an RPO of less than 1 minute and an RTO of less than 5 minutes. Which solution meets these requirements?

A.Create a cross-Region read replica in the secondary Region and promote it during failover.
B.Use automated backups and restore to a new DB instance in the secondary Region.
C.Use Amazon Aurora Global Database with a secondary Region cluster.
D.Take manual snapshots of the DB instance and copy them to the secondary Region every hour.
AnswerC

Aurora Global Database replicates data from the primary Region to a secondary cluster using a dedicated storage-based replication channel with typical latency under one second, ensuring an RPO of under one minute. Failover can be initiated either manually or automatically, and a promoted secondary cluster becomes available in minutes without the need to restore from a backup or apply transaction logs. This is the only option that inherently satisfies both the 1-minute RPO and a recovery time objective measured in minutes.

Why this answer

Amazon Aurora Global Database is designed for low-latency cross-Region replication with a typical RPO of 1 second and RTO of 1 minute or less, meeting the <1 minute RPO and <5 minute RTO requirements. It uses a dedicated storage-level replication channel that keeps the secondary cluster fully synchronized without impacting primary performance, and failover involves promoting the secondary cluster to primary in under a minute.

Exam trap

The trap here is that candidates confuse a cross-Region read replica (Option A) with Aurora Global Database, assuming both provide similar failover speed, but the read replica's promotion process is slower and less reliable for meeting strict RTO/RPO targets.

How to eliminate wrong answers

Option A is wrong because a cross-Region read replica for Aurora MySQL uses asynchronous replication with a typical RPO of several seconds to minutes, but the promotion process can take longer than 5 minutes due to the need to apply remaining redo logs and reconfigure endpoints, failing the RTO requirement. Option B is wrong because automated backups are taken once per day (default retention of 1-35 days) and restoring to a new instance in a secondary Region requires copying the backup across Regions, which can take hours and far exceeds both the RPO and RTO limits. Option D is wrong because manual snapshots taken every hour provide an RPO of up to 60 minutes, which violates the <1 minute RPO requirement, and restoring from a snapshot in a secondary Region also takes significantly longer than 5 minutes.

490
Multi-Selectmedium

A company runs a stateful web application on EC2 instances that store session data locally. They want to migrate to a stateless architecture for better resilience. Which TWO actions should they take?

Select 2 answers
A.Use Amazon CloudFront to cache session data at the edge.
B.Use Amazon DynamoDB to store session data.
C.Use Amazon S3 to store session data as objects.
D.Use ElastiCache for Redis to store session data externally.
E.Use Amazon EFS to store session data as files.
AnswersB, D

Amazon DynamoDB is a fully managed NoSQL key-value database that delivers single-digit millisecond read/write performance at any scale, making it an excellent external session store. Its fine-grained access control, encryption, backup, and TTL support for automatic item expiration align well with session lifecycle needs. Because DynamoDB is serverless and horizontally scalable, EC2 instances can share session state without statefulness, enabling the application tier to scale freely behind a load balancer.

Why this answer

DynamoDB provides a fully managed, low-latency, highly available NoSQL database that is ideal for storing session state externally. By moving session data to DynamoDB, the EC2 instances become stateless, allowing any instance to handle any request without relying on local storage, which improves resilience and scalability.

Exam trap

The trap here is that candidates may confuse stateless session storage with caching or file storage, incorrectly choosing S3 or EFS because they are persistent, while overlooking the need for low-latency, high-throughput, and consistent access that only DynamoDB or ElastiCache can provide.

491
MCQeasy

A company runs a serverless application using AWS Lambda functions behind an Amazon API Gateway. The application processes user uploads stored in an S3 bucket. The Lambda function writes results to a DynamoDB table. Recently, the function started timing out when processing large files. What should the DevOps engineer do to improve resilience for large file processing?

A.Increase the Lambda function memory to improve CPU performance.
B.Use S3 event notifications to trigger an AWS Step Functions workflow that processes the file asynchronously.
C.Increase the Lambda function timeout to the maximum 15 minutes.
D.Add Amazon ElastiCache to cache processed results and reduce Lambda execution time.
AnswerB

S3 event notifications triggering Step Functions decouple processing from the API request, letting large files be handled asynchronously with retries and state tracking. This removes the Lambda timeout constraint imposed by synchronous invocation behind API Gateway.

Why this answer

Using S3 event notifications to trigger an AWS Step Functions workflow enables asynchronous processing of large files, decoupling the upload from the processing and avoiding Lambda's timeout limits. Option A (increasing memory) may improve CPU performance but does not address the timeout issue for large files. Option C (increasing timeout) can extend up to 15 minutes, but large files may still exceed this limit and it does not provide a resilient architecture.

Option D (ElastiCache) caches processed results but does not solve the initial timeout problem; it is irrelevant to the large file processing issue.

492
MCQeasy

A DevOps engineer is using AWS CodeBuild to build a container image and push it to Amazon ECR. The buildspec.yml file includes a post_build phase that runs `docker push`. The build fails with an error indicating that the Docker daemon is not available. The CodeBuild project uses the `aws/codebuild/standard:5.0` image and has privileged mode disabled. Which action should the engineer take to resolve the issue?

A.Enable privileged mode in the CodeBuild project configuration.
B.Add a pre_build phase to start the Docker daemon manually using `sudo service docker start`.
C.Use the `docker buildx` command instead of `docker build` to avoid needing privileged mode.
D.Switch to a custom build image that includes Docker pre-installed and configured.
AnswerA

Docker requires privileged mode to run inside a container. AWS CodeBuild projects must have privileged mode enabled to build Docker images. Enabling it allows the Docker daemon to run, resolving the error. This is a common configuration for building container images in CodeBuild and is the minimal change needed to fix the issue.

Why this answer

Docker requires privileged mode in CodeBuild to run the Docker daemon. Enabling privileged mode in the project configuration allows Docker commands to execute, resolving the error. Other options do not address the underlying permission issue and would not allow Docker to run.

Exam trap

The trap here is assuming that Docker can run in CodeBuild without privileged mode, leading to attempts to start the daemon manually or change images.

493
MCQhard

A company runs a multi-region application on Amazon EC2 instances across us-east-1 and eu-west-1. The application uses an Amazon Aurora global database for writes in us-east-1 and reads in eu-west-1. The DevOps team wants to monitor the replication lag between the primary and secondary regions. They have set up a CloudWatch alarm on the AuroraReplicaLag metric in both regions. However, they notice that the alarm in eu-west-1 sometimes triggers false positives when the lag spikes briefly but then recovers. The team wants to reduce false alarms while still being alerted to sustained high lag that could impact read replicas. The team is already using a standard CloudWatch alarm with a period of 1 minute and evaluation periods of 1. What should the team change to reduce false positives?

A.Increase the alarm threshold to a higher value, such as 10 seconds.
B.Reduce the metric period to 30 seconds to get more granular data.
C.Increase the number of evaluation periods to 3, so the alarm triggers only if the lag is high for 3 consecutive minutes.
D.Create a composite alarm that triggers when both the AuroraReplicaLag and CPUUtilization metrics are high.
AnswerC

Increasing the number of evaluation periods to 3 with a 1-minute period means the alarm must observe the replication lag exceeding the threshold for three consecutive datapoints (three consecutive minutes) before entering ALARM state. This creates a temporal smoothing effect that filters out brief, self-correcting spikes in AuroraReplicaLag, which are often caused by momentary write bursts or replica catch-up delays. CloudWatch evaluates all three most recent datapoints when determining alarm state, so sustained high lag triggers the alarm, while isolated outliers do not, directly reducing false positives without changing the sensitivity to genuine long-duration lag.

Why this answer

The false positives occur because a single 1-minute evaluation period triggers on brief, transient lag spikes. Increasing the number of evaluation periods to 3 means the alarm fires only when the lag exceeds the threshold for 3 consecutive 1-minute periods, filtering out short spikes while still catching sustained high lag. This is the standard CloudWatch approach to reduce false positives without losing sensitivity to real issues.

Exam trap

DOP-C02 often tests the difference between changing a threshold (sensitivity) and changing evaluation periods/datapoints (transient filtering), and candidates frequently pick threshold changes or composite alarms when the real fix is sustained-breach evaluation.

How to eliminate wrong answers

Option A is wrong because raising the threshold to 10 seconds may reduce some false positives but also risks missing real sustained lag that is still impactful — it changes sensitivity rather than filtering transients. Option B is wrong because reducing the period to 30 seconds increases granularity and would make the alarm more sensitive to spikes, worsening false positives. Option D is wrong because a composite alarm combining AuroraReplicaLag with CPUUtilization adds an unrelated condition and does not address the transient-spike problem; it could also suppress real lag alerts when CPU is normal.

494
MCQhard

A company runs a three-tier application on Amazon EC2 behind an Application Load Balancer. During an incident, users report intermittent HTTP 502 errors, and the engineer finds that some targets are failing health checks and being removed and re-added repeatedly. The application writes large log files to the instance store and the engineer suspects the health check configuration. Application startup takes about 90 seconds. Which configuration change should the engineer make to resolve the flapping targets?

A.Enable cross-zone load balancing on the Application Load Balancer to spread traffic more evenly.
B.Set the health check interval to 5 seconds and the unhealthy threshold to 1 so failures are detected faster.
C.Increase the health check interval and unhealthy threshold, and use a longer health check grace period or slow start so targets are not removed during startup.
D.Change the target group to use TCP health checks instead of HTTP health checks.
AnswerC

Flapping occurs when targets are marked unhealthy during normal startup or heavy I/O. Lengthening the interval and unhealthy threshold tolerates transient failures, while a health check grace period or slow start keeps new or busy targets in service until they are ready, which directly stops the repeated removal and re-addition causing 502s.

Why this answer

The targets flap because health checks fail during startup or heavy disk I/O, so the load balancer repeatedly removes and re-adds them, producing intermittent 502 responses. Extending the health check interval and unhealthy threshold and applying a grace period or slow start keeps targets in service until they are genuinely ready, which stops the flapping and the resulting errors.

Exam trap

The trap here is treating flapping targets as a load-balancing distribution problem and enabling cross-zone balancing, when the real cause is health check timing relative to application startup.

495
MCQhard

A company runs a critical application on Amazon EC2 instances behind an Application Load Balancer. The application is deployed using AWS CodeDeploy with an in-place deployment configuration. During a recent deployment, the deployment failed because the new application version caused a health check failure, and CodeDeploy did not automatically roll back. What should the engineer do to ensure automatic rollback on health check failure?

A.Set up an EC2 instance lifecycle hook to trigger a rollback script when the instance enters a pending state
B.Configure an Amazon SQS queue to monitor health checks and invoke a rollback Lambda function
C.Enable automatic rollback in the CodeDeploy deployment group and set up a CloudWatch alarm for the ALB health check
D.Modify the Auto Scaling group to replace unhealthy instances automatically
AnswerC

Enabling automatic rollback in the deployment group, combined with a CloudWatch alarm on the ALB health check, lets CodeDeploy detect the failed health check and revert to the last known-good revision automatically, satisfying the requirement for rollback without manual intervention.

Why this answer

CodeDeploy can automatically roll back a deployment when a CloudWatch alarm, such as one monitoring ALB health check failures, enters the ALARM state. By enabling automatic rollback in the deployment group and associating the CloudWatch alarm, the deployment will revert to the previous version as soon as the health check fails, without manual intervention.

Exam trap

The trap here is that candidates often assume Auto Scaling group health checks or lifecycle hooks can handle deployment rollbacks, but they operate at the instance level and do not revert application code, whereas CodeDeploy's native automatic rollback with CloudWatch alarms is the correct, integrated solution.

How to eliminate wrong answers

Option A is wrong because EC2 instance lifecycle hooks are designed to pause an instance during launch or termination for custom actions, not to trigger rollbacks based on health check failures; they operate at the instance lifecycle level, not the deployment level. Option B is wrong because SQS queues are message brokers and cannot directly monitor health checks or invoke rollbacks; while a Lambda function could be triggered, this approach adds unnecessary complexity and is not the native, supported mechanism for automatic rollback in CodeDeploy. Option D is wrong because Auto Scaling group health checks replace unhealthy instances but do not revert the application version; they would launch a new instance with the same failing code, perpetuating the failure rather than rolling back the deployment.

496
MCQhard

A team uses AWS CodePipeline to deploy a serverless application using AWS SAM. The pipeline includes a build stage that runs 'sam build' and a deploy stage that runs 'sam deploy'. The deployment fails with an error: 'The security token included in the request is invalid.' What is the MOST likely cause?

A.The build stage did not produce the correct output artifact.
B.The SAM template has a syntax error.
C.The IAM role used in the deploy stage does not have permission to assume the CloudFormation execution role.
D.The 'sam deploy' command is missing the '--capabilities' parameter.
AnswerC

CodePipeline's CloudFormation deploy action uses a service role to issue `sts:AssumeRole` for the CloudFormation execution role specified in the pipeline configuration. If the service role's policy lacks `sts:AssumeRole` permission on that execution role (or the execution role's trust policy does not allow the service role), the STS call fails with 'AccessDenied' or an invalid-token indication because the temporary credentials cannot resolve the requested role. This precisely matches the reported error, making it the root cause.

Why this answer

The error 'The security token included in the request is invalid' typically occurs when the IAM role used by CodePipeline in the deploy stage lacks the necessary trust relationship or permissions to assume the CloudFormation execution role. In AWS SAM deployments via CodePipeline, the deploy action uses a specified IAM role to call CloudFormation, and if that role cannot assume the CloudFormation service role (or the CloudFormation role itself is misconfigured), the security token becomes invalid. This is a common misconfiguration when the pipeline's IAM role does not include the 'sts:AssumeRole' permission for the CloudFormation execution role ARN.

Exam trap

The trap here is that candidates often confuse IAM permission errors with template syntax or missing parameters, but the specific 'security token invalid' error points directly to an STS trust or assumption failure, not to CloudFormation validation or artifact issues.

How to eliminate wrong answers

Option A is wrong because the build stage not producing the correct output artifact would cause a different error, such as 'Artifact not found' or a missing file error, not an invalid security token. Option B is wrong because a SAM template syntax error would result in a CloudFormation validation error (e.g., 'Template format error') or a 'sam build' failure, not a security token error. Option D is wrong because missing the '--capabilities' parameter would cause a CloudFormation error like 'Requires capabilities: [CAPABILITY_IAM]', not an invalid security token error.

497
MCQeasy

A company stores application logs in Amazon CloudWatch Logs. During an incident, an engineer needs to search across multiple log groups for a specific request ID from the last hour and then preserve the findings for a post-incident review. Which approach meets both needs with the least operational effort?

A.Create a CloudWatch Logs subscription filter that streams matching events to AWS Lambda for processing.
B.Enable CloudTrail logging for the log groups and search the trail for the request ID.
C.Use CloudWatch Logs Insights to query the log groups for the request ID and save the query results for the post-incident review.
D.Export all log groups to Amazon S3 and use Amazon Athena to search for the request ID.
AnswerC

CloudWatch Logs Insights queries multiple log groups at once with a time range and supports saving or exporting results, directly matching the search and preservation requirements. It requires no infrastructure, so it is the lowest-effort option for finding the request ID and retaining the findings for review.

Why this answer

CloudWatch Logs Insights is designed to run interactive queries across one or more log groups over a specified time range, which fits searching the last hour of logs for a request ID. Because it also lets you save or export query results, it satisfies the need to preserve findings for a post-incident review with minimal setup and no additional infrastructure.

Exam trap

The trap here is reaching for S3 export and Athena for a quick, time-bounded log search, when CloudWatch Logs Insights already queries multiple log groups directly and can retain the results.

498
MCQeasy

A company uses AWS CodeDeploy to deploy a new version of an application to EC2 instances. They want to minimize downtime and roll back quickly if the deployment fails. Which deployment type should they use?

A.Canary deployment
B.Linear deployment
C.Blue/green deployment
D.In-place deployment
AnswerC

Blue/green deployment in AWS CodeDeploy is the only deployment type that provisions a new, separate environment (green) alongside the existing one (blue), then shifts production traffic to the green environment. This architecture enables instant rollback by simply switching traffic back to the blue environment if the deployment fails or exhibits issues, with no need to reinstall or reconfigure instances. CodeDeploy manages this through lifecycle hooks like AllowTraffic and redirecting traffic via the load balancer, providing the fastest and most reliable rollback path.

Why this answer

Blue/green deployment creates two separate environments (blue and green) and shifts traffic from the old to the new after testing. This minimizes downtime because traffic is switched instantly, and rollback is achieved by reverting traffic to the original environment. Option A (Canary) is a traffic shifting pattern used within blue/green deployments, not a standalone deployment type that offers immediate rollback.

Option B (Linear) is also a traffic shifting pattern for blue/green. Option D (In-place) updates existing instances, causing downtime during deployment and requiring a manual rollback process.

499
MCQeasy

A company hosts a static website on Amazon S3 with CloudFront as the CDN. Users report that they see an old version of the website even after the DevOps team updated the S3 objects. The team verified that the new objects are in the S3 bucket and are publicly accessible. The CloudFront distribution has a default TTL of 24 hours. To immediately serve the new content to users, the team needs to invalidate the CloudFront cache. Which of the following is the CORRECT approach to achieve this with minimal impact?

A.Create a CloudFront invalidation request for the path '/*'.
B.Change the CloudFront origin path to point to a new S3 bucket.
C.Update the CloudFront distribution's default TTL to 0 and wait for the changes to propagate.
D.Delete the S3 objects and re-upload them with different names.
AnswerA

An invalidation request for the '/*' path removes all objects from CloudFront's edge caches across every region, which forces the distribution to return to the S3 origin on the next request and fetch the updated content. This is the standard, immediate method for clearing cached content when you need to publish new website changes, and it does not require changing URLs or reconfiguring any origin settings.

Why this answer

A CloudFront invalidation for the path '/*' tells all edge locations to stop serving cached objects matching that pattern and fetch fresh copies from the S3 origin on the next request. This is the standard, immediate way to purge stale content without changing the distribution configuration or object keys. It has minimal impact because it only affects cached objects and does not require re-uploading or renaming anything.

Exam trap

The trap is confusing TTL changes with cache invalidation — candidates pick 'set TTL to 0' thinking it purges existing cached objects, but TTL changes only affect future caching, not objects already stored at edge locations.

How to eliminate wrong answers

Option B is wrong because changing the origin path to a new S3 bucket would require the new content to actually exist in that bucket and would break existing URLs; it is a disruptive configuration change, not a cache purge. Option C is wrong because setting the default TTL to 0 only affects future cache behavior — objects already cached at edge locations will still be served until their existing TTL expires, so it does not immediately serve new content. Option D is wrong because deleting and re-uploading objects with different names changes the object keys, breaking existing links and requiring HTML updates; it also does not invalidate the old cached keys.

500
MCQeasy

A company uses AWS CodeDeploy to deploy applications to an Auto Scaling group. The deployment fails because the new version of the application crashes the instances. The DevOps engineer needs the Auto Scaling group to automatically replace the unhealthy instances with the previous working version. Which deployment configuration should the engineer use?

A.In-place deployment with a deployment group that has a failure threshold of 0.
B.Blue/Green deployment with a load balancer to switch traffic only after health checks pass.
C.Canary deployment that shifts 10% of traffic to the new version, then 100% after 10 minutes.
D.Linear deployment that shifts 10% of traffic every 10 minutes.
AnswerB

Blue/Green deployment creates a completely separate green environment and reroutes traffic only after the new instances pass health checks. The load balancer is the key component: it keeps traffic anchored to the blue environment until the green is verified healthy, and if health checks fail, you simply do not cut over or you can switch back to blue instantly. This gives an automatic, low-risk rollback path because the original environment remains intact and available. Thus, it satisfies the need to revert to the original version when issues are detected.

Why this answer

A blue/green deployment with a load balancer health check ensures that the new (green) instances are validated before any traffic is routed to them. If the new version crashes, the health checks fail, the load balancer keeps traffic on the old (blue) instances, and the Auto Scaling group can automatically terminate the unhealthy green instances and replace them with the previous working version by reverting to the original launch configuration or template.

Exam trap

The trap here is that candidates often confuse deployment strategies (in-place, canary, linear) with rollback mechanisms, assuming that any traffic-shifting method automatically replaces unhealthy instances with the previous version, when in fact only blue/green deployments inherently isolate the new environment and allow a clean revert without affecting the old instances.

How to eliminate wrong answers

Option A is wrong because an in-place deployment with a failure threshold of 0 means the deployment will stop as soon as any single instance fails, but it does not automatically replace unhealthy instances with the previous working version; it simply halts the deployment, leaving the failed instances in place. Option C is wrong because a canary deployment shifts a small percentage of traffic to the new version and then fully shifts after a time window, but if the new version crashes instances, the canary instances become unhealthy and the deployment may still proceed to full rollout if the health check grace period expires, failing to automatically revert to the previous version. Option D is wrong because a linear deployment incrementally shifts traffic in steps, but like the canary, it does not inherently replace crashed instances with the previous working version; it only controls traffic shifting, not instance recovery or rollback.

501
MCQhard

A company uses AWS CodeDeploy to deploy a web application to an Auto Scaling group. The deployment fails during the 'ValidateService' lifecycle event. The CloudWatch Agent reports that the target process is running but the health check endpoint returns HTTP 503. The CodeDeploy agent logs show no errors. What is the most likely cause of the failure?

A.The Auto Scaling group is not healthy
B.The CodeDeploy agent is not installed on the instances
C.The application is not fully functional due to missing configuration files
D.The target process is not listening on the expected port
AnswerC

A running process with an HTTP 503 health endpoint indicates the application started but cannot serve requests, typically because required configuration files are absent. The CodeDeploy agent logs no errors, confirming the failure is application-level, not deployment-mechanism-level.

Why this answer

The 'ValidateService' lifecycle event in CodeDeploy runs a health check against the application endpoint. A 503 HTTP status indicates the web server is running (the target process is up) but the application itself is not fully functional, often due to missing configuration files, environment variables, or dependencies. The CloudWatch Agent confirming the process is running and the CodeDeploy agent logs showing no errors further isolate the issue to the application layer, not the deployment infrastructure.

Exam trap

The trap here is that candidates may assume a running process (confirmed by CloudWatch Agent) means the application is fully functional, but the 503 status explicitly indicates the application layer is failing, not the process or network layer.

How to eliminate wrong answers

Option A is wrong because an unhealthy Auto Scaling group would cause the instance to be terminated or fail health checks at the EC2 level, not specifically result in a 503 from the application health check endpoint during CodeDeploy's ValidateService hook. Option B is wrong because if the CodeDeploy agent were not installed, the deployment would fail much earlier (e.g., during the 'DownloadBundle' or 'Install' events) and the agent logs would show errors or be absent, not report no errors. Option D is wrong because the CloudWatch Agent reports the target process is running, which implies the process is listening on its expected port; a port mismatch would typically cause a connection refused (e.g., ECONNREFUSED) or timeout, not an HTTP 503.

502
Multi-Selectmedium

A company is using Amazon CloudWatch Logs to collect logs from multiple EC2 instances. They need to filter logs in real time and send specific log events to a custom application for processing. Which TWO services can they use to achieve this?

Select 2 answers
A.Use Amazon Kinesis Data Analytics to process the log stream.
B.Configure a CloudWatch Logs subscription filter that invokes an AWS Lambda function.
C.Create a CloudWatch Events rule to capture log events and send them to Amazon SQS.
D.Configure a CloudWatch Logs subscription filter that sends data to Amazon Kinesis Data Firehose.
E.Use Amazon S3 event notifications to trigger a Lambda function on new log files.
AnswersB, D

A CloudWatch Logs subscription filter can be configured with a Lambda function as its destination, enabling real-time processing of log events as they arrive. When a log event matches the filter pattern, CloudWatch Logs invokes the Lambda function asynchronously, passing the event payload (base64-encoded and gzipped) for your custom logic to parse and forward. This is a native, low-latency pattern that satisfies the requirement to filter and stream logs in real time, and it is a widely recommended approach for building log-processing pipelines.

Why this answer

Option B is correct because a CloudWatch Logs subscription filter performs real-time filtering of log events and can deliver matching events directly to AWS Lambda, which then runs the custom application logic for processing. Option D is correct because a subscription filter can also stream filtered log events to Amazon Kinesis Data Firehose, which reliably delivers them to a custom destination for processing. Option A is incorrect because Kinesis Data Analytics analyzes streaming data but is not a CloudWatch Logs subscription destination for real-time log filtering.

Option C is incorrect because CloudWatch Events (EventBridge) rules react to AWS service events, not individual CloudWatch Logs log events, and cannot filter log events to SQS. Option E is incorrect because S3 event notifications only trigger on object creation in S3 and do not provide real-time filtering of CloudWatch Logs streams.

Exam trap

DOP-C02 often tests the difference between CloudWatch Logs subscription filters and CloudWatch Events; candidates may confuse the two and select CloudWatch Events for log processing.

503
MCQeasy

A company deploys an AWS CloudFormation stack that creates an Auto Scaling group and a launch template. The team wants every code change merged to the main branch to automatically update the stack, and they want to be able to roll back to the previous stack state if the update fails. Which combination of services should they use?

A.Amazon EventBridge with a rule that matches CodeCommit events and targets an AWS Step Functions state machine that calls cloudformation update-stack.
B.AWS CodeCommit with a repository trigger that invokes an AWS Lambda function to call cloudformation create-stack.
C.AWS CodePipeline with a CodeCommit source and a CloudFormation deploy action configured to create or update the stack.
D.AWS CodeBuild with a buildspec that runs the AWS CLI cloudformation deploy command and ignores failures.
AnswerC

CodePipeline can trigger on commits to the main branch through the CodeCommit source action, and the CloudFormation deploy action performs create or update semantics on the stack. CloudFormation automatically rolls back to the previous stack state if the update fails, satisfying both the automation and rollback requirements natively.

Why this answer

The requirement is automated deployment on main branch commits plus automatic rollback on failure. CodePipeline with a CodeCommit source and a CloudFormation deploy action provides both: the source action triggers on branch changes, and the deploy action creates or updates the stack with CloudFormation's built-in rollback on failure, requiring minimal custom code.

Exam trap

The trap here is choosing a custom orchestration with Lambda or Step Functions when the native pipeline and CloudFormation deploy action already provide triggering and rollback.

504
Multi-Selecthard

A company has a microservices architecture running on Amazon ECS with Fargate launch type. Each service is deployed in multiple Availability Zones. The services communicate via REST APIs. Recently, a downstream service experienced a partial outage, causing upstream services to time out and leading to cascading failures. The team wants to improve resilience against such failures. Which combination of actions should the DevOps engineer take? (Choose TWO.)

Select 2 answers
A.Increase the HTTP timeout values for all service-to-service calls.
B.Implement circuit breaker patterns in the service clients.
C.Remove all retry logic from service calls.
D.Adopt an asynchronous communication pattern using Amazon SQS or Amazon EventBridge.
E.Configure Auto Scaling for all services based on request count.
AnswersB, D

Circuit breakers in service clients detect repeated downstream failures and trip open, failing fast instead of holding threads until timeout. This halts the cascade at the caller, satisfying the resilience requirement by preventing upstream exhaustion when a downstream service partially fails.

Why this answer

Option B is correct because a circuit breaker in the service clients detects repeated failures from a downstream dependency and trips open, failing fast instead of letting upstream threads block on slow REST calls, which directly prevents the timeout propagation and cascading failures described. Option D is correct because moving to asynchronous communication with Amazon SQS or Amazon EventBridge decouples services: upstream services enqueue or publish events and return immediately, so a partial outage in a downstream consumer no longer blocks callers, and messages can be retried or buffered until the consumer recovers. Option A is not appropriate because increasing HTTP timeouts makes callers wait longer, consuming threads and connections and worsening cascading failures rather than containing them.

Option C is wrong because removing retry logic eliminates a useful resilience mechanism for transient errors; retries should be bounded and combined with circuit breakers and backoff, not deleted. Option E is not the right fix because scaling on request count does not address a downstream dependency that is failing or slow, and could even amplify load against the impaired service.

Exam trap

DOP-C02 often tests whether candidates confuse 'make the timeout longer' with resilience — longer timeouts amplify cascading failures rather than preventing them.

505
Multi-Selectmedium

A DevOps engineer needs to set up centralized logging for an application running on multiple EC2 instances across different AWS accounts. The logs must be aggregated in a single S3 bucket and also be analyzed in near real-time. Which TWO services should be used together to achieve this?

Select 2 answers
A.Amazon Simple Queue Service (SQS)
B.Amazon Kinesis Data Firehose
C.AWS CloudTrail
D.Amazon CloudWatch Logs subscription
E.AWS Lambda
AnswersB, D

Amazon Kinesis Data Firehose is a fully managed streaming ingestion service that can receive log events directly from a CloudWatch Logs subscription and deliver them near-real-time to centralized destinations such as Amazon S3, Amazon Redshift, or Amazon OpenSearch Service. It automatically handles buffering, compression, and encryption, which reduces cost and operational overhead. This makes it the ideal component for aggregating logs from multiple AWS accounts into a single, queryable data lake for centralized analysis.

Why this answer

Option B, Amazon Kinesis Data Firehose, is correct because it can ingest streaming log data and reliably deliver it to a single S3 bucket, with optional near real-time transformation and buffering, satisfying both the aggregation and near real-time analysis requirements. Option D, Amazon CloudWatch Logs subscription, is correct because a subscription filter on a CloudWatch log group can stream log events in near real time to Kinesis Data Firehose, enabling centralized collection from EC2 instances across multiple AWS accounts. Together, CloudWatch Logs subscriptions feed Firehose, which delivers the aggregated logs to S3.

Option A, Amazon SQS, is not appropriate because it is a message queue, not a log ingestion or delivery service to S3. Option C, AWS CloudTrail, records API activity rather than application logs, so it does not meet the application logging requirement. Option E, AWS Lambda, is compute for processing events and is not the primary service for aggregating logs into S3.

Exam trap

DOP-C02 often tests the confusion between CloudWatch Logs (application logs) and CloudTrail (API audit logs), and between SQS (decoupling) and Firehose (delivery to S3) — candidates must recognize that only the CloudWatch Logs subscription + Firehose pairing provides both aggregation and near real-time delivery.

506
MCQmedium

An application log excerpt shows repeated HTTP 500 errors for the /api/orders endpoint, with occasional successful health checks. The application runs on EC2 instances behind an ALB. What is the MOST likely cause of this pattern?

A.The backend service that the /api/orders endpoint depends on is unavailable or failing.
B.The EC2 instances are running out of memory and the application is crashing.
C.The ALB is misconfigured and routing requests to the wrong target group.
D.The EC2 instances are not passing health checks and are being deregistered from the target group.
AnswerA

The HTTP 500 on /api/orders is generated by the application after the ALB successfully forwards the request, meaning the instance is reachable and the web server process is handling traffic. However, the /api/orders endpoint depends on a downstream service (e.g., database, internal microservice, cache) whose failure causes the application to throw an unhandled exception and return a 500. Health checks succeed because the configured health check path (typically /health or /) exercises simple connectivity and doesn't call the failing dependency, so the instance remains in service. This is a classic partial-failure scenario where the dependency, not the instance, is unhealthy.

Why this answer

The pattern of repeated HTTP 500 errors for /api/orders with occasional successful health checks strongly indicates that the backend service dependency (e.g., a database, cache, or another microservice) is intermittently failing or unavailable. HTTP 500 errors are server-side errors, meaning the application code is running but cannot complete the request due to a downstream failure. Successful health checks confirm the EC2 instances themselves are healthy and in-service, ruling out instance-level or ALB misconfiguration issues.

Exam trap

The trap here is that candidates confuse HTTP 500 errors with instance-level failures (like OOM or health check failures), but the key differentiator is that successful health checks prove the instances are operational, shifting the root cause to a failing backend dependency rather than the compute layer.

How to eliminate wrong answers

Option B is wrong because running out of memory typically causes the application process to crash (e.g., OOM killer), leading to connection timeouts or immediate 503 errors, not repeated HTTP 500 errors with successful health checks. Option C is wrong because a misconfigured ALB routing to the wrong target group would cause requests to reach instances that don't serve the /api/orders endpoint, resulting in 404 or 503 errors, not 500 errors from the application itself. Option D is wrong because if instances were failing health checks and being deregistered, they would be removed from the target group and stop receiving traffic entirely, which contradicts the observed pattern of occasional successful health checks and persistent 500 errors on /api/orders.

507
MCQeasy

A company uses Amazon CloudWatch to monitor its production environment. The DevOps team wants to receive an email notification whenever the average CPU utilization of any EC2 instance exceeds 90% for 5 consecutive minutes. Which steps should be taken to set up this notification?

A.Install the CloudWatch Logs agent on each EC2 instance and configure a metric filter to trigger an SNS notification
B.Create a CloudWatch alarm on CPUUtilization with a threshold of 90% for 5 consecutive periods, and configure an SNS topic to send email
C.Use AWS CloudTrail to monitor CPU utilization and send notifications via SNS
D.Use AWS Config to create a rule that triggers an SNS notification when CPU utilization exceeds 90%
AnswerB

CloudWatch automatically emits the CPUUtilization metric for each EC2 instance, expressed as a percentage, so an alarm can directly evaluate it. Configuring the alarm with a threshold of 90% for five consecutive evaluation periods (5-minute periods, or 1-minute if detailed monitoring is enabled) ensures the alert only fires on sustained load, not transient spikes. The action sends an SNS notification to a topic with an email subscription, which is the canonical AWS approach for metric-based alerting.

Why this answer

Create a CloudWatch alarm on the CPUUtilization metric with a threshold of 90% for 5 consecutive periods (5 minutes assuming 1-minute periods), and configure an SNS topic to send email notifications. Option A is incorrect because the CloudWatch Logs agent collects log data, not metrics, and metric filters are for log analysis, not setting alarms on CPU utilization. Option C is incorrect because AWS CloudTrail logs API calls, not CPU utilization metrics.

Option D is incorrect because AWS Config is for resource configuration auditing and compliance, not for monitoring CPU utilization metrics.

508
MCQhard

A company runs a critical application on an Amazon RDS for MySQL DB instance. The application experiences intermittent connection timeouts. The DevOps team notices that the DB instance's CPU and memory metrics are normal. What should the team check NEXT to diagnose the issue?

A.Enable Enhanced Monitoring to check OS-level metrics
B.Examine the slow query log to identify long-running queries
C.Verify that the DB instance's storage is not full
D.Check the 'DatabaseConnections' CloudWatch metric to see if the connection count is near the max_connections limit
AnswerD

The DatabaseConnections CloudWatch metric directly tracks the number of client connections currently established to the RDS instance, and when this value approaches the max_connections parameter, the server begins rejecting or timing out new connection attempts. This pattern perfectly matches the symptom of intermittent connection timeouts while CPU and memory remain normal, because the connection ceiling is a hard limit unaffected by resource utilization. Comparing this metric against the configured max_connections (from the parameter group) confirms whether connection exhaustion is the cause.

Why this answer

The most likely cause of intermittent connection timeouts when CPU and memory are normal is that the number of database connections has reached the max_connections limit. The 'DatabaseConnections' CloudWatch metric shows the current number of connections; if it's near the limit, new connections will be refused or time out. Checking this metric is the next logical step to diagnose the issue.

Exam trap

DOP-C02 often tests the misconception that connection timeouts are always due to slow queries or resource exhaustion, but they can also be caused by connection limits, which are not reflected in CPU/memory metrics.

How to eliminate wrong answers

Option A is wrong because Enhanced Monitoring provides OS-level metrics, but CPU and memory are already normal, so OS-level metrics are less likely to reveal the cause. Option B is wrong because slow queries would typically increase CPU or memory usage, and they cause slow responses, not necessarily connection timeouts. Option C is wrong because a full storage would cause write failures, not intermittent connection timeouts, and it would likely be accompanied by other errors.

509
MCQmedium

A development team uses AWS CodePipeline with multiple stages including source, build, and deploy. The pipeline uses an Amazon S3 source action that triggers on changes to a specific bucket. Recently, the pipeline stopped triggering automatically. The IAM role for CodePipeline has the necessary permissions. What is the most likely cause?

A.The IAM role for CodePipeline does not have s3:GetObject permission.
B.The S3 bucket policy denies access to CodePipeline.
C.The S3 bucket does not have event notifications configured.
D.AWS CloudTrail is not configured to deliver S3 data events to CloudWatch Logs.
AnswerD

CodePipeline uses Amazon EventBridge rules that match S3 object-created events recorded by AWS CloudTrail as data events. For these events to reach EventBridge, CloudTrail must be enabled to log S3 data events for the source bucket and deliver them to a CloudWatch Logs log group; the EventBridge rule then forwards matching events to the pipeline. If CloudTrail is not configured to deliver S3 data events to CloudWatch Logs, the pipeline's trigger remains silent even though the source code changes, causing the pipeline to appear stalled or never start. This is the root cause, as the IAM role, bucket policy, and S3 event notifications are all in order.

Why this answer

CodePipeline's S3 source action does not rely on S3 event notifications to trigger pipeline executions. Instead, it uses Amazon CloudWatch Events (now Amazon EventBridge) to detect changes to the S3 bucket. For this to work, AWS CloudTrail must be configured to deliver S3 data events (specifically `PutObject` API calls) to CloudWatch Logs, which then generates the event that triggers the pipeline.

Without CloudTrail data event logging, CodePipeline cannot detect object uploads, even if the IAM role has proper permissions.

Exam trap

The trap here is that candidates assume S3 event notifications are required for CodePipeline triggers, but AWS actually uses CloudTrail and EventBridge, so the correct answer focuses on CloudTrail configuration rather than S3 notifications.

How to eliminate wrong answers

Option A is wrong because `s3:GetObject` permission is required for CodePipeline to read the source artifact during the source stage, but the issue is about the pipeline not triggering automatically, not about failing to retrieve objects. Option B is wrong because a bucket policy denying access would cause a permission error when CodePipeline tries to access the bucket, but the question states the IAM role has necessary permissions, and the pipeline stopped triggering, not failing on access. Option C is wrong because S3 event notifications are not used by CodePipeline's S3 source action; CodePipeline uses CloudTrail and EventBridge for change detection, so missing event notifications would not prevent automatic triggers.

510
Multi-Selectmedium

A company runs a microservices application on Amazon ECS with Fargate. The application includes a service that processes orders and stores them in an RDS PostgreSQL database. The company wants to ensure that the order service is resilient to AZ failures and can handle a sudden increase in order volume. Which TWO actions should the DevOps engineer take? (Choose TWO.)

Select 2 answers
A.Increase the CPU and memory limits for the ECS task definition.
B.Place an Amazon CloudFront distribution in front of the order service.
C.Deploy the RDS instance in a Multi-AZ configuration.
D.Configure the ECS service to run tasks in multiple Availability Zones.
E.Use RDS Proxy to manage database connections.
AnswersC, D

Deploying RDS in a Multi-AZ configuration creates a synchronous standby replica in a different Availability Zone, and Amazon RDS automatically fails over to the standby if the primary instance becomes unhealthy or the AZ fails. The DNS name stays the same, so the ECS order service can reconnect without code changes, and the standby is continuously updated with synchronous replication to prevent data loss. This directly addresses the database as a single point of failure, which is necessary because the order service depends on durable transactions to record orders.

Why this answer

Deploying the RDS instance in a Multi-AZ configuration provides automatic failover to a standby replica in a different Availability Zone, ensuring database resilience to AZ failures. Option D is correct because configuring the ECS service to run tasks in multiple Availability Zones distributes the order processing workload across AZs, improving both fault tolerance and scalability during sudden traffic spikes.

Exam trap

The trap here is that candidates often confuse connection pooling (RDS Proxy) with high availability (Multi-AZ) or assume that vertical scaling (increasing task limits) is sufficient for both resilience and sudden load, when in fact horizontal distribution across AZs is required for fault tolerance and elasticity.

511
MCQhard

A team uses AWS CodePipeline to deploy a containerized application to Amazon ECS. The pipeline uses a source stage from CodeCommit, a build stage that builds a Docker image and pushes it to Amazon ECR, and a deploy stage that updates an ECS service. The team wants to add a manual approval step before the deploy stage to allow QA to verify the image. What is the BEST way to implement this?

A.Configure an AWS Lambda function in the pipeline that checks a DynamoDB table for approval status and pauses until approved.
B.Use an Amazon SNS topic to send a notification to QA, and have them manually trigger the deploy stage by clicking a link in the email.
C.Use Amazon CloudWatch Events to trigger a custom action that waits for an approval signal.
D.Add a manual approval stage in CodePipeline between the build and deploy stages, and configure SNS to notify approvers.
AnswerD

The native CodePipeline manual approval action is the correct pattern: it creates a gate that pauses the pipeline after the build stage and does not proceed to deploy until an approved or rejected decision is recorded. When you add the action, you configure an SNS topic for notifications, and approvers with the proper IAM policy respond through the console or CLI with `put-approval-result`. This integrated workflow provides explicit audit trails and automatically resumes only on approval, which is far more reliable than any external workaround. It is purpose-built to block stage transitions and supports both email SNS notifications and custom SNS topics for team alerting.

Why this answer

CodePipeline natively supports manual approval actions that pause the pipeline at a specified stage and wait for an approver to manually approve or reject the transition. By adding a manual approval stage between the build and deploy stages, the pipeline will automatically halt after the build completes, and you can configure Amazon SNS to notify the QA team via email or other endpoints when their approval is required. This approach requires no custom infrastructure, integrates directly with the pipeline's state machine, and provides a built-in audit trail of approvals.

Exam trap

The trap here is that candidates often over-engineer a solution by introducing custom polling, Lambda functions, or external triggers, when AWS CodePipeline already provides a fully managed, native manual approval action that handles pausing, notification, and resumption without any custom code.

How to eliminate wrong answers

Option A is wrong because using a Lambda function to poll a DynamoDB table for approval status introduces unnecessary complexity, latency, and custom code; CodePipeline already provides a native manual approval action that handles pausing and resuming the pipeline without custom polling logic. Option B is wrong because SNS notifications alone cannot pause the pipeline or trigger the deploy stage; clicking a link in an email cannot programmatically resume a CodePipeline execution without a custom webhook or API integration, and the pipeline would continue past the deploy stage immediately if no blocking mechanism is in place. Option C is wrong because CloudWatch Events can trigger actions based on pipeline state changes but cannot natively pause a pipeline and wait for an approval signal; the manual approval action in CodePipeline is the correct mechanism for inserting a human-in-the-loop gate.

512
MCQmedium

You are a DevOps engineer for a company that uses AWS CodePipeline to deploy a microservice to Amazon ECS with Fargate. The pipeline has a source stage (CodeCommit), a build stage (CodeBuild) that builds a Docker image and pushes it to Amazon ECR, and a deploy stage that uses an ECS task definition update. Recently, the deploy stage started failing intermittently with the error 'The task definition does not have a compatibilities attribute set correctly.' The task definition is generated dynamically during the build stage and uses the 'FARGATE' launch type. The error occurs only when a new task definition revision is created. You suspect the issue is related to how the task definition is generated. Upon reviewing the buildspec, you see that the task definition JSON is created using environment variables for the image URI. What is the MOST likely cause and solution?

A.The task definition is missing the 'executionRoleArn' field, which is required for Fargate.
B.The task definition JSON does not include the 'requiresCompatibilities' field with the value '["FARGATE"]'.
C.The task definition specifies 'networkMode' as 'bridge', but Fargate requires 'awsvpc'.
D.The task definition does not specify 'cpu' and 'memory' values, which are required for Fargate.
AnswerB

The `requiresCompatibilities` array must explicitly contain the value `"FARGATE"` so that ECS can verify the task definition is eligible to launch on Fargate. When this field is missing, any attempt to run the task with `launchType: FARGATE` produces an `InvalidParameterException` stating that the task definition is not compatible with the requested launch type. This field is an explicit declaration independent of networkMode, CPU, or memory, and is the exact missing piece that triggers the complaint about compatibility.

Why this answer

The 'requiresCompatibilities' attribute must be explicitly set to 'FARGATE' for Fargate tasks. Option A is incorrect because the error is about compatibilities, not execution role. Option C is incorrect because network mode should be 'awsvpc', but that is not the error.

Option D is incorrect because CPU and memory values are required but would cause a different error.

513
MCQhard

A company requires that all secrets (e.g., database passwords) used by Lambda functions be rotated automatically every 30 days. Which combination of services should be used?

A.AWS CloudHSM and AWS Lambda
B.AWS Secrets Manager and AWS Lambda
C.AWS Systems Manager Parameter Store and AWS Lambda
D.AWS KMS and AWS Lambda
AnswerB

AWS Secrets Manager is purpose-built for storing, retrieving, and automatically rotating database credentials and other sensitive secrets, including integration with Amazon RDS, Redshift, and DocumentDB. The rotation feature uses an AWS Lambda function to update credentials on a schedule, ensuring that database passwords are cycled without manual intervention. Lambda can also retrieve secrets at runtime via the AWS SDK using IAM permissions, making it the correct and most appropriate pairing for the company's requirement.

Why this answer

AWS Secrets Manager is the correct choice because it natively supports automatic secret rotation on a configurable schedule (e.g., every 30 days) using a Lambda function as the rotation handler. Secrets Manager directly integrates with Lambda to invoke the rotation logic, updating the secret value and propagating the change to the target database or service without custom infrastructure. CloudHSM, Parameter Store, and KMS do not provide built-in, scheduled rotation of secrets with automatic Lambda invocation.

Exam trap

The trap here is that candidates confuse AWS Systems Manager Parameter Store (which can store secrets but lacks automatic rotation) with AWS Secrets Manager (which is purpose-built for rotation), or they assume KMS or CloudHSM can manage secrets directly when they only handle encryption keys.

How to eliminate wrong answers

Option A is wrong because AWS CloudHSM is a hardware security module for key generation and cryptographic operations, not a service for storing or rotating secrets like database passwords; it lacks any built-in rotation scheduling or Lambda integration for secret rotation. Option C is wrong because AWS Systems Manager Parameter Store can store secrets but does not natively support automatic rotation; any rotation would require custom orchestration and polling, whereas Secrets Manager provides managed rotation with a single API call. Option D is wrong because AWS KMS is a key management service for encryption keys, not a secret store; it cannot store or rotate secrets like database passwords, and while it can encrypt secrets stored elsewhere, it does not provide rotation logic.

514
Multi-Selectmedium

A DevOps team is designing a CI/CD pipeline for a containerized application. Which THREE components are essential for a complete pipeline? (Choose three.)

Select 3 answers
A.AWS CodeDeploy
B.Artifact storage
C.Amazon CloudWatch
D.Build and test automation
E.Source control repository
AnswersB, D, E

Artifact storage is the durable, versioned repository that holds the output of the build stage, such as Docker images in Amazon ECR or packages in S3, until the deployment stage consumes them. Without this component, builds are ephemeral and you cannot reliably reproduce a release, roll back to a known-good version, or audit exactly what was deployed. In a container context, the image registry is a distinct stateful service that integrates with IAM and image scanning, making it a non-negotiable pipeline component.

Why this answer

Artifact storage (Option B) is essential because a complete CI/CD pipeline must store build outputs (e.g., Docker images, JAR files) in a durable, versioned repository. Without artifact storage, subsequent deployment stages cannot reliably retrieve the exact build artifact that passed testing, breaking traceability and rollback capabilities. Services like Amazon ECR or S3 serve this role, ensuring immutability and consistent delivery across environments.

Exam trap

The trap here is that candidates confuse 'deployment service' (CodeDeploy) with a pipeline component, or mistake monitoring (CloudWatch) as essential, when the question specifically asks for the three core stages that form a complete pipeline: source, build/test, and artifact storage.

515
Multi-Selectmedium

A company is designing a highly available architecture for a stateless web application using AWS services. Which TWO steps should they take to achieve high availability?

Select 2 answers
A.Store session state in an EBS volume attached to each instance
B.Deploy EC2 instances in multiple Availability Zones
C.Use a single NAT instance in a public subnet
D.Use only M5 instance types for better performance
E.Use an Application Load Balancer to distribute traffic
AnswersB, E

Distributing EC2 instances across multiple Availability Zones ensures the application can tolerate a complete failure of one physical data center, because the remaining instances continue to serve traffic. Each AZ has independent power, cooling, and network connectivity, so an outage in one AZ does not affect the others. This redundancy is the fundamental building block of high availability on AWS and is a necessary condition for achieving a higher service-level agreement.

Why this answer

Option B is correct because deploying EC2 instances across multiple Availability Zones ensures the application survives an AZ-level failure, which is a fundamental requirement for high availability in AWS. Option E is correct because an Application Load Balancer distributes incoming traffic across healthy targets in multiple AZs, performs health checks, and automatically routes around failed instances, directly supporting high availability for a stateless web tier. Option A is incorrect because storing session state on an EBS volume tied to a single instance creates a single point of failure and is unnecessary for a stateless application.

Option C is incorrect because a single NAT instance is itself a single point of failure and cannot provide high availability. Option D is incorrect because choosing a specific instance type like M5 improves performance but does nothing to increase availability.

Exam trap

The trap is thinking that a single NAT instance or EBS-backed session state can provide HA; both introduce single points of failure.

516
MCQeasy

A company uses AWS CodePipeline with multiple stages: Source, Build, Test, Deploy. The Test stage runs integration tests using AWS CodeBuild. If the Test stage fails, what happens to the pipeline execution?

A.The pipeline continues to the next stage but marks the Test stage as failed.
B.The pipeline execution stops and the status is set to 'Failed'.
C.The pipeline skips the Test stage and proceeds to Deploy.
D.The pipeline automatically retries the Test stage up to three times.
AnswerB

This is correct. When any action in the Test stage fails, CodePipeline immediately transitions the entire pipeline execution to the Failed status, and no later stages run. The execution history preserves the failure details for inspection, but you must resolve the root cause and retry the failed action or start a new execution manually; the pipeline does not recover on its own.

Why this answer

In AWS CodePipeline, by default, if a stage (such as Test) fails, the pipeline execution immediately stops and the overall pipeline status is set to 'Failed'. This is because CodePipeline treats each stage as a sequential gate; a failure in any stage blocks progression to subsequent stages unless explicitly configured otherwise (e.g., with a 'Blocker' or 'Manual Approval' action). Option B correctly describes this default behavior.

Exam trap

The trap here is that candidates may confuse the default behavior with optional features like automatic retries or stage skipping, assuming CodePipeline behaves like a CI/CD tool that allows failures to pass through (e.g., Jenkins with 'unstable' status) or automatically retries failed jobs.

How to eliminate wrong answers

Option A is wrong because CodePipeline does not continue to the next stage after a failure; it halts execution and marks the pipeline as 'Failed', not just the stage. Option C is wrong because CodePipeline does not skip a failed stage; it stops entirely, preventing the Deploy stage from running. Option D is wrong because CodePipeline does not automatically retry a failed stage; retry behavior must be explicitly configured using the 'Retry' feature in the pipeline settings or via manual intervention.

517
MCQmedium

An e-commerce platform uses Amazon DynamoDB as its primary database. The platform experiences occasional read throttling during flash sales. The operations team needs to ensure that read traffic is handled without errors, while keeping costs low. What should a DevOps engineer recommend?

A.Enable DynamoDB Accelerator (DAX) to cache frequently read data.
B.Increase the read capacity units for the table during flash sale events.
C.Use DynamoDB Streams to replicate reads to a separate table.
D.Implement Global Tables to distribute read traffic across multiple regions.
AnswerA

DAX is a fully managed, in-memory caching service placed in front of DynamoDB, returning cached items with microsecond latency. By writing through and caching the frequently read flash-sale items, it absorbs the burst of read traffic before it reaches the table, which directly reduces consumed read capacity units and throttling events. This requires no costly rearchitecture or constantly adjusting provisioned throughput, making it the most appropriate solution for unpredictable read spikes.

Why this answer

DynamoDB Accelerator (DAX) is an in-memory cache for DynamoDB that reduces read load on the table by serving repeated read requests from cache. During flash sales, read traffic spikes on popular items; DAX absorbs these reads, preventing throttling and improving latency, while keeping costs low because it reduces the need to over-provision read capacity. This is the most cost-effective solution for read-heavy, repetitive access patterns.

Exam trap

The trap is thinking that increasing RCUs or using Global Tables is the best fix for read throttling, when the question emphasizes cost and handling read traffic without errors; DAX is the purpose-built caching solution.

How to eliminate wrong answers

Option B is wrong because increasing read capacity units (RCUs) during flash sales is a manual, reactive approach that can be costly if over-provisioned and still may not handle sudden spikes quickly enough; it also does not reduce costs. Option C is wrong because DynamoDB Streams capture item-level changes for replication or triggers, not for serving reads, and replicating to another table does not offload read traffic from the original table. Option D is wrong because Global Tables are for multi-region active-active replication, which increases cost and complexity, and does not directly solve read throttling in a single region unless reads are distributed across regions, which may introduce latency.

518
MCQeasy

A company is deploying a critical application on Amazon EC2 instances behind an Application Load Balancer (ALB) across multiple Availability Zones. The application must be resilient to the failure of an entire Availability Zone. Which design should the company implement?

A.Launch EC2 instances in at least two Availability Zones and place them behind an Application Load Balancer with cross-zone load balancing enabled.
B.Use one EC2 instance in a single Availability Zone behind a Network Load Balancer.
C.Launch EC2 instances in one Availability Zone and use an Application Load Balancer to distribute traffic.
D.Deploy EC2 instances in two Availability Zones but use a single Application Load Balancer in one AZ.
AnswerA

This is the correct approach. An Application Load Balancer is a regional service; by enabling subnets in at least two Availability Zones, you create redundant ALB nodes. The ALB performs health checks and automatically routes traffic to healthy EC2 instances across both AZs. Cross-zone load balancing ensures each instance receives an equal share of requests, so the architecture tolerates an AZ failure and even an instance failure without manual intervention.

Why this answer

Deploying EC2 instances across at least two Availability Zones (AZs) behind an Application Load Balancer (ALB) with cross-zone load balancing enabled ensures that if an entire AZ fails, the ALB can route traffic to healthy instances in the remaining AZs. Cross-zone load balancing allows the ALB to distribute incoming requests evenly across all registered instances in all enabled AZs, which improves fault tolerance and resource utilization. This design meets the requirement for resilience to an AZ failure by eliminating a single point of failure at the AZ level.

Exam trap

The trap here is that candidates often assume that simply placing instances in multiple AZs behind a load balancer is sufficient, but they overlook the critical requirement that the load balancer itself must be deployed across multiple AZs to avoid being a single point of failure.

How to eliminate wrong answers

Option B is wrong because using a single EC2 instance in one AZ behind a Network Load Balancer (NLB) does not provide resilience to an AZ failure; if that AZ goes down, the application becomes unavailable. Option C is wrong because launching EC2 instances in only one AZ behind an ALB still creates a single point of failure at the AZ level; the ALB cannot route traffic to healthy instances if the entire AZ fails. Option D is wrong because deploying EC2 instances in two AZs but using a single ALB in one AZ means the ALB itself is a single point of failure; if that AZ fails, the ALB becomes unavailable, and traffic cannot be distributed to instances in the other AZ.

519
MCQmedium

Refer to the exhibit. A DevOps engineer runs the above commands. The build project 'my-project' uses an S3 bucket as source and another S3 bucket for artifacts. The build fails with an 'Access Denied' error when trying to download the source code. What is the most likely cause?

A.The encryption key is a KMS key that the role cannot access
B.The service role does not have s3:GetObject permission on the source bucket
C.The source type is S3, but the project expects CodeCommit
D.The source location is incorrect
AnswerB

The service role is the IAM role that CodePipeline assumes to perform actions on your behalf. To pull source artifacts from S3, the role must have an IAM policy allowing s3:GetObject (and typically s3:ListBucket) on the specified bucket and prefix. The error indicates the pipeline cannot download the object, which is a direct consequence of the role lacking this permission. Adding a policy statement with s3:GetObject on the source bucket ARN will resolve the stage failure.

Why this answer

The build project 'my-project' uses an S3 bucket as the source. When CodeBuild downloads source code from S3, the service role must have the s3:GetObject permission on the source bucket. The 'Access Denied' error indicates that the role lacks this permission, making option B the most likely cause.

Exam trap

The trap here is that candidates may confuse 'Access Denied' with other S3 errors like 'NoSuchKey' or 'BucketNotFound', or incorrectly attribute the error to KMS encryption when the error message does not reference it.

How to eliminate wrong answers

Option A is wrong because the error message does not mention KMS or encryption key issues; an 'Access Denied' for KMS would typically include a specific message about the key. Option C is wrong because the project is configured to use an S3 source, not CodeCommit, and the error is about access, not source type mismatch. Option D is wrong because an incorrect source location would result in a 'NoSuchKey' or '404' error, not an 'Access Denied' error.

520
MCQmedium

A DevOps engineer is designing a configuration management solution for a fleet of EC2 instances. The instances are ephemeral and frequently replaced by an Auto Scaling group. The engineer needs to ensure that newly launched instances are automatically configured with the latest software packages and settings. Which AWS service should be used?

A.AWS CodeDeploy
B.AWS OpsWorks Stacks
C.AWS CloudFormation
D.AWS Systems Manager State Manager
AnswerD

AWS Systems Manager State Manager is the correct choice because it enforces desired configuration on managed instances through associations. An association defines a set of actions, such as running a custom SSM document or a Chef recipe, and State Manager executes it according to a schedule to ensure the instance remains compliant. If an instance drifts from the desired state, the next scheduled execution brings it back, providing continuous configuration management for both EC2 and hybrid instances.

Why this answer

AWS Systems Manager State Manager is the correct choice because it provides a configuration management solution that ensures EC2 instances maintain a desired state. It uses associations to define the software packages, settings, and policies that should be applied to instances, and it automatically applies these configurations to newly launched instances, including those in an Auto Scaling group, without requiring custom scripts or manual intervention.

Exam trap

The trap here is that candidates often confuse AWS Systems Manager State Manager with AWS CodeDeploy or AWS CloudFormation, mistakenly thinking that deployment or provisioning tools also handle ongoing configuration management, but State Manager is specifically designed for maintaining desired state on running instances.

How to eliminate wrong answers

Option A is wrong because AWS CodeDeploy is a deployment service for automating application code deployments, not a configuration management tool for maintaining desired state across ephemeral instances. Option B is wrong because AWS OpsWorks Stacks uses Chef or Puppet for configuration management but requires a persistent stack and agent management, making it less suitable for ephemeral instances that are frequently replaced by Auto Scaling groups. Option C is wrong because AWS CloudFormation is an Infrastructure as Code (IaC) service for provisioning and managing AWS resources, not for ongoing configuration management of software packages and settings on running instances.

521
MCQmedium

A company uses AWS CloudFormation to manage its infrastructure. The DevOps team wants to ensure that stack updates do not accidentally delete critical resources like a database. Which CloudFormation stack policy should they apply to protect the database resource?

A.Create a stack policy that denies delete actions on the logical resource ID of the database.
B.Apply an IAM policy that denies cloudformation:DeleteStack on the database.
C.Use an S3 bucket policy to deny deletion of the database snapshot.
D.Enable termination protection on the CloudFormation stack.
AnswerA

A stack policy is a JSON document attached to a CloudFormation stack that controls which update, replacement, or deletion actions are allowed on the stack's resources. By writing a statement that denies the Delete action on the logical resource ID of the database (e.g., the AWS::RDS::DBInstance resource), CloudFormation will refuse to remove or replace that resource during any stack update. This protects the database from being deleted as a side effect of an update, while still permitting other modifications to proceed.

Why this answer

A CloudFormation stack policy allows you to define resource-level permissions that prevent specific resources (identified by their logical resource ID) from being updated or deleted during a stack update. By creating a policy that denies delete actions on the database's logical resource ID, the DevOps team ensures that even if the template or parameters change, the database resource cannot be accidentally removed.

Exam trap

The trap here is that candidates confuse termination protection (which prevents stack deletion) with resource-level protection during updates, leading them to choose option D instead of understanding that stack policies are needed for granular resource safeguards.

How to eliminate wrong answers

Option B is wrong because an IAM policy denying cloudformation:DeleteStack would block the entire stack deletion, not protect individual resources like a database during an update; it does not prevent resource-level deletion within an update. Option C is wrong because an S3 bucket policy controls access to S3 buckets and objects, not CloudFormation resources; database snapshots are not managed by S3 bucket policies in this context. Option D is wrong because termination protection only prevents the entire stack from being deleted, not individual resources from being replaced or removed during a stack update.

522
MCQmedium

A company's application running on EC2 instances behind an Application Load Balancer (ALB) is returning intermittent 504 errors. The instances are in an Auto Scaling group with a health check grace period of 300 seconds. What should the DevOps engineer check first to troubleshoot the issue?

A.Review the Auto Scaling group scaling policies.
B.Verify the target group health checks are passing.
C.Check security group rules for the ALB.
D.Check ALB access logs for target response times.
AnswerD

ALB access logs are the authoritative source for diagnosing 504 timeouts because they record the HTTP status code and three timing dimensions: request_processing_time, target_processing_time, and response_processing_time. The target_processing_time field directly measures how long the ALB waited for the target to begin sending a response; values at or near the idle timeout threshold precisely identify the request and target instance responsible for the 504. This option provides concrete, per-request evidence of backend latency, unlike health checks or security group rules, which only indicate availability and reachability.

Why this answer

A 504 error indicates the load balancer did not receive a response from the target within the idle timeout period. Checking ALB access logs for target response times is the first step to determine if the backend is slow or unresponsive. Option A is wrong because scaling policies affect the number of instances, not response times.

Option B is wrong because health checks verify instance availability, but intermittent slow responses may not cause health check failures. Option C is wrong because security group rules would cause different errors (e.g., connection timeouts) rather than 504s.

523
MCQeasy

A DevOps engineer is troubleshooting a Lambda function that processes S3 events. The function has been running successfully for months, but today it started timing out. The engineer checks CloudWatch Logs and sees 'Task timed out after 3.01 seconds' errors. The function is configured with a 3-second timeout. What should the engineer do to resolve the issue?

A.Increase the Lambda function reserved concurrency.
B.Increase the memory allocation for the Lambda function.
C.Increase the Lambda function timeout to 10 seconds.
D.Configure a dead-letter queue (DLQ) for the Lambda function.
AnswerC

The Lambda function is failing because its actual execution time is longer than the configured timeout. Raising the timeout to 10 seconds directly expands the allowed execution window, giving the code enough time to complete its work — as long as it stays below the 15-minute maximum. This is the only option that addresses the root cause of a timeout error rather than changing resources or failure handling.

Why this answer

The error 'Task timed out after 3.01 seconds' indicates the function is exceeding its configured 3-second timeout. Since the function has been running successfully for months and only recently started timing out, the workload has likely grown (larger S3 objects, downstream latency, or cold starts). Increasing the timeout to 10 seconds gives the function enough time to complete its processing, directly addressing the timeout error.

Exam trap

The trap here is confusing timeout with concurrency or memory: candidates often assume more memory or concurrency will fix a timeout, but the timeout is a separate, explicitly configured limit that must be raised.

How to eliminate wrong answers

Option A is wrong because reserved concurrency controls how many concurrent invocations can run, not how long a single invocation may run; it would not prevent a timeout. Option B is wrong because increasing memory can improve CPU and slightly speed execution, but it does not extend the timeout window and is not the direct fix for a timeout error. Option D is wrong because a DLQ captures failed asynchronous invocations after they fail; it does not prevent the timeout and would only record the failure.

524
MCQhard

Refer to the exhibit. An IAM policy is attached to a user. The user tries to push a commit to the 'develop' branch of 'MyRepo' using Git. What is the outcome?

A.The push is denied only if the user is pushing to the 'main' branch.
B.The push succeeds because the Allow statement grants GitPush.
C.The push is denied because the user is not pushing to the 'main' branch.
D.The push is denied because the Deny statement blocks all GitPush actions.
AnswerC

This is correct because the Deny statement is scoped to any branch reference that is not `main`. When the user pushes to a branch like `develop`, the condition `StringNotEquals` on `codecommit:References` evaluates to true, causing an explicit Deny that blocks the `GitPush` action. Under IAM's evaluation logic, an explicit Deny overrides the earlier Allow, so the push is denied. A push to `main` would not trigger the Deny and would be allowed, confirming the condition is the decisive factor.

Why this answer

The policy allows GitPull and GitPush on the repo, but denies GitPush when the reference is not 'refs/heads/main'. Since the user is pushing to 'develop', the condition is met, and the Deny applies. An explicit Deny overrides any Allow, so the push is denied.

525
MCQmedium

A DevOps engineer is designing a multi-Region active-active architecture for a stateless web application using Route 53 latency-based routing and DynamoDB global tables. The application must continue to serve traffic even if an entire AWS Region becomes unavailable. Which additional step is MOST critical for resilience?

A.Use an Auto Scaling group with a scheduled scaling policy
B.Enable DynamoDB Accelerator (DAX) in each Region
C.Place a CloudFront distribution in front of the application
D.Configure Route 53 health checks and associate them with the latency records
AnswerD

Route 53 health checks associated with the latency records let DNS stop returning the failed Region's endpoints, redirecting users to a healthy Region. Without this, latency records would continue resolving to the unavailable Region, defeating the active-active resilience requirement.

Why this answer

In a multi-Region active-active architecture using Route 53 latency-based routing, the routing policy alone does not detect or react to a Region failure. Route 53 health checks must be created and associated with each latency record so that unhealthy endpoints are removed from DNS responses. Without health checks, Route 53 would continue directing traffic to a failed Region, defeating the resilience goal.

Exam trap

The trap is that candidates assume latency-based routing automatically fails over when a Region goes down, when in fact Route 53 requires explicitly configured health checks to remove unhealthy endpoints from DNS responses.

How to eliminate wrong answers

Option A is wrong because scheduled scaling adjusts capacity based on time, not health or demand, and does nothing to redirect traffic away from a failed Region. Option B is wrong because DAX is an in-memory cache for DynamoDB that improves read latency but does not provide cross-Region failover or traffic redirection. Option C is wrong because CloudFront caches content at edge locations but does not perform origin health-based failover across Regions unless paired with origin groups and health checks — and it does not replace Route 53 health checks for latency-based routing.

Page 6

Page 7 of 18

Page 8