Courseiva

CCNA Incident and Event Response Questions

33 of 183 questions · Page 3/3 · Incident and Event Response · Answers revealed

151
MCQeasy

A DevOps engineer receives an alert that an EC2 instance's CPU utilization has been above 90% for the last 30 minutes. The engineer needs to investigate the root cause. Which AWS service should the engineer use to get OS-level process details and identify which process is consuming the CPU?

A.AWS Config
B.AWS CloudTrail
C.AWS Systems Manager Run Command
D.Amazon CloudWatch
AnswerC

AWS Systems Manager Run Command is part of AWS Systems Manager and lets you remotely and securely execute shell commands or PowerShell scripts on EC2 instances (and on-premises machines) via the SSM Agent. You can run a command like `ps aux` or `Get-Process` to enumerate running processes, capture output, and store it in S3 or CloudWatch Logs. Because the SSM Agent runs inside the instance as a guest process, it has direct access to OS-level state, making it the appropriate service for collecting process-level data. It also supports rate control and error handling for fleet-wide execution.

Why this answer

AWS Systems Manager Run Command allows you to run commands (e.g., 'top', 'ps') remotely on EC2 instances to obtain OS-level process details and identify which process is consuming CPU. Option A is wrong because AWS Config records configuration changes, not OS-level processes. Option B is wrong because AWS CloudTrail logs API calls, not system-level metrics.

Option D is wrong because Amazon CloudWatch provides aggregated CPU utilization metrics but cannot provide process-level details.

152
MCQmedium

A company runs a microservices application on Amazon ECS with Fargate. The application includes a service that processes messages from an Amazon SQS queue. Recently, the processing time has increased, and the SQS queue depth is growing. The CloudWatch metrics show that the ECS service's CPU utilization is consistently around 70%, memory utilization is 80%, and the number of running tasks is at the maximum allowed (10). The service is configured with a target tracking scaling policy based on CPU utilization with a target value of 50%. However, the auto scaling does not seem to be adding tasks. The engineer checks the ECS service events and finds no scaling activity. What is the MOST likely reason the auto scaling is not working, and what action should be taken to resolve the issue?

A.The service has reached the maximum number of tasks defined in the auto scaling configuration; increase the maximum tasks.
B.The scaling policy is not properly configured; recreate it with a lower target value.
C.The CloudWatch metric is not being emitted correctly; check the metric namespace.
D.The ECS service is using Fargate, which does not support target tracking scaling policies.
AnswerA

The ECS service's Application Auto Scaling target tracking policy is capped by a maximum capacity of 10 tasks. Since the service is already running at that desired count of 10, the policy cannot scale out any further even though the CloudWatch metric is above the target value. To resolve the issue, increase the maximum tasks (for example, from 10 to 15) in the scaling configuration, which allows the policy to add more task instances.

Why this answer

The auto scaling is not adding tasks because the ECS service has already reached the maximum number of tasks defined in the auto scaling configuration (10). With CPU utilization at 70% and the target tracking policy set to 50%, the policy would normally trigger scale-out actions, but since the maximum task count is already hit, no scaling activity occurs. The engineer must increase the maximum tasks in the auto scaling configuration to allow further scale-out.

Exam trap

The trap here is that candidates may assume the scaling policy itself is misconfigured or that Fargate lacks support for target tracking, when in reality the issue is the hard cap on the maximum number of tasks preventing any scale-out action.

How to eliminate wrong answers

Option B is wrong because the target value of 50% is appropriate; lowering it would not resolve the issue since the policy is not being triggered due to the max task limit, not the target value. Option C is wrong because CloudWatch metrics are being emitted correctly (CPU utilization is visible at 70%), so the metric namespace is not the problem. Option D is wrong because Fargate fully supports target tracking scaling policies for ECS services; this is a supported and common configuration.

153
MCQeasy

A company uses CloudWatch Logs to store application logs. The security team requires that logs be encrypted at rest using a customer-managed KMS key. What must be done to enable this?

A.Enable encryption on the log group using the default AWS managed key.
B.Use a third-party encryption tool before sending logs to CloudWatch.
C.Create a new log group in a region where KMS is enabled.
D.Associate a customer-managed KMS key with the log group and update the key policy to allow CloudWatch Logs to use it.
AnswerD

To encrypt a CloudWatch Logs log group with a customer-managed KMS key, you must create or select a KMS key, update its key policy to grant the CloudWatch Logs service principal permissions such as kms:Encrypt, kms:Decrypt, kms:GenerateDataKey*, and kms:DescribeKey, and then associate that key with the log group using the console or the associate-kms-key API. This server-side encryption protects log data at rest and allows CloudWatch Logs to decrypt the data internally for features like log queries.

Why this answer

CloudWatch Logs supports encryption at rest using a customer-managed KMS key. To enable this, you must associate the KMS key with the log group via the CloudWatch Logs console or API, and you must update the key policy to grant CloudWatch Logs the necessary permissions (kms:Encrypt, kms:Decrypt, kms:ReEncrypt*, kms:GenerateDataKey*, and kms:DescribeKey). Without this key policy update, CloudWatch Logs cannot use the key to encrypt the log data at rest.

Exam trap

The trap here is that candidates often assume encryption is automatically applied when a KMS key exists in the account, but they overlook the critical step of updating the key policy to grant CloudWatch Logs service principal permissions to use the key.

How to eliminate wrong answers

Option A is wrong because using the default AWS managed key does not meet the security team's requirement for a customer-managed KMS key; the default key is AWS-owned and not customer-managed. Option B is wrong because using a third-party encryption tool before sending logs to CloudWatch would result in encrypted log data that CloudWatch Logs cannot index, search, or process natively, defeating the purpose of centralized logging. Option C is wrong because KMS is available in all AWS regions where CloudWatch Logs is supported; creating a new log group in a different region does not enable customer-managed KMS encryption—you must explicitly associate a customer-managed key with the log group.

154
MCQmedium

A company uses AWS Systems Manager Patch Manager to patch EC2 instances. During a patching window, some instances fail to apply patches. The engineer checks the SSM Agent logs and sees 'ERROR: Failed to download patch files from the source.' What is the most likely cause?

A.The IAM instance profile does not grant ssm:UpdateInstanceInformation.
B.The SSM Agent is outdated.
C.The patch baseline is configured incorrectly.
D.The security group or NACL is blocking outbound HTTPS traffic (port 443).
AnswerD

The SSM Agent downloads patch binaries directly from the configured patch repositories, such as Windows Update or Linux package mirrors, using HTTPS on TCP port 443. If the instance's security group or subnet NACL blocks outbound HTTPS, the agent cannot reach these repositories, leading to a patch download failure. This is the most common network-level cause, especially when the instance can otherwise communicate with the Systems Manager API but fails specifically during patch retrieval.

Why this answer

The error 'Failed to download patch files from the source' indicates that the SSM Agent on the instance cannot reach the patch source repositories (e.g., Windows Update, Amazon Linux repos, or custom patch sources). Systems Manager Patch Manager requires outbound HTTPS (port 443) connectivity to download patch metadata and binaries. If a security group or NACL blocks this traffic, the download fails, producing this exact error in the agent logs.

Exam trap

The trap here is that candidates often assume the error is due to IAM permissions or patch baseline misconfiguration, overlooking that the specific 'Failed to download' message is a classic symptom of network egress blocking, not authorization or configuration issues.

How to eliminate wrong answers

Option A is wrong because ssm:UpdateInstanceInformation is required for the instance to register and send heartbeat data to Systems Manager, but it does not control the ability to download patch files; the error is about download failure, not registration. Option B is wrong because an outdated SSM Agent would typically produce errors about agent version incompatibility or missing features, not a specific 'Failed to download patch files from the source' message; the agent can still attempt downloads. Option C is wrong because a misconfigured patch baseline might cause patches to be incorrectly approved or rejected, but the error message points to a network connectivity issue preventing download, not a baseline configuration problem.

155
MCQeasy

An application running on Amazon RDS for PostgreSQL is experiencing slow query performance. The DevOps team suspects a specific query is causing high CPU usage. Which tool should they use to identify the problematic query?

A.Amazon RDS Event Subscriptions
B.Amazon RDS Performance Insights
C.Amazon CloudWatch Logs with metric filters
D.Amazon RDS Enhanced Monitoring
AnswerB

Amazon RDS Performance Insights gives you an at-a-glance dashboard of database load (measured in average active sessions), broken down by wait events, SQL statements, hosts, and users. This allows you to see exactly which query is consuming the most resources and why—whether it is blocked on locks, I/O, or CPU. Its historical and real-time views enable you to pinpoint and correlate the slow query with the performance problem, making it the correct choice for this scenario.

Why this answer

Amazon RDS Performance Insights is the correct tool because it provides a database performance tuning and monitoring feature that visualizes database load and lets you identify specific queries causing high CPU usage. It breaks down the database load by wait events, SQL statements, hosts, and users, making it straightforward to pinpoint the problematic query. This directly addresses the DevOps team's need to find the query responsible for high CPU consumption.

Exam trap

The trap here is that candidates often confuse Enhanced Monitoring (OS-level metrics) with Performance Insights (database-level query analysis), assuming OS metrics alone can identify the specific SQL query, but Enhanced Monitoring lacks the query-level granularity needed to pinpoint the problematic statement.

How to eliminate wrong answers

Option A is wrong because Amazon RDS Event Subscriptions are used to receive notifications about database events (e.g., failover, maintenance) and do not provide query-level performance data or CPU usage analysis. Option C is wrong because Amazon CloudWatch Logs with metric filters can monitor log data and extract metrics, but they require custom log parsing and do not natively surface database query performance or CPU load by SQL statement. Option D is wrong because Amazon RDS Enhanced Monitoring provides OS-level metrics (e.g., CPU, memory, disk I/O) but does not identify which specific SQL queries are causing high CPU usage.

156
MCQmedium

A company is using AWS CloudFormation to deploy infrastructure. An engineer needs to ensure that any changes to the production stack are reviewed and approved before they are applied. The engineer also wants to prevent unauthorized changes. Which solution should the engineer implement?

A.Use CloudFormation StackSets to manage the production stack across multiple accounts.
B.Use CloudFormation Change Sets and require manual approval to execute the change set.
C.Use AWS Service Catalog to create a product for the stack and require approval for any portfolio changes.
D.Use AWS CodePipeline to deploy the stack and require manual approval at the deploy stage.
AnswerB

Change Sets generate a preview of proposed resource modifications without applying them, and the manual approval step gates execution until reviewers accept. This satisfies the requirement for review before production changes and blocks unauthorised updates.

Why this answer

CloudFormation Change Sets allow you to preview how proposed changes to a stack will impact running resources before you apply them. By requiring manual approval to execute the change set, the engineer ensures that all modifications are reviewed and approved, preventing unauthorized changes. This directly meets the requirement for a review-and-approval workflow without introducing unnecessary complexity.

Exam trap

The trap here is that candidates often confuse the purpose of StackSets (multi-account deployment) or CodePipeline (CI/CD pipeline) with the need for a simple change review mechanism, overlooking the direct and built-in capability of CloudFormation Change Sets to preview and require approval before applying changes.

How to eliminate wrong answers

Option A is wrong because CloudFormation StackSets are designed to deploy stacks across multiple accounts and regions, not to enforce a review-and-approval workflow for changes to a single production stack. Option C is wrong because AWS Service Catalog products and portfolio changes control the provisioning of pre-defined templates, not the approval of changes to an already-deployed stack; it does not provide a change review mechanism for existing stacks. Option D is wrong because while CodePipeline can include a manual approval stage, it is a CI/CD orchestration tool that adds unnecessary overhead and complexity for a simple change review requirement; CloudFormation Change Sets provide a more direct and lightweight solution.

157
Multi-Selectmedium

A company is experiencing a DDoS attack on their web application hosted on Amazon EC2 behind an Application Load Balancer (ALB). The attack is causing high CPU utilization on the instances. The security team needs to mitigate the attack with minimal disruption to legitimate users. Which TWO actions should the team take? (Choose two.)

Select 2 answers
A.Configure AWS WAF rate-based rules to block excessive requests from specific IP addresses.
B.Enable AWS Shield Advanced on the ALB for additional DDoS protection.
C.Enable VPC Flow Logs to analyze traffic patterns and identify the source of the attack.
D.Scale up the EC2 instances by increasing their instance size.
E.Place an Amazon CloudFront distribution in front of the ALB to cache content.
AnswersA, C

AWS WAF rate-based rules track the number of requests from each client IP over a rolling evaluation window (typically 5 minutes) and, when a configured threshold is exceeded, automatically block that IP for a specified duration. This provides immediate, in-place mitigation on the existing ALB without DNS changes, making it the fastest L7 DDoS countermeasure. You can tune the rate limit and scope (e.g., by URI or session) to allow legitimate bursts while dropping attack traffic.

Why this answer

AWS WAF rate-based rules are designed to automatically block IP addresses that exceed a specified request rate, which directly mitigates DDoS attacks by limiting excessive traffic from specific sources. This approach minimizes disruption to legitimate users because it only blocks IPs that exceed the threshold, preserving access for normal traffic patterns.

Exam trap

The trap here is that candidates often confuse AWS Shield Advanced as a direct mitigation for application-layer DDoS attacks, when it primarily protects against infrastructure-layer attacks (e.g., SYN floods) and requires WAF for application-layer control.

158
Multi-Selecteasy

A DevOps engineer needs to receive notifications when an EC2 instance's status check fails. Which TWO services should the engineer use? (Choose TWO.)

Select 2 answers
A.AWS Lambda
B.Amazon Simple Notification Service (SNS)
C.AWS CloudTrail
D.AWS Config
E.Amazon CloudWatch Alarm
AnswersB, E

Amazon Simple Notification Service (SNS) is a fully managed pub/sub messaging service that delivers messages to subscribers such as email, SMS, Lambda, or HTTP endpoints. When a CloudWatch alarm transitions to the ALARM state for the StatusCheckFailed metric, it publishes a message to an SNS topic, which then fans out the notification to all subscribed endpoints. This makes SNS the essential delivery mechanism for alerting the DevOps engineer promptly without requiring polling or custom integration code, and it integrates seamlessly with CloudWatch Alarms.

Why this answer

Amazon CloudWatch Alarms (Option E) can monitor EC2 instance status checks (both system and instance checks) and trigger an action when the alarm state changes to ALARM. Amazon SNS (Option B) is the service that delivers the notification by publishing messages to subscribers (e.g., email, SMS, HTTP endpoints) when the CloudWatch alarm triggers. Together, they provide a complete monitoring and notification pipeline for status check failures.

Exam trap

The trap here is that candidates often select AWS Lambda or AWS Config because they associate them with automation or compliance, but the question explicitly asks for services to 'receive notifications' when a status check fails, which requires a notification delivery service (SNS) and a monitoring service (CloudWatch Alarm), not compute or configuration tracking.

159
Multi-Selectmedium

A company runs a production database on Amazon RDS for MySQL. The database experiences a sudden spike in connections, causing the application to time out. The DevOps team needs to diagnose the issue quickly. Which combination of actions should be taken? (Choose two.)

Select 2 answers
A.Check CloudWatch metrics for DatabaseConnections and CPUUtilization.
B.Immediately scale up the RDS instance to handle the load.
C.Analyze VPC Flow Logs to identify the source IPs of connections.
D.Use the RDS console to view the number of active connections per user.
E.Enable Performance Insights and review the top SQL statements.
AnswersA, E

This is the correct first step because CloudWatch provides two directly relevant metrics for an RDS for MySQL instance: DatabaseConnections shows the number of client sessions currently established, and CPUUtilization reflects aggregate CPU consumption. If DatabaseConnections spikes while CPUUtilization remains normal, the problem is connection exhaustion or a connection leak; if CPUUtilization also rises, it likely indicates heavy query load. These metrics give you a high-level, time-aligned picture of whether the symptom is connection-bound or compute-bound, which drives the next diagnostic step (e.g., enabling Performance Insights).

Why this answer

Option A is correct because Amazon CloudWatch publishes RDS metrics such as DatabaseConnections (the current number of client sessions) and CPUUtilization, which directly reveal whether the timeout is caused by connection saturation or CPU exhaustion on the instance. Option E is correct because Performance Insights, once enabled, provides a database load view broken down by wait events and top SQL statements, letting the team pinpoint which queries or sessions are consuming resources during the spike. Option B is not appropriate as a diagnostic step because scaling up is a remediation action taken after the root cause is identified, and it may not resolve connection-limit or query-level problems.

Option C is wrong because VPC Flow Logs capture IP-level network traffic metadata, not database session counts or query behavior, so they cannot explain an RDS connection spike. Option D is wrong because the RDS console does not provide a per-user active connection breakdown; that level of session detail requires querying performance_schema or using Performance Insights.

Exam trap

DOP-C02 often tests the difference between diagnosis and remediation — candidates pick 'scale up immediately' because it feels proactive, but the question asks how to diagnose quickly, not how to fix.

160
MCQeasy

A DevOps engineer is investigating a security incident where an EC2 instance was used to launch an outbound DDoS attack. Which AWS service can provide details about the source IP addresses and network traffic from the instance?

A.VPC Flow Logs
B.AWS CloudTrail
C.Amazon GuardDuty
D.AWS Config
AnswerA

VPC Flow Logs are correct because they capture IP traffic metadata at the elastic network interface level, including source/destination IP addresses, ports, protocol, and the accept/reject action for each flow. This data directly shows which connections were made, from where, and to which destination, enabling a DevOps engineer to trace the communication patterns involved in the incident. These logs can be enabled per VPC, subnet, or ENI and delivered to CloudWatch Logs or S3 for analysis.

Why this answer

VPC Flow Logs capture IP traffic metadata for network interfaces in a VPC, including source and destination IP addresses, ports, protocol, and accept/reject status. For an EC2 instance launching an outbound DDoS attack, flow logs provide the source IPs and traffic patterns needed to identify the attack destinations and volume. This is the service designed for network-level traffic visibility.

Exam trap

DOP-C02 often tests the distinction between network telemetry (VPC Flow Logs) and API audit (CloudTrail) — the trap is choosing CloudTrail because it 'logs activity' when the question asks for source IPs and network traffic.

How to eliminate wrong answers

Option B is wrong because AWS CloudTrail records API activity (who called which AWS API), not network packet flows, so it cannot show source IPs of outbound traffic. Option C is wrong because Amazon GuardDuty is a threat-detection service that generates findings; it can alert on malicious activity but does not provide the raw source IP and traffic detail the question asks for. Option D is wrong because AWS Config tracks resource configuration changes and compliance, not network traffic metadata.

161
MCQeasy

A DevOps engineer receives a CloudWatch alarm indicating that an EC2 instance's CPU utilization has exceeded 90% for 10 minutes. The instance is part of an Auto Scaling group behind an Application Load Balancer. What is the MOST efficient initial step to troubleshoot the high CPU usage?

A.Review the EC2 instance's CloudWatch metrics for CPU credit balance and network utilization.
B.Modify the Auto Scaling group to use a larger instance type.
C.Check the ALB's HTTP 5xx error rate metric for the target group.
D.Immediately increase the desired capacity of the Auto Scaling group.
AnswerA

Reviewing the instance's CPU credit balance is the correct first step because a T-series instance that exhausts its earned credits will be throttled to the baseline CPU, causing slowdowns even when the alarm threshold is breached. The network utilization metric is also essential since high network throughput can drive CPU overhead from interrupt handling and packet processing, which may be the actual root cause. Together, these metrics reveal whether the alarm reflects genuine resource contention or a misconfigured alarm threshold.

Why this answer

Reviewing the EC2 instance's CloudWatch metrics for CPU credit balance and network utilization is the most efficient initial step to diagnose high CPU usage. CPU credit balance is critical for burstable performance instances (e.g., T2/T3), as a depleted credit balance directly causes sustained high CPU utilization. Network utilization metrics can reveal if the high CPU is driven by excessive traffic or a DDoS-like pattern, allowing targeted remediation without unnecessary scaling or configuration changes.

Exam trap

The trap here is that candidates often jump to scaling actions (Options B or D) or application-layer metrics (Option C) without first checking the instance's foundational health metrics, specifically CPU credit balance for burstable instances, which is the most efficient diagnostic step per AWS Well-Architected best practices.

How to eliminate wrong answers

Option B is wrong because modifying the Auto Scaling group to use a larger instance type is a reactive, long-term solution that does not diagnose the root cause and may incur unnecessary cost; the immediate need is to understand why CPU is high, not to blindly resize. Option C is wrong because checking the ALB's HTTP 5xx error rate metric for the target group focuses on application-layer errors, which are a symptom of high CPU but do not reveal the underlying cause (e.g., CPU credit exhaustion, process spike, or network saturation). Option D is wrong because immediately increasing the desired capacity of the Auto Scaling group is a scaling action that treats the symptom (high CPU) without investigating the cause, potentially leading to over-provisioning or masking a deeper issue like a memory leak or misconfigured application.

162
Multi-Selecthard

A company experiences a security incident where an IAM user's access key is compromised. Which THREE steps should the DevOps engineer take immediately?

Select 3 answers
A.Review AWS CloudTrail logs for any unauthorized API calls
B.Rotate the access key by creating a new key and deleting the old one
C.Change the IAM user's password
D.Delete the IAM user and recreate it
E.Revoke any temporary security credentials issued to the user
AnswersA, B, E

AWS CloudTrail records all IAM user and role API activity as events, including the source IP address, user agent, event name, and whether the call was authorized. Analyzing CloudTrail logs with querying or visualization tools helps you identify which unauthorized or anomalous calls occurred, when they occurred, and what resources were accessed, so you can scope the impact and determine whether other remediation steps are needed. It is the first responder's forensic tool for understanding the security incident.

Why this answer

Option A is correct because AWS CloudTrail records all API activity in the account, so reviewing its logs lets the engineer identify unauthorized calls made with the compromised access key and assess the incident's scope. Option B is correct because rotating the access key—creating a new key pair and deleting the compromised key—immediately invalidates the leaked credential so it can no longer be used for API authentication. Option E is correct because if the compromised IAM user had assumed roles or obtained temporary credentials via STS, those session tokens remain valid until expiry, so they must be explicitly revoked (e.g., by attaching a deny-all policy or using AWSRevokeOlderSessions) to cut off the attacker's access.

Option C does not belong because changing the IAM user's console password does not affect the compromised access key, which is used for programmatic API calls rather than console sign-in. Option D does not belong because deleting and recreating the IAM user is a disruptive, unnecessary step that would break existing permissions and resource associations; rotating the key and revoking sessions is sufficient to contain the incident.

Exam trap

DOP-C02 often tests whether candidates conflate console password compromise with access key compromise — the trap is choosing 'change the password' when the credential at risk is the API key.

163
MCQeasy

A company has a legacy application running on an EC2 instance that is not part of an Auto Scaling group. The instance is experiencing a memory leak. The DevOps engineer needs to collect memory metrics to analyze the issue without modifying the application. What should the engineer do?

A.Install the CloudWatch agent on the instance and configure it to collect memory metrics.
B.Use the AWS Management Console to view memory metrics from the EC2 monitoring tab.
C.Use EC2Rescue to generate a memory dump and analyze it.
D.Enable CloudWatch detailed monitoring on the instance.
AnswerA

The default EC2 monitoring only exposes hypervisor-level metrics like CPU, network, and disk I/O; memory utilization is a guest-OS metric that AWS cannot see without an in-guest component. Installing the unified CloudWatch agent (with the `mem_used_percent` and similar metrics in the agent's JSON config) and starting the `amazon-cloudwatch-agent` service enables the agent to publish memory metrics to CloudWatch, making them available for alarms and dashboards. This is required because no amount of instance-level monitoring settings can surface guest-OS memory.

Why this answer

The CloudWatch agent is required to collect custom metrics like memory utilization from an EC2 instance because the standard EC2 monitoring only captures hypervisor-level metrics (CPU, network, disk I/O). By installing and configuring the CloudWatch agent, the engineer can collect memory metrics without modifying the application code, directly addressing the memory leak analysis requirement.

Exam trap

The trap here is that candidates often assume the EC2 monitoring tab or detailed monitoring includes memory metrics, but AWS does not provide OS-level metrics (memory, disk space, swap usage) without the CloudWatch agent.

How to eliminate wrong answers

Option B is wrong because the AWS Management Console EC2 monitoring tab only displays default metrics (CPU, network, disk, status checks) and does not include memory metrics, which require a custom agent. Option C is wrong because EC2Rescue is a tool for troubleshooting and repairing common EC2 issues (e.g., OS boot failures, disk corruption), not for collecting ongoing memory metrics; it can generate a memory dump but that is a one-time snapshot, not a continuous metric stream for trend analysis. Option D is wrong because enabling CloudWatch detailed monitoring only increases the frequency of default metric collection (from 5 minutes to 1 minute) but does not add memory metrics, which are not available at the hypervisor level.

164
MCQeasy

A company uses Amazon RDS for MySQL as its database. The operations team notices that the database CPU utilization is consistently above 90% during peak hours, causing slow query responses. The team needs to quickly reduce CPU load without changing the application code. Which action should the team take?

A.Enable Multi-AZ deployment.
B.Modify the DB parameter group to increase max_connections.
C.Add a read replica to offload read traffic.
D.Enable Performance Insights and analyze the top queries.
AnswerD

Performance Insights delivers a comprehensive, real-time view of database load, breaking down utilization by waits, SQL statement, and host. By drilling into the 'Top SQL' section, you can pinpoint the exact queries consuming the most CPU, along with statistics such as rows examined and temp tables. This evidence-based approach allows you to optimize indexes or rewrite expensive statements, directly addressing the observed CPU spike.

Why this answer

Enabling Performance Insights allows the team to identify the specific queries that are consuming CPU resources. By analyzing these top queries, the team can take targeted actions such as optimizing queries or adding indexes to reduce CPU load without changing application code. Option A is incorrect because Multi-AZ deployment provides high availability and failover support but does not reduce CPU utilization.

Option B is incorrect because increasing max_connections allows more concurrent connections, which can actually increase CPU load rather than reduce it. Option C is incorrect because adding a read replica offloads read traffic but does not reduce CPU load on the primary instance, and typically requires application changes to route read queries to the replica.

165
MCQmedium

A company is using AWS Lambda to process events from an Amazon SQS queue. The Lambda function is configured with a batch size of 10 and a maximum concurrency of 5. Recently, the function started experiencing high error rates and the SQS queue's ApproximateNumberOfMessagesVisible metric is increasing. The CloudWatch logs show that the function is timing out after 30 seconds. The function makes calls to an external API that sometimes takes more than 30 seconds to respond. The DevOps engineer needs to reduce the backlog and prevent message loss. The engineer is considering the following actions: A) Increase the Lambda function timeout to 60 seconds and increase the SQS visibility timeout to 90 seconds. B) Decrease the batch size to 1 to avoid processing multiple messages at once. C) Increase the Lambda function reserved concurrency to 100 to allow more concurrent executions. D) Use a dead-letter queue to capture messages that fail processing after all retries. Which combination of actions should the engineer take?

A.Use a dead-letter queue to capture messages that fail processing after all retries.
B.Decrease the batch size to 1 to avoid processing multiple messages at once.
C.Increase the Lambda function timeout to 60 seconds and increase the SQS visibility timeout to 90 seconds.
D.Increase the Lambda function reserved concurrency to 100 to allow more concurrent executions.
AnswerC

This is correct because an SQS-triggered Lambda invocation has a maximum execution window set by the function timeout, and the SQS visibility timeout controls when unacknowledged messages become visible again for redelivery. If the function timeout is too short, valid work gets aborted, and if the visibility timeout is shorter than the processing time, the message is re-delivered before the first attempt finishes, causing duplicate work and retries that inflate the backlog. Setting the visibility timeout to 90 seconds (longer than the 60-second function timeout) ensures the message stays hidden until the Lambda function either succeeds or itself times out, giving the function the full time it needs.

Why this answer

The correct action because increasing the Lambda function timeout to 60 seconds allows the function to wait longer for the external API, and increasing the SQS visibility timeout to 90 seconds prevents messages from becoming visible again before the function completes. This reduces unnecessary retries and helps clear the backlog. Option A (DLQ) is useful for capturing failed messages but does not address the timeout issue.

Option B (decrease batch size) reduces throughput and worsens the backlog. Option D (increase concurrency) may lead to more timeouts if the function still cannot complete within the existing timeout.

166
Multi-Selecthard

A company has a multi-account AWS organization. The security team needs to detect and respond to security incidents across all accounts centrally. Which THREE services should the team use together? (Choose three.)

Select 3 answers
A.AWS Security Hub
B.Amazon Inspector
C.Amazon Macie
D.Amazon GuardDuty
E.Amazon Detective
AnswersA, D, E

AWS Security Hub is the correct answer because it is designed as a multi-account, multi-region aggregation service that centralizes security findings from AWS services and partner products. It enables a delegated administrator to view a consolidated security posture across the entire AWS Organizations hierarchy, evaluate compliance against standards like CIS and NIST, and automate responses via custom actions and AWS Config rules.

Why this answer

AWS Security Hub (A) is correct because it acts as the central aggregation and prioritization layer, ingesting findings from GuardDuty, Inspector, Macie, and other services across all accounts in the organization and normalizing them to the AWS Security Finding Format (ASFF) for a single-pane-of-glass view. Amazon GuardDuty (D) is correct because it provides the continuous threat detection foundation, analyzing CloudTrail management events, VPC Flow Logs, and DNS logs across every account to surface malicious or unauthorized activity. Amazon Detective (E) is correct because it complements detection with investigation, automatically building behavior graphs from GuardDuty, CloudTrail, and VPC Flow Logs so the security team can triage and root-cause incidents centrally.

Amazon Inspector (B) is not one of the three because it is a vulnerability management service scoped to EC2, ECR, and Lambda resources rather than a cross-account incident detection and response hub. Amazon Macie (C) is not one of the three because it is a data-security service that discovers and classifies sensitive data in S3, which is useful but not part of the core centralized detect-and-respond trio.

Exam trap

DOP-C02 often tests the distinction between detection/aggregation/investigation services (GuardDuty, Security Hub, Detective) and specialized scanning services (Inspector for vulnerabilities, Macie for sensitive data) — candidates include Inspector or Macie thinking 'security' means all of them, but the question asks for the central detection-and-response trio.

167
MCQhard

A company runs a web application on Amazon EC2 instances behind an Application Load Balancer (ALB). The application experiences intermittent 503 errors. The engineer suspects the ALB is returning these errors because the target instances are unhealthy. Which metric should the engineer monitor to confirm this suspicion?

A.RequestCount
B.UnhealthyHostCount
C.HealthyHostCount
D.TargetResponseTime
AnswerC

HealthyHostCount directly reflects the number of targets that are passing the configured health checks. For a target group with EC2 instances, it is the authoritative CloudWatch metric for monitoring target health; when all instances fail, this value drops to zero, leaving the ALB with no registered and healthy targets to serve traffic, which causes it to return HTTP 503 Service Unavailable. Therefore, it is the correct metric to identify the described incident.

Why this answer

The ALB publishes 'HealthyHostCount' metric showing the number of healthy targets. When this count drops to zero, the ALB cannot forward requests and returns 503 errors. Option A (RequestCount) is incorrect because it measures total requests, not health.

Option B (UnhealthyHostCount) is a valid metric but does not directly confirm the suspicion that targets are unhealthy; a decreasing HealthyHostCount is more direct. Option D (TargetResponseTime) measures latency, not health status.

168
MCQmedium

An organization uses AWS Systems Manager to manage its EC2 instances. After a security incident, the security team wants to ensure that all future API calls to Systems Manager are logged and monitored. What is the MOST efficient way to achieve this?

A.Enable S3 server access logging on the Systems Manager log bucket
B.Enable AWS CloudTrail for the Systems Manager service
C.Install the CloudWatch Logs agent on each instance to capture Systems Manager logs
D.Create an AWS Config rule to monitor Systems Manager usage
AnswerB

AWS CloudTrail is the authoritative service for auditing API calls, and it natively records Systems Manager management events such as SendCommand, RunCommand, and StartSession. When a trail is enabled (or via the default event history), each event includes the IAM principal, source IP address, event time, request parameters, and response elements, yielding a complete 'who did what' record. This is precisely what is needed to audit and govern SSM usage across an EC2 fleet.

Why this answer

Enabling CloudTrail for Systems Manager logs all API calls made to the Systems Manager service. Option A is incorrect because S3 server access logging only logs access to S3 buckets, not Systems Manager API calls. Option C is incorrect because the CloudWatch Logs agent captures instance logs, not API calls to Systems Manager.

Option D is incorrect because AWS Config rules track configuration changes, not API calls. Therefore, CloudTrail is the most efficient way to log and monitor all future API calls to Systems Manager.

169
MCQhard

A company uses AWS Lambda functions behind an Amazon API Gateway REST API. During an incident, the API returns 502 Bad Gateway errors. The Lambda function logs show no errors. What is the most likely cause?

A.The Lambda function is throwing an unhandled exception
B.The Lambda function is returning a response that exceeds the API Gateway payload size limit
C.The API Gateway has reached its maximum concurrency limit
D.The Lambda function is timing out and API Gateway is not handling the timeout correctly
AnswerB

API Gateway imposes a hard 10 MB payload size limit for REST API responses (and 4 MB for HTTP APIs). When a Lambda function returns a response exceeding this threshold, API Gateway cannot process it and returns a 502 Bad Gateway error to the client, with no error logged from the Lambda side because the function already completed successfully. This is the classic 'silent' 502 cause and matches the symptoms in the question.

Why this answer

When an API Gateway REST API returns 502 Bad Gateway errors but the Lambda function logs show no errors, the most likely cause is that the Lambda function is returning a response that exceeds the API Gateway payload size limit. API Gateway has a maximum payload size of 10 MB for REST APIs, and if the Lambda function returns a response larger than this, API Gateway will reject it and return a 502 error without the Lambda function itself throwing an exception or logging an error.

Exam trap

AWS often tests the distinction between different HTTP status codes (502 vs 504 vs 429) and the specific conditions under which each is returned, leading candidates to incorrectly attribute 502 errors to Lambda timeouts or API Gateway throttling instead of payload size limits.

How to eliminate wrong answers

Option A is wrong because an unhandled exception in the Lambda function would cause the function to fail and log an error in Amazon CloudWatch Logs, but the question states that the Lambda function logs show no errors. Option C is wrong because API Gateway does not have a maximum concurrency limit; it scales automatically, and reaching a concurrency limit would result in 429 Too Many Requests errors, not 502 Bad Gateway errors. Option D is wrong because if the Lambda function were timing out, the Lambda service would log a timeout error in CloudWatch Logs, and API Gateway would typically return a 504 Gateway Timeout error, not a 502 Bad Gateway error.

170
Multi-Selectmedium

A DevOps engineer is designing an incident response plan for a multi-region application. The application runs on EC2 instances behind an Application Load Balancer (ALB) and uses Amazon RDS for MySQL with Multi-AZ. Which TWO actions should the engineer include to ensure high availability and fast failover during a regional incident?

Select 2 answers
A.Set up an Amazon RDS read replica in a second region and promote it during failover.
B.Create an Auto Scaling group that can launch instances in multiple regions.
C.Deploy an Application Load Balancer that spans both regions.
D.Configure Amazon RDS Multi-AZ in a second region.
E.Use Amazon Route 53 with health checks to fail over DNS to a secondary region.
AnswersA, E

For a true cross-region disaster recovery, an Amazon RDS read replica in a second region can be promoted to a standalone primary instance during failover. The replica is typically created with asynchronous replication, so some data loss may occur, but it provides a writable database endpoint in the secondary region. This is a standard DR pattern for maintaining business continuity.

Why this answer

Options A and E are correct. Amazon RDS read replicas can be created in a different region and promoted to a standalone primary instance during a regional incident, providing a disaster recovery solution. Amazon Route 53 with health checks can automatically fail over DNS traffic to a secondary region by routing to a healthy endpoint.

Option B is incorrect because Auto Scaling groups are regional and cannot launch instances across multiple regions directly; you would need separate Auto Scaling groups per region. Option C is incorrect because an Application Load Balancer is regional and cannot span regions; you would need separate ALBs in each region and Route 53 to route traffic. Option D is incorrect because Multi-AZ RDS replicates synchronously within a single region only; for cross-region disaster recovery, you need a read replica or a separate Multi-AZ deployment in the other region, not Multi-AZ in a second region (Multi-AZ in a second region is not a standard feature; you would need a separate RDS instance).

171
MCQeasy

A DevOps engineer notices that an EC2 instance running a web application is unresponsive. CloudWatch alarms are not triggering. What is the FIRST step the engineer should take to diagnose the issue?

A.Terminate the instance and launch a new one from the latest AMI.
B.Review the EC2 instance system log and CloudWatch Logs for error messages.
C.Restart the EC2 instance immediately to restore service.
D.Create a new CloudWatch alarm with a lower threshold to get alerted quicker next time.
AnswerB

The EC2 system log (console output) is a hypervisor-accessible snapshot of the instance's serial port, capturing kernel panics, OOM killer events, and boot-time failures that may be invisible from inside the OS. Pairing that with CloudWatch Logs—where the CloudWatch agent streams Apache, Nginx, or custom application errors—gives you a non-disruptive, evidence-based starting point to pinpoint whether the web service stopped due to memory exhaustion, a crashed process, or an external dependency. These sources are available via the EC2 console or the get-console-output CLI call and require no downtime, making them the correct first step for diagnosis.

Why this answer

When an EC2 instance is unresponsive but CloudWatch alarms are not triggering, the first diagnostic step is to check the instance system log (console output) and CloudWatch Logs for error messages. This approach follows the principle of gathering evidence before taking action, as the logs may reveal application crashes, kernel panics, or resource exhaustion that caused the unresponsiveness without breaching CloudWatch alarm thresholds.

Exam trap

The trap here is that candidates often jump to immediate remediation (restart or replace) instead of following the incident response process of first gathering diagnostic data from logs and system output.

How to eliminate wrong answers

Option A is wrong because terminating the instance destroys forensic evidence and prevents root cause analysis; the correct first step is to diagnose, not destroy. Option C is wrong because restarting the instance without investigation may temporarily restore service but loses volatile diagnostic data (e.g., memory dumps, process states) and does not address the underlying issue. Option D is wrong because creating a new alarm with a lower threshold does not help diagnose the current unresponsive instance; it only changes future alerting behavior and does not provide any immediate diagnostic information.

172
MCQeasy

A company uses AWS CloudFormation to manage infrastructure. During an incident, a stack update fails with the error 'The following resource(s) failed to create: [AWS::RDS::DBInstance]'. Which AWS service should the engineer use to view detailed error messages for the failed resource creation?

A.AWS Config timeline
B.AWS CloudFormation console Events tab
C.AWS Service Catalog
D.AWS CloudTrail event history
AnswerB

The CloudFormation console Events tab lists every operation performed on a stack in chronological order, one per resource action, including statuses like CREATE_FAILED and UPDATE_FAILED. For each failed event, the Status Reason field contains the exact error message returned by the underlying service, such as an IAM permission issue or a missing S3 bucket. This is the authoritative source for diagnosing stack deployment problems.

Why this answer

The CloudFormation console's Events tab shows the chronological sequence of stack events, including the specific status reason for each resource failure. When a resource like AWS::RDS::DBInstance fails to create, the Events tab displays the underlying error message returned by the RDS API (e.g., invalid parameter, insufficient privileges, or subnet issues). This is the fastest and most direct way to diagnose the failure.

Exam trap

DOP-C02 often tests whether candidates conflate CloudTrail (API audit) with CloudFormation Events (resource-level deployment diagnostics), leading them to pick CloudTrail when the question asks for the specific failure reason.

How to eliminate wrong answers

Option A is wrong because AWS Config timeline tracks resource configuration changes and compliance over time, not CloudFormation stack deployment errors — it will not show why a CREATE_FAILED occurred. Option C is wrong because AWS Service Catalog is a governance and provisioning portal for approved products; it does not surface CloudFormation resource-level error messages. Option D is wrong because CloudTrail event history records API calls (who called what and when) but does not include the detailed resource creation failure reasons that CloudFormation surfaces in its own Events tab.

173
Multi-Selecthard

Which TWO metrics should be monitored in Amazon CloudWatch to detect a potential memory leak in an EC2 instance? (Choose two.)

Select 2 answers
A.DiskReadOps
B.MemoryUtilization (custom metric published via CloudWatch Agent).
C.NetworkIn
D.SwapUsage (custom metric published via CloudWatch Agent).
E.CPUUtilization
AnswersB, D

MemoryUtilization, published as a custom metric by the CloudWatch Agent, is calculated from /proc/meminfo or OS-level memory counters and represents the percentage of actual RAM in use. A memory leak causes this value to climb continuously over time as allocations are never freed, even when the workload is stable. Monitoring this metric reveals the leak as a monotonic upward trend, making it one of the most direct and essential signals for detection.

Why this answer

MemoryUtilization is a custom metric that must be published via the CloudWatch Agent because EC2 does not expose memory metrics by default. Monitoring this metric over time can reveal a steady upward trend in memory usage that does not drop after processes complete, which is a classic symptom of a memory leak.

Exam trap

The trap here is that candidates assume EC2 provides memory metrics by default (like CPUUtilization), but they must be explicitly enabled via the CloudWatch Agent, and they overlook SwapUsage as a complementary indicator of memory pressure from a leak.

174
MCQhard

A company's production environment consists of EC2 instances in an Auto Scaling group behind an Application Load Balancer (ALB). The instances run a web application that stores session data in an ElastiCache Redis cluster. The company has enabled detailed CloudWatch metrics and set up a dashboard. The operations team notices that the average CPU utilization across the Auto Scaling group spikes to 95% every 15 minutes, coinciding with a high number of Redis connections. What is the MOST likely cause?

A.The application is using Memcached instead of Redis, causing increased load.
B.The Auto Scaling group's scaling policy is based on memory utilization instead of CPU.
C.The ALB has session stickiness enabled, causing traffic to be routed to the same instances.
D.The ElastiCache cluster is not large enough to handle the number of requests.
AnswerC

With session stickiness enabled on the ALB target group, the load balancer consistently routes a given client's requests to the same EC2 instance for the stickiness duration. When a few power users or long-lived connections generate heavy traffic, those specific instances absorb a disproportionate share of load, causing CPU utilization to spike even if the aggregate fleet average remains moderate. This creates a hot-spot pattern where some targets are saturated while others idle, and the Auto Scaling group, if scaling on average CPU, may not react in time.

Why this answer

ALB session stickiness (sticky sessions) causes the load balancer to route requests from the same client to the same EC2 instance. When a large number of clients connect simultaneously (e.g., due to a periodic batch job), they may all be routed to the same few instances, causing CPU spikes on those instances. The high number of Redis connections corresponds to the sessions being stored in Redis.

This explains the 15-minute periodic spikes. Option A is incorrect because Memcached is not used; the question states ElastiCache Redis. Option B is incorrect because the scaling policy is based on CPU utilization, not memory.

Option D is incorrect because the cluster size doesn't cause periodic spikes; it would cause consistent high utilization.

175
MCQhard

A company has a multi-account strategy using AWS Organizations. The security team needs to respond to incidents across all accounts. They want to ensure that all CloudTrail trails are enabled and logging to a central S3 bucket in the management account. What is the MOST efficient way to monitor compliance?

A.Create a CloudTrail organization trail and use CloudTrail Insights to detect configuration changes.
B.Use AWS Config conformance packs with a managed rule to check CloudTrail is enabled.
C.Set up CloudWatch Events rules in each account to detect trail disabling.
D.Use AWS Trusted Advisor to check CloudTrail configuration in each account.
AnswerB

AWS Config conformance packs provide a way to deploy a collection of AWS Config rules and remediation actions across all accounts and Regions in an organization. A managed rule such as `cloudtrail-enabled` can be included in a conformance pack to verify that CloudTrail trails are configured and enabled, and the results are aggregated centrally in the AWS Config console for the entire organization. This approach gives a single, policy-as-code mechanism to enforce and audit CloudTrail enablement consistently across every account.

Why this answer

AWS Config conformance packs with managed rules can be deployed across multiple accounts using StackSets or directly via AWS Organizations to check that CloudTrail trails are enabled and logging to a central S3 bucket. This provides centralized, automated compliance monitoring without manual per-account setup. Option A is wrong because CloudTrail Insights detects unusual API activity, not configuration compliance.

Option C is wrong because setting up CloudWatch Events rules in each account is less efficient and harder to maintain than Config conformance packs. Option D is wrong because Trusted Advisor checks are per-account and cannot be centrally enforced across an organization.

Exam trap

Candidates might mistakenly believe that CloudTrail organization trails or Trusted Advisor can monitor compliance across all accounts, but neither provides the centralized rule enforcement and automated remediation that AWS Config conformance packs offer.

176
MCQeasy

A company experiences an unexpected spike in network traffic to a web application hosted on EC2 instances behind an Application Load Balancer. The DevOps team needs to investigate the source IP addresses generating the traffic. Which AWS service should they use to capture the traffic?

A.Amazon CloudWatch Logs
B.AWS Config
C.AWS CloudTrail
D.VPC Flow Logs
AnswerD

VPC Flow Logs capture IP traffic metadata for every network interface in a VPC, subnet, or at the interface level, recording source and destination IP addresses, ports, protocol, packet/byte counts, and whether the action was accepted or rejected by security groups and network ACLs. Enabling flow logs on the affected VPC or subnets would immediately provide the raw data needed to identify the top talkers, unusual port usage, or malicious sources behind the spike. This makes it the correct service for diagnosing an unexpected increase in network traffic.

Why this answer

VPC Flow Logs capture IP traffic information, including source and destination IPs, ports, and protocols, allowing investigation of source IPs. Option A (CloudWatch Logs) is wrong because it captures application logs, not network traffic. Option B (AWS Config) is wrong because it records resource configuration changes.

Option C (CloudTrail) is wrong because it logs API calls, not network traffic.

177
MCQeasy

A company runs a critical application on Amazon EC2 instances behind an Application Load Balancer. During a security incident, the security team needs to isolate a compromised instance for forensic analysis without affecting the application's availability. What is the MOST effective action to take?

A.Deregister the instance from the target group and stop the instance for forensic analysis.
B.Modify the security group of the instance to deny all inbound and outbound traffic.
C.Terminate the compromised instance immediately to prevent further damage.
D.Change the subnet route table to route traffic away from the compromised instance.
AnswerA

Deregistering the instance from the Target Group stops new traffic from the Application Load Balancer while leaving the instance running, which preserves volatile memory for forensic collection. Stopping the instance (not terminating it) retains the EBS volumes, allowing investigators to create snapshots or attach them to a forensic workstation for offline analysis. This approach keeps the remaining fleet operational and maintains availability while containing the compromise.

Why this answer

Deregistering the instance from the target group removes it from the Application Load Balancer's routing, ensuring no new traffic is sent to it while existing connections drain (connection draining). Stopping the instance preserves its memory and disk state for forensic analysis without impacting application availability, as the remaining healthy instances continue to serve traffic.

Exam trap

The trap here is that candidates confuse network-level isolation (security groups or route tables) with application-level isolation (target group deregistration), failing to recognize that the ALB continues to route traffic to a registered instance regardless of its security group or subnet routing.

How to eliminate wrong answers

Option B is wrong because modifying the security group to deny all traffic only blocks network-level access; the instance remains registered in the target group, and the ALB may still attempt to route traffic to it, potentially causing connection timeouts or errors. Option C is wrong because terminating the instance immediately destroys volatile data (e.g., memory contents, running processes) needed for forensic analysis and could cause a sudden loss of capacity if the instance was handling active requests. Option D is wrong because changing the subnet route table affects all instances in that subnet, not just the compromised one, and does not prevent the ALB from sending traffic to the instance via its private IP; route tables control layer-3 routing, not load balancer target group membership.

178
MCQmedium

A company uses AWS Organizations with multiple accounts. The security team wants to ensure that all IAM roles in member accounts have a maximum session duration of 1 hour. They need a way to detect any roles that violate this policy. What should they do?

A.Use IAM Access Analyzer to validate the roles against a policy template.
B.Use AWS Config with the managed rule iam-role-max-session-duration to evaluate roles.
C.Run AWS Trusted Advisor and check the IAM report for roles with long session durations.
D.Enable AWS CloudTrail and create a metric filter to detect role creation with session duration greater than 1 hour.
AnswerB

The AWS Config managed rule iam-role-max-session-duration evaluates every IAM role in the account, comparing each role's MaxSessionDuration setting against the rule's maxSessionDuration parameter. This rule is triggered proactively on configuration changes and periodically, so it detects both existing and newly modified roles, flagging any role whose allowed session duration exceeds the defined threshold as noncompliant. It integrates with AWS Organizations and can be remediated automatically or through Config conformance packs.

Why this answer

AWS Config provides a managed rule called `iam-role-max-session-duration` that specifically evaluates IAM roles to ensure their `MaxSessionDuration` setting does not exceed a specified threshold (default 1 hour). This rule can be deployed across all member accounts in AWS Organizations using a conformance pack or AWS Config aggregator, allowing the security team to continuously detect and report any roles that violate the policy without manual intervention.

Exam trap

The trap here is that candidates often confuse AWS Config's ability to evaluate resource configurations (like IAM role session duration) with CloudTrail's event logging or IAM Access Analyzer's policy analysis, leading them to choose options that detect creation events rather than continuously assess the current state of all roles.

How to eliminate wrong answers

Option A is wrong because IAM Access Analyzer is designed to analyze resource-based policies (like S3 bucket policies or KMS key policies) for unintended public or cross-account access, not to validate IAM role session duration settings against a policy template. Option C is wrong because AWS Trusted Advisor checks for IAM use (e.g., unused IAM users, MFA on root) but does not include a specific check for IAM role maximum session duration. Option D is wrong because while CloudTrail can log `CreateRole` and `UpdateAssumeRolePolicy` events, a metric filter cannot directly evaluate the `MaxSessionDuration` parameter from the event; it would require complex custom parsing and still not provide ongoing compliance evaluation like AWS Config.

179
MCQeasy

A DevOps engineer receives a CloudWatch alarm that an Auto Scaling group has been in an 'Insufficient data' state for 20 minutes. What does this indicate?

A.All instances in the Auto Scaling group are unhealthy
B.The Auto Scaling group needs to scale up
C.The alarm has not received enough metric data to evaluate
D.The CloudWatch agent is not installed on the instances
AnswerC

CloudWatch alarms use the state INSUFFICIENT_DATA when the number of metric data points available in the evaluation period is less than the number required to determine the alarm state. This commonly occurs for a newly created alarm before the first metric points arrive, when an instance is stopped and stops publishing metrics, or when the metric name is missing. The alarm cannot evaluate to ALARM or OK until the metric stream provides enough valid points within the configured period.

Why this answer

The 'Insufficient data' state in CloudWatch alarms indicates that the alarm has not received enough metric data points to determine whether the threshold has been breached. This can occur when the metric is not being published, the data collection period is too short, or the metric namespace is misconfigured. It does not directly indicate instance health, scaling needs, or agent installation status.

Exam trap

The trap here is that candidates confuse 'Insufficient data' with a problem state (like unhealthy instances or scaling failures), when it actually means the alarm simply lacks enough data to make a determination.

How to eliminate wrong answers

Option A is wrong because 'Insufficient data' does not imply unhealthy instances; it means the alarm lacks metric data to evaluate, whereas unhealthy instances would trigger 'ALARM' state if health check metrics are configured. Option B is wrong because the alarm state does not indicate a scaling need; scaling decisions are based on threshold breaches, not insufficient data. Option D is wrong because the CloudWatch agent is not required for all metrics; many metrics (e.g., EC2 basic monitoring) are published automatically without an agent, and 'Insufficient data' can occur even with the agent installed if data is not flowing.

180
MCQeasy

A DevOps engineer receives an alert that an Amazon EC2 instance’s CPU utilization has been above 90% for the past hour. The instance is part of an Auto Scaling group with a step scaling policy based on average CPU. The engineer checks the CloudWatch alarm and sees that it is in the ALARM state. What should the engineer do to verify that the Auto Scaling group is scaling out properly?

A.Ensure the scaling policy is configured to scale in
B.Check the CloudWatch Logs for the instance
C.Verify that the CloudWatch alarm is in INSUFFICIENT_DATA state
D.Review the Auto Scaling group’s activity history in the EC2 console
AnswerD

The Auto Scaling group’s activity history is the authoritative record of every scale-out and scale-in event, including the time, instance ID, reason code, and status (e.g., 'InProgress', 'Successful', or 'Failed'). By reviewing it, you can determine whether the CloudWatch alarm triggered the scaling policy, whether the capacity was blocked by the maximum instance limit, a service quota, or a launch template error. This directly shows why the instance did not receive the expected scale-out action, making it the correct first troubleshooting step.

Why this answer

Reviewing the Auto Scaling group's activity history in the EC2 console shows whether scaling actions were triggered and if new instances were launched. The CloudWatch alarm is in ALARM state (indicating high CPU), and the step scaling policy should scale out, so checking the activity history confirms the scaling action occurred. Option A is incorrect because the scaling policy is for scaling out on high CPU, not scaling in.

Option B is incorrect because CloudWatch Logs show instance-level logs, not Auto Scaling actions. Option C is incorrect because the alarm is in ALARM state, not INSUFFICIENT_DATA.

181
MCQhard

A company uses AWS Organizations with multiple accounts. The security team needs a centralized solution to detect and respond to EC2 instances that are publicly accessible with SSH open to 0.0.0.0/0. Which combination of services provides the most automated detection and remediation?

A.AWS CloudTrail and Amazon EventBridge
B.Amazon GuardDuty and AWS Lambda
C.AWS Config and Amazon Simple Notification Service (SNS)
D.AWS Config and AWS Systems Manager Automation
AnswerD

This pair provides end-to-end compliance: AWS Config continuously evaluates security group resources against rules like restricted-common-ports, and when an SG is non-compliant, Config's remediation action triggers an AWS Systems Manager Automation runbook (e.g., AWS-DisablePublicSecurityGroupIngress or a custom runbook). The Automation step performs the actual modification, such as revoking offending ingress rules, thus closing the loop. Config can remediate automatically via SSM Automation without manual intervention, making this the correct solution.

Why this answer

AWS Config and AWS Systems Manager Automation together provide automated detection and remediation. AWS Config rules can detect EC2 instances with security groups that allow SSH from 0.0.0.0/0, and the rule can trigger an SSM Automation document to remediate the issue, such as by modifying the security group or stopping the instance. This combination is centralized and works across multiple accounts in AWS Organizations when Config and SSM are configured with a delegated administrator.

Exam trap

DOP-C02 often tests the difference between detection-only services (Config with SNS, GuardDuty) and detection-plus-automated-remediation combinations (Config with SSM Automation), leading candidates to choose a notification-only solution.

How to eliminate wrong answers

Option A is wrong because CloudTrail and EventBridge can detect API calls but do not evaluate resource configuration state or provide built-in remediation for publicly accessible instances. Option B is wrong because GuardDuty detects threats and anomalous behavior, not misconfigured security groups, and Lambda alone would require custom code for both detection and remediation. Option C is wrong because AWS Config with SNS only notifies; it does not automatically remediate the issue.

182
MCQhard

A company runs a containerized microservices application on Amazon ECS with Fargate launch type. The application uses an Application Load Balancer to route traffic to the ECS service. Recently, the DevOps team noticed that the ECS service is failing to deploy new tasks during a rolling update. The CloudWatch Logs for the ECS service show that new tasks are failing to start because they cannot pull the container image from Amazon ECR. The error message indicates 'AccessDenied' when attempting to pull the image. The task execution role has the necessary permissions, and the image URI is correct. The VPC has a VPC endpoint for ECR configured. The security group for the tasks allows outbound traffic to the VPC endpoint. What is the MOST likely cause of the access denied error?

A.The task execution role does not have the 'ecr:GetAuthorizationToken' permission.
B.The security group for the ALB does not allow inbound traffic to the ECS tasks.
C.The VPC endpoint for ECR does not have 'Private DNS names enabled' selected.
D.The task role does not have the 'ecr:BatchGetImage' permission.
AnswerC

This option is correct because when the 'Private DNS names enabled' option is not selected for an Amazon ECR VPC endpoint, the default DNS endpoint (ecr.<region>.amazonaws.com) continues to resolve to public IP addresses. Since the ECS task runs in a private subnet with no internet route, the ECS agent sends the request to a public IP that is unreachable, and the VPC endpoint responds with AccessDenied because the request is not being handled by the endpoint. Enabling private DNS names creates a Route 53 private hosted zone that associates the endpoint's DNS with its private IP, ensuring traffic destined for ECR is routed through the VPC endpoint and bypasses the public network.

Why this answer

For ECS tasks using Fargate to pull images from ECR via a VPC endpoint, the private DNS names must be enabled on the endpoint. If not enabled, the task's DNS resolution of the ECR repository URL returns a public IP, which may be blocked by security groups or route tables, causing an 'AccessDenied' error despite correct IAM permissions. Option A is incorrect because 'ecr:GetAuthorizationToken' is needed for authentication, but the error occurs after authentication (the task execution role has permissions).

Option B is irrelevant as the ALB security group does not affect image pulling. Option D is incorrect because the task role is for application-level permissions, not for pulling images; that is handled by the task execution role.

183
MCQhard

A DevOps engineer is troubleshooting an issue where an Amazon RDS instance's CPU utilization is consistently high. The engineer has enabled Performance Insights and sees that the top SQL query is a SELECT statement that scans many rows. What is the best course of action to reduce CPU utilization?

A.Create a read replica to offload read traffic.
B.Increase the allocated storage to improve I/O.
C.Increase the DB instance size to handle the load.
D.Add appropriate indexes to optimize the query.
AnswerD

Adding an appropriate index transforms a full table scan into an index seek, drastically reducing the number of rows the database engine must read and process, which directly lowers the CPU cycles spent on evaluating rows, join operations, and sort operations. An index that matches the query predicate (e.g., on the columns used in WHERE, JOIN, or ORDER BY) allows the optimizer to access only relevant pages, shrinking buffer pool I/O and improving overall query latency. This is the root-cause fix because it eliminates the unnecessary CPU work rather than merely hiding the symptom.

Why this answer

Adding appropriate indexes to optimize the query is the best course of action because the high CPU utilization is caused by a SELECT statement that scans many rows. Indexes allow the database to locate rows efficiently without scanning the entire table, reducing CPU load and improving query performance. This directly addresses the root cause identified by Performance Insights.

Exam trap

The trap is that candidates may choose to scale up the instance (option C) as a quick fix, but the question asks for the best course of action to reduce CPU utilization, which is to address the root cause with indexing.

How to eliminate wrong answers

Option A is wrong because creating a read replica offloads read traffic but does not optimize the inefficient query; the same full scan would still occur on the replica, and the primary would still be impacted if the query runs there. Option B is wrong because increasing allocated storage improves I/O capacity, not CPU utilization caused by row scanning. Option C is wrong because increasing the DB instance size provides more CPU resources but does not fix the underlying inefficient query, which will continue to consume excessive CPU and may still be a problem as data grows.

← PreviousPage 3 of 3 · 183 questions total

Ready to test yourself?

Try a timed practice session using only Incident and Event Response questions.