Be able to map a stated RTO and RPO to a specific DR strategy and AWS service, and to configure metric alarms and centralized log queries. The single most important thing: read the RTO/RPO numbers first, then pick the cheapest strategy that still meets both.
Start practicing
Operations and Support — choose a session length
Free · No account required
Domain overview
Operations and Support covers running cloud workloads after deployment: monitoring, alerting, logging, backup, and disaster recovery. Questions present a business requirement (RTO, RPO, region, cost) and ask you to select the AWS service, replication strategy, or alert configuration that satisfies it. Expect scenario-based items tying metrics and logs to concrete actions.
Exam objectives
Selecting RTO/RPO-appropriate DR strategies: backup/restore, pilot light, warm standby, multi-site
Using CloudWatch alarms on metrics like CPUUtilization with thresholds and evaluation periods
Centralizing logs with CloudWatch Logs, subscriptions, and CloudWatch Logs Insights queries
Cross-region replication options: S3 CRR, RDS read replicas, and snapshot copying
Confusing RTO with RPO: RTO is acceptable downtime, RPO is acceptable data loss window; swapping them changes the correct DR strategy.
Choosing multi-site active-active for every scenario, ignoring cost when warm standby or pilot light already meets the stated RTO/RPO.
Assuming CloudWatch alarms act alone; they need SNS topics or Auto Scaling actions attached to actually notify or remediate.
Click any question to see the full explanation and answer options, or start a focused practice session above.
A cloud engineer is setting up automated patching for Linux instances in AWS. They need to define a maintenance window during which patches are applied. Which service should they use?
2A cloud administrator is troubleshooting a network connectivity issue between two subnets. They suspect a security group or NACL is blocking traffic. Which tool should they use to analyze the traffic flow?
3A cloud engineer needs to ensure that an auto-scaling group does not launch new instances immediately after a scale-in event to allow metrics to stabilize. Which feature should they configure?
4A company wants to reduce costs by identifying underutilized EC2 instances. Which tool should they use to get rightsizing recommendations?
5A cloud operations team is implementing structured logging for better querying. They have decided to use JSON format. What is a key benefit of structured logging over unstructured logging?
6A cloud architect is designing a disaster recovery plan that includes testing. Which TWO activities are commonly performed as part of DR testing?
7A cloud engineer is troubleshooting a performance issue in a microservices application. Which THREE tools can help with distributed tracing and latency diagnosis?
8A cloud administrator needs to monitor CPU utilization for a fleet of EC2 instances and receive notifications when utilization exceeds 80%. Which AWS service should be used to create a metric alarm that triggers an SNS notification?
9A company wants to centralize logs from multiple AWS services and analyze them using SQL-like queries. Which service should they use?
10A cloud administrator needs to perform a disaster recovery test for a critical application running in a different AWS region. The RTO is 1 hour, and the RPO is 15 minutes. Which replication strategy should be used to meet the RPO?
11A company uses AWS and wants to optimize costs by receiving recommendations to downsize over-provisioned EC2 instances. Which tool provides rightsizing recommendations?
12A cloud engineer is troubleshooting a performance issue in a multi-tier application on AWS. The web tier shows high latency, but the application logs indicate no errors. The engineer wants to trace a request end-to-end across services. Which AWS service should be used?
13A cloud administrator needs to apply security patches to a fleet of EC2 instances running Windows Server. The patches must be applied during a maintenance window to minimize downtime. Which AWS service can automate patching?
14Which GCP service provides centralized log management and analysis with the ability to create log-based metrics and alerts?
15A company wants to implement a tagging strategy for their cloud resources to track costs by department and project. Tags must be applied to resources such as virtual machines and storage buckets. Which of the following is a best practice for cost attribution using tags?
16A cloud administrator is configuring an alert for an Azure virtual machine. The alert should trigger when the average CPU percentage exceeds 90% for more than 10 minutes. Which Azure service should be used to create this metric alert?
17A cloud engineer is configuring an auto-scaling group with a lifecycle hook to run a custom script when instances are launched. The script installs software and registers the instance with a load balancer. The engineer wants to ensure the instance does not receive traffic until the script completes successfully. What should the engineer do?
18A cloud administrator is configuring a notification channel for critical alerts. Which TWO of the following are commonly used notification channels in cloud monitoring systems? (Select TWO.)
19A cloud team is planning a disaster recovery drill for their application running in a public cloud. They want to validate that the recovery process meets the defined RTO and RPO. Which THREE activities should be included in the DR drill? (Select THREE.)
20A cloud administrator needs to receive real-time notifications when CPU utilization exceeds 90% on a production server. Which AWS service should be used to trigger an alert based on a metric threshold?
21A company uses a multi-cloud environment with AWS and Azure. They want to centralize log collection and enable advanced querying for troubleshooting. Which combination of services should they use?
22A cloud operations team is implementing a tagging strategy for cost attribution. They need to track costs by environment (dev, test, prod), project, and team. Which approach should they use?
23A cloud engineer notices that an auto-scaling group is adding and removing instances too frequently, causing instability. Which configuration parameter should be adjusted to reduce this behavior?
24A company wants to implement automated patching for their Windows and Linux servers in AWS. They need to schedule patching during a maintenance window and have a rollback plan. Which service should they use?
25A cloud administrator wants to troubleshoot network connectivity issues between two VPCs. Which AWS feature provides detailed logs of IP traffic for analysis?
26A cloud architect is designing a disaster recovery plan. The application requires a Recovery Time Objective (RTO) of 15 minutes and a Recovery Point Objective (RPO) of 1 hour. Which strategy best meets these requirements?
27A company uses GCP and wants to implement alerting based on anomaly detection for their Compute Engine instances. Which GCP service should they use?
28A company wants to optimize cloud costs by identifying underutilized EC2 instances. Which AWS service provides rightsizing recommendations?
29A cloud operations team is setting up log-based alerting for security events. They want to use structured logging to facilitate querying. Which TWO practices support effective log-based alerting? (Choose TWO.)
30A cloud administrator needs to detect unusual spikes in CPU usage across a fleet of EC2 instances. Which AWS service should be used to create an alarm that triggers when CPU utilization exceeds an expected baseline?
31A cloud engineer wants to receive real-time notifications when a CloudWatch Alarm enters the ALARM state. Which notification channel can be configured directly within the CloudWatch Alarm action?
32A company uses AWS and wants to analyze cost trends and identify the top services contributing to monthly spending. Which AWS tool provides a pre-built dashboard for this purpose?
33An organization uses GCP and wants to implement a tagging strategy to track costs by project and environment. Which GCP feature should be used to assign metadata to resources for cost attribution?
34A cloud administrator is designing an auto-scaling policy for a web application that experiences predictable traffic spikes during business hours. The administrator wants to ensure that the application scales out before the start of business hours to avoid performance degradation. Which scaling policy type should be used?
35A cloud engineer needs to apply security patches to a group of Linux VMs running in Azure. The engineer wants to automate the patching process and ensure that patches are applied during a predefined maintenance window. Which Azure service should be used?
36A cloud administrator is troubleshooting network connectivity issues between two VPCs in AWS. The administrator wants to examine traffic flow logs to identify dropped packets. Which AWS feature provides detailed network traffic logs for VPCs?
37A cloud engineer needs to collect and query log data from multiple cloud services in a centralized location. Which cloud service should be used for centralized log management?
38A company wants to reduce costs by identifying underutilized EC2 instances and receiving recommendations to downsize them. Which AWS service provides rightsizing recommendations based on historical utilization metrics?
39A cloud administrator is setting up auto-scaling for a web application that uses an SQS queue for incoming requests. The administrator wants to scale the number of EC2 instances based on the queue depth. Which two metrics are appropriate for this auto-scaling policy? (Choose TWO.)
40A company uses AWS and wants to implement a structured logging format to simplify querying and analysis of application logs. Which three best practices should be followed when implementing structured logging? (Choose THREE.)
41A cloud administrator needs to monitor CPU utilization of a group of virtual machines and automatically add more instances when utilization exceeds 80% for 5 minutes. Which cloud service should the administrator use to define this scaling policy?
42A company wants to implement a disaster recovery strategy with an RTO of 15 minutes and an RPO of 1 hour for a critical application running on AWS. Which approach would best meet these requirements?
43A cloud administrator is troubleshooting a performance issue where an application occasionally experiences high latency. The application runs on AWS and uses EC2, ELB, and RDS. Which combination of tools would best help trace the request flow and identify the bottleneck?
44A company wants to be notified when their monthly AWS spending exceeds $10,000. Which AWS service should they use to set up this alert?
45A cloud administrator needs to apply security patches to a fleet of 50 Linux servers running on AWS without interrupting business hours. Which approach should the administrator use to schedule patching during a maintenance window?
46A company is using Azure VMs and wants to centralize logs from multiple applications for security analysis. The logs must be retained for 2 years. Which Azure service should they use?
47A cloud administrator notices that an Auto Scaling group is launching and terminating instances too frequently, causing instability. What should the administrator adjust to reduce this flapping behavior?
48A company uses AWS and wants to receive alerts when CPU utilization of an EC2 instance exceeds 90% for 10 minutes. Which AWS service should be used to create this alarm?
49A cloud engineer needs to implement a solution to automatically scale an application based on the number of messages in an SQS queue. The goal is to keep the queue length short. Which Auto Scaling policy type should the engineer use?
50A company uses GCP and wants to ensure that log entries from Compute Engine instances are automatically exported to BigQuery for analysis. The logs must include structured JSON data. Which GCP service should be configured to route logs?
51A cloud administrator is configuring cost management for a multi-account cloud environment. The company wants to allocate costs by department and project. Which TWO steps should the administrator take to achieve this? (Choose two.)
52A cloud engineer is planning a disaster recovery drill for a critical application that spans multiple availability zones. The drill must validate RTO and RPO without affecting production. Which THREE actions should the engineer include? (Choose three.)
53A company uses AWS CloudFormation to manage infrastructure. The operations team needs to be alerted when a stack update fails. Which TWO methods can be used to send notifications? (Choose two.)
54A cloud architect needs to monitor CPU utilization across a fleet of EC2 instances and receive an alert when the average CPU exceeds 80% for 10 minutes. Which AWS service should be used to collect the metric and trigger the alert?
55A company wants to track cloud spending by department and project. Which strategy should be implemented to enable cost attribution?
56A cloud administrator needs to centralize logs from multiple AWS services, including VPC flow logs and application logs, to enable searching and querying. Which solution should be used?
57An organization wants to automate patching of their EC2 instances running Windows Server. They need to schedule patching during a maintenance window and ensure minimal downtime. Which AWS service should they use?
58A cloud engineer needs to troubleshoot network connectivity issues between two subnets. Which feature can help capture and analyze network traffic metadata?
59A company uses Azure and wants to set up an alert that triggers when the average CPU of a virtual machine exceeds 90% for the past 15 minutes. The alert should send an email to the operations team. Which Azure resources are needed?
60A cloud operations team wants to analyze application performance and identify slow database queries. They need a distributed tracing solution. Which service should they use?
61A cloud administrator is configuring auto-scaling for a batch processing application that uses an SQS queue. The number of jobs varies unpredictably. Which metric is most appropriate for scaling the worker instances?
62A company wants to reduce cloud costs by identifying underutilized EC2 instances. Which AWS service provides rightsizing recommendations?
63A cloud operations team is designing a disaster recovery plan that includes regular testing. Which TWO activities should be part of the DR testing process? (Select TWO.)
64A company uses AWS and wants to implement structured logging for their applications to improve queryability. Which THREE practices should they follow? (Select THREE.)
65An organization is using Azure and wants to implement a patch management strategy with minimal disruption. Which TWO actions should they take? (Select TWO.)
66A company uses a multi-cloud strategy with workloads in AWS and Azure. The cloud team wants a centralized log management solution to correlate security events across both platforms. Which approach is most suitable?
67A cloud architect is designing a disaster recovery plan for a critical application with an RTO of 15 minutes and an RPO of 1 minute. The application runs on AWS EC2 instances with data stored on EBS volumes. Which replication strategy best meets these requirements?
68An organization wants to reduce cloud costs by identifying underutilized EC2 instances. Which AWS service provides rightsizing recommendations?
69A cloud administrator is configuring an auto-scaling group for a web application. The application experiences predictable traffic spikes every weekday at 9 AM. Which scaling policy is most appropriate?
70A cloud engineer is troubleshooting a network connectivity issue between two VPCs in AWS. To analyze traffic patterns and identify dropped packets, which feature should be enabled?
71A cloud administrator notices that an application's latency has increased. The application is distributed across multiple microservices. Which tool can help trace requests across services to identify the bottleneck?
72A company is planning to migrate to AWS and wants to achieve the lowest possible compute costs for a steady-state workload that will run 24/7. Which purchasing option should be recommended?
73A cloud administrator needs to set up a centralized logging solution to collect logs from multiple projects. Which cloud service should be used?
74A cloud operations team is preparing for a disaster recovery drill for a multi-tier application. Which TWO activities are essential for verifying the effectiveness of the DR plan? (Select TWO.)
75A cloud administrator is tasked with monitoring CPU utilization across a fleet of virtual machines. Which cloud service should be used to collect and visualize this metric?
76A company is migrating on-premises workloads to a public cloud. The disaster recovery plan requires an RTO of 15 minutes and an RPO of 5 minutes. Which replication strategy should be used for a critical database?
77A cloud administrator needs to centralize logs from multiple cloud provider accounts and on-premises servers for security analysis. Which approach should be used?
78A cloud engineer is implementing a tagging strategy for cost allocation. Which tags should be applied to resources to track costs by business unit and environment?
79During a scheduled DR drill, the cloud team fails over a critical application to the secondary region. After the drill, the application is failed back. The application's RTO was 2 hours, but the actual failover took 2.5 hours. Which action should be taken to improve future failover times?
80A cloud administrator needs to apply security patches to a group of Windows servers during a maintenance window to minimize disruption. Which type of service should be used?
81A cloud architect is designing an auto-scaling policy for a web application. The application's traffic spikes predictably every weekday at 9 AM and decreases after 5 PM. Which scaling policy is most cost-effective?
82A cloud administrator needs to ensure that log data is retained for one year to meet compliance requirements. Which action should be taken for the log group in CloudWatch Logs?
83A company uses a hybrid cloud environment with workloads in AWS and on-premises. They want to use a single monitoring dashboard to view metrics from both environments. Which solution should they implement?
84A company is performing a disaster recovery test for a critical application. The test reveals that the application's RTO of 1 hour is not being met due to slow database restoration. Which THREE actions could help improve the restoration time? (Select THREE.)
85A cloud engineer wants to be notified when the average CPU utilization of an auto-scaling group exceeds 80% for 5 minutes. Which alerting mechanism should be used?
86A company wants to implement a disaster recovery strategy with an RTO of 15 minutes and an RPO of 1 minute for a critical database. Which approach should be used?
87A cloud administrator wants to analyze network traffic to troubleshoot connectivity issues between VMs. Which feature should be enabled?
88A cloud team wants to automatically scale an application based on the number of pending messages in a message queue. Which scaling policy type should be used?
89A company wants to ensure that logs from their application are easily searchable and structured for analysis. Which logging format should be recommended?
90A cloud administrator notices that an auto-scaling group is frequently adding and removing instances due to brief spikes in CPU usage. What should be adjusted to stabilize the scaling activity?
91A cloud engineer wants to view a dashboard showing cost breakdown by department. Which tool provides pre-built billing dashboards?
92A company is experiencing intermittent performance issues in a microservices application. Which TWO tools can help diagnose latency problems through distributed tracing? (Choose TWO)
93A cloud administrator wants to choose an auto-scaling policy that can respond to changing demand patterns. Which TWO policy types support dynamic adjustments based on real-time metrics? (Choose TWO)
94A cloud administrator is investigating a sudden increase in cost for a production environment. The administrator wants to identify the sources of the cost increase and implement a tagging strategy for cost allocation. Which TWO actions should the administrator take? (Choose two.)
95A company is designing a disaster recovery plan for a critical database that requires a recovery point objective (RPO) of 1 minute and a recovery time objective (RTO) of 15 minutes. The database runs on a cloud virtual machine. Which backup strategy should the administrator implement to meet these requirements?
96A cloud application is experiencing intermittent high latency. The operations team has enabled distributed tracing using AWS X-Ray but is unable to pinpoint the source. Which additional step should the team take to identify the root cause of the latency?
97A cloud operations team runs a fleet of Amazon EC2 instances behind an Application Load Balancer. During a load test, the team notices that healthy targets are being marked unhealthy and removed from rotation whenever a deployment briefly pushes CPU utilization above 90 percent. The team wants the load balancer to remove an instance only when the application stops responding to HTTP requests, not when it is merely busy. Which action should the team take?
98A cloud operations engineer is responsible for a fleet of Amazon EC2 instances that run a stateless web application behind an Application Load Balancer. The engineer needs to perform a rolling replacement of the instances with a new AMI while ensuring that the application remains available and that the deployment automatically rolls back if a specified Amazon CloudWatch alarm enters the ALARM state. Which AWS deployment service should the engineer use?
99A cloud operations team manages a multi-tier web application on Google Cloud. The application logs are being written to Cloud Logging, and the team needs to be alerted whenever the number of HTTP 500 errors exceeds 50 in a 5-minute window. Which action should the team take to meet this requirement?
100A cloud administrator is managing a Microsoft Azure environment. The administrator needs to enforce a policy that prevents the creation of any Azure Storage account without HTTPS-only traffic enabled and without a minimum TLS version of 1.2. The policy must apply to all current and future subscriptions in the tenant and must be evaluated when resources are created or updated. Which Azure feature should the administrator use?
101A cloud engineer is responsible for a Kubernetes cluster on AWS EKS. The cluster runs a stateful application that requires persistent storage. The engineer must ensure that storage volumes are automatically provisioned when persistent volume claims are created, and that the storage remains available if the pod is rescheduled to a different node. Which solution should the engineer implement?
102A cloud operations team runs a three-tier web application on AWS. During a recent incident, the on-call engineer received hundreds of Amazon CloudWatch alarms within minutes and could not identify the root cause. The team wants to reduce alarm fatigue while still capturing meaningful signals. Which action should the team take FIRST?
103A cloud engineer is investigating intermittent latency in a three-tier application hosted in Google Cloud. The engineer suspects that a specific Compute Engine instance is experiencing packet loss to its database backend. The engineer needs to capture and analyze the traffic at the packet level on the instance without installing third-party agents on the instance and without disrupting production traffic. Which Google Cloud feature should the engineer use?
104A cloud administrator manages an Amazon EC2 Auto Scaling group behind an Application Load Balancer. Users report intermittent 502 errors during scale-in events. Logs show that instances are terminated while still serving in-flight requests. Which configuration change should the administrator make to resolve this?
105A cloud operations team runs a fleet of Amazon EC2 instances behind an Application Load Balancer. During a recent incident, the team discovered that a single unhealthy instance continued to receive traffic for several minutes before being removed. The team wants to reduce the time it takes for the load balancer to detect and stop routing traffic to unhealthy targets. Which action should the administrator take to meet this requirement?
106A cloud administrator is responsible for a set of Linux virtual machines in AWS. The administrator needs to run a script on all of the instances at a scheduled time each night to rotate application logs. The script must run without the administrator logging in to each instance, and the administrator wants to avoid managing SSH keys for this task. Which AWS service should the administrator use?
107A cloud operations team runs a containerized workload on Amazon ECS with tasks spread across an Auto Scaling group of EC2 instances. During a peak-traffic event, the team observes that a single task repeatedly restarts with an out-of-memory error while the host instance still shows 40% free memory. The team wants the scheduler to stop placing new tasks on that host when its committed memory is exhausted. Which action should the team take?
108A cloud engineer is responsible for a fleet of Amazon EC2 instances running a stateless web application. The engineer needs to ensure that if an instance fails a status check, it is automatically replaced without manual intervention. Which AWS feature should be used to meet this requirement?
109A cloud administrator needs to ensure that an Amazon S3 bucket containing regulated data logs every object-level access attempt, including reads and writes, for audit purposes. Which action should the administrator take?
110A cloud operations team manages a fleet of Amazon EC2 instances behind an Application Load Balancer. During a recent incident, several instances stopped passing their ELB health checks but the Auto Scaling group did not replace them. The team wants the Auto Scaling group to automatically terminate and replace instances that fail ELB health checks, not just EC2 status checks. Which action should the team take?
111A cloud operations team is designing a backup strategy for a set of Amazon RDS for MySQL databases that support a production application. The team needs to be able to restore the database to any point in time within the last 35 days and must also retain a copy of the database for seven years for regulatory compliance. The team wants to minimize operational overhead. Which TWO actions should the team take? (Choose two.)
112A cloud operations team manages a multi-account AWS environment with AWS Organizations. They need a centralized, near-real-time view of security findings across all accounts and want the ability to automatically suppress findings that match approved exceptions. Which service should they use to aggregate and manage these findings?
113A cloud operations team needs to ensure that all Amazon S3 buckets in their AWS account have server access logging enabled. They want to automatically detect and remediate any bucket that does not have logging enabled. Which combination of AWS services should they use?
114A cloud administrator needs to give the security team read-only visibility into all API activity across an AWS account, including who made each call, when, and from which IP address. The records must be retained for 365 days for compliance. Which service should the administrator use?
115A cloud administrator is deploying a containerized workload to Google Kubernetes Engine. The workload must automatically scale based on the number of incoming HTTP requests per second rather than CPU utilization. Which GKE feature should the administrator configure to meet this requirement?
116A cloud operations team is deploying a three-tier application across multiple availability zones. To meet a strict recovery time objective, they want the application to keep serving traffic if an entire availability zone becomes unavailable. Which TWO design actions should the team take? (Choose two.)
117A cloud operations team deploys a containerized workload to a Kubernetes cluster managed by Amazon EKS. The application pods intermittently fail during peak traffic hours, and the team suspects that the pods are being terminated because they exceed their configured resource limits. Which action should the team take FIRST to confirm this suspicion?
118A cloud engineer is troubleshooting a Microsoft Azure virtual machine that becomes unresponsive under sustained load. The engineer suspects a storage performance bottleneck but needs to confirm whether the issue is IOPS throttling or throughput throttling on the managed disk. Which Azure Monitor metrics should the engineer examine to distinguish between these two causes?
119A cloud administrator is asked to give the security team read-only visibility into all objects stored in an Amazon S3 bucket used for application logs, without granting the ability to delete or overwrite any object. The security team authenticates as an IAM role. Which action should the administrator take?
120A cloud engineer is investigating a sudden increase in egress charges. The engineer suspects that a misconfigured Amazon S3 bucket is being read frequently from the internet. Which tool should the engineer use to identify the source IP addresses and request patterns for that bucket?
121A cloud administrator manages a fleet of Amazon EC2 instances that must receive operating system patches on a defined schedule. The administrator wants to automate patching, control the maintenance window, and receive compliance reports showing which instances are missing patches. Which AWS service should the administrator use?
122A cloud operations team runs a containerized API on Amazon ECS with the Fargate launch type. During peak hours, CPU utilization on the tasks regularly reaches 95 percent and response latency doubles. The team wants the service to add tasks automatically before users notice degradation, and to remove them when demand drops. Which action should the team take?
123A cloud operations team manages a containerized microservices application running on an Amazon EKS cluster. During peak hours, the team observes that pods are frequently being terminated and restarted, and node CPU utilization is consistently above 90 percent. The team wants to automatically scale the number of pods based on CPU utilization while ensuring the cluster has enough nodes to schedule the pods. Which combination of actions should the team take to meet these requirements?
124A cloud operations team runs a three-tier application on Google Cloud and wants to reduce the mean time to recovery for incidents. They want automated actions to run when a Cloud Monitoring alert fires, and they want to capture the exact configuration state at the moment of the incident for later analysis. Which TWO approaches should the team implement? (Choose two.)
125A cloud operations team must ensure that a critical workload continues to run even if an entire AWS Region becomes unavailable. The workload's data is stored in Amazon S3, and the team wants the data available in a second Region with minimal operational effort and automatic replication. Which action should the team take?
126A cloud engineer is investigating why an application hosted on Amazon EC2 cannot connect to an Amazon RDS for MySQL database in the same VPC. The database security group allows traffic on port 3306 from the application's security group. The engineer confirms the application is using the correct endpoint and credentials. Which action should the engineer take NEXT to identify the cause?
127A cloud administrator manages workloads in Microsoft Azure. The security team requires that all virtual machines apply operating system updates automatically during a defined window without the administrator logging in to each machine. Which Azure feature should the administrator use to meet this requirement?
128A cloud operations team runs a Kubernetes cluster on Google Kubernetes Engine (GKE). They need to ensure that a critical payment microservice is automatically restarted if its container process fails, and that a new Pod is created if the node hosting it becomes unhealthy. Which Kubernetes object should they configure to meet these requirements?
129A cloud operations team runs a three-tier application on Amazon EC2 instances behind an Application Load Balancer. During a peak-traffic event, users report intermittent 503 errors, and the operations team wants to automatically add capacity when the average CPU utilization of the Auto Scaling group exceeds 70 percent for five consecutive minutes, then remove capacity when it drops below 30 percent. Which TWO configuration elements must the team define to accomplish this? (Choose two.)
130A cloud administrator manages a fleet of Linux virtual machines on Google Cloud. A compliance rule requires that an interactive SSH session to any of these instances be brokered through an identity-aware proxy so that sessions are authenticated and auditable, and that no external IP addresses be assigned to the instances. Which solution should the administrator implement?
131A cloud operations team manages a three-tier application on Google Cloud. After a deployment, users report intermittent 503 errors from the HTTP(S) load balancer, and backend health checks are flapping between healthy and unhealthy. The team suspects the backend instances are being overwhelmed during health check bursts. Which TWO actions should the team take to stabilize the health checks and reduce false failures? (Choose two.)
132A cloud operations team needs to reduce the mean time to recovery for a microservices application running on Amazon EKS. They want to detect service degradation earlier and automatically replace unhealthy pods without manual intervention. Which TWO actions should the team take? (Choose two.)
133A cloud engineer is responsible for a set of Amazon EC2 instances that run a stateless web application. The engineer must ensure that the application can automatically recover from instance-level failures and that new instances are launched in multiple Availability Zones to maintain high availability. Which combination of AWS services should the engineer use?
134A cloud engineer is deploying a containerized workload to Google Kubernetes Engine. The security team requires that the container run as a non-root user, that the root filesystem be mounted read-only, and that privilege escalation be disallowed. The engineer wants to enforce these controls at the pod level so that any violating pod is rejected during admission. Which action should the engineer take?
135A cloud operations team runs a three-tier web application on Amazon EC2 instances behind an Application Load Balancer. Users report intermittent 502 errors, and the operations team wants to identify whether the issue originates from unhealthy targets before the load balancer removes them. Which AWS feature should the team enable to actively probe target health at a configurable interval?
136A cloud engineer is investigating why a nightly batch job on Amazon EC2 takes far longer than expected. CloudWatch shows the instance's CPU and memory usage are low throughout the run, but the job performs many small reads against an Amazon EBS gp3 volume. The engineer wants to reduce the time the job spends waiting on storage. Which action should the engineer take?
137A cloud operations team supports a latency-sensitive application running on Amazon EC2 instances. Users in a remote region report that responses are slow even though the application's own metrics show normal processing times. The team wants to continuously measure the network path between the users' region and the application endpoint, capturing round-trip latency and packet loss without modifying the application. Which AWS service should they use?
138A cloud administrator is responsible for a production account and needs to ensure that an Amazon S3 bucket containing sensitive data cannot be made public, even by an administrator. The administrator wants a preventive control that blocks public access at the bucket and account level. Which action should the administrator take?
139A cloud administrator is responsible for a Microsoft Azure environment. The administrator needs to ensure that virtual machine (VM) disks are backed up daily and that backups are retained for 30 days. The administrator also needs to be able to restore individual files from the backups. Which TWO actions should the administrator take to meet these requirements? (Choose two.)
140A cloud engineer manages a Kubernetes cluster on Google Kubernetes Engine. A production Deployment repeatedly enters CrashLoopBackOff after a configuration change, and the engineer needs to inspect why the container is terminating without modifying the running workload. Which action should the engineer take?
141A cloud administrator is responsible for a Microsoft Azure environment where several production virtual machines must be backed up nightly. The recovery requirements state that backups must be retained for 90 days, that individual files must be restorable without recovering the entire VM, and that the backup data must be encrypted at rest. Which Azure Backup configuration should the administrator implement?
142A cloud administrator needs to grant a third-party auditing firm read-only access to compliance reports in an Amazon S3 bucket for a limited period. The firm's identity provider supports SAML 2.0. Which approach best meets the requirement with least administrative overhead?
143A cloud operations team is using AWS and needs to monitor the CPU utilization of a fleet of Amazon EC2 instances. The team wants to receive an alert when the average CPU utilization exceeds 80% for 5 consecutive minutes. Which AWS service should they use to create the alarm?
144A cloud administrator is standardizing infrastructure provisioning across teams and wants to enforce that all deployed resources carry a mandatory cost-center tag. The administrator needs non-compliant deployments to be rejected automatically across multiple accounts in an AWS Organization. Which control should be implemented?
145A cloud administrator is responsible for a Microsoft Azure environment with a hub-and-spoke network topology. The administrator needs to ensure that all traffic from the spoke virtual networks to the internet is routed through a network virtual appliance (NVA) in the hub virtual network for inspection. Which Azure feature should the administrator configure?
146A cloud engineer is troubleshooting a performance issue in a Microsoft Azure environment. The application runs on an Azure Virtual Machine Scale Set (VMSS) behind an Azure Load Balancer. Users report intermittent slow response times. The engineer suspects that the VMSS instances are experiencing high CPU utilization due to uneven traffic distribution. The engineer needs to collect and analyze performance data to identify the root cause. Which Azure feature should the engineer use to gain deep visibility into the performance of the VMSS instances and the load balancer?
147A cloud operations team is configuring cost anomaly detection for a multi-account AWS organization. They want to be notified proactively when spending deviates from expected patterns and to attribute the deviation to the right team. Which TWO actions should they take? (Choose two.)
148A cloud engineer must ensure that a critical Azure virtual machine automatically restarts if the guest operating system becomes unresponsive, even when the Azure host is healthy. Which Azure feature should be configured?
149A cloud administrator is responsible for a set of Amazon EC2 instances that must be patched on a recurring schedule. The administrator wants to define a maintenance window, register the target instances, and have the system apply operating system patches automatically with a defined baseline. Which AWS service should the administrator use to accomplish this with the least operational overhead?
150A cloud operations team supports a microservices application running on Amazon ECS with AWS Fargate tasks. Users report that some requests fail intermittently, and the team suspects that containers are being terminated because they exceed resource limits or fail health checks. The team wants to collect the relevant diagnostic data to confirm the cause. (Choose two.)
151A cloud administrator is responsible for an application hosted on Amazon EC2 that stores session data in memory. The business requires that, in the event of an instance failure, a replacement instance can resume serving users with the existing session data intact and with minimal interruption. Which action should the administrator take to meet this requirement?
152A cloud operations team manages a fleet of Amazon EC2 instances running a stateless web tier behind an Application Load Balancer. The team wants to replace instances automatically when an instance fails an Elastic Load Balancing health check, without manual intervention, while keeping the desired capacity constant. Which AWS feature should the team configure to meet this requirement?
153A cloud administrator needs to centrally collect, search, and retain application and system logs from hundreds of Amazon EC2 instances and AWS services for troubleshooting and compliance. The administrator wants a managed service that stores logs in durable storage and allows ad hoc queries using a query language. Which AWS service should the administrator use?
154A cloud engineer manages an application running on Amazon ECS with the Fargate launch type. The application occasionally experiences task failures during deployment. The engineer wants to inspect the container's standard output and standard error to determine why a task stopped, without modifying the application to write to a file. Which action should the engineer take?
155A cloud operations team is responsible for a multi-tier application on AWS. They need to improve observability by correlating metrics, logs, and traces to diagnose performance issues across services. The team wants to use AWS services that natively support this correlation. Which TWO actions should the team take? (Choose two.)
Be able to map a stated RTO and RPO to a specific DR strategy and AWS service, and to configure metric alarms and centralized log queries. The single most important thing: read the RTO/RPO numbers first, then pick the cheapest strategy that still meets both.
The Courseiva CV0-004 question bank contains 155 questions in the Operations and Support domain, covering the 27% of the exam attributed to this domain in the official CompTIA blueprint. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Operations and Support domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included