Courseiva

CV0-004 · domain

Troubleshooting

The Troubleshooting domain covers diagnosing cloud workloads across compute, storage, and networking. Questions present symptoms like failed health checks, container restarts, VM latency, or high CPU ready time, then ask you to identify the cause, sequence remediation steps, or select the right tool. Expect scenario-based items referencing orchestrators, hypervisors, load balancers, and cloud monitoring services.

73 questions20 easy27 medium26 hard

Focused practice

Practice Troubleshooting questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Troubleshooting

You must read logs, metrics, and health check output to find the true fault, then choose the correct fix or sequence. The single most important thing: distinguish symptoms from root cause before changing configuration or scaling resources.

Diagnosing container exit codes and orchestrator restart loops using logs and health check configuration

Reading hypervisor CPU ready time and memory ballooning metrics to identify host contention

Using cloud monitoring, metrics, and load balancer health checks to isolate intermittent latency

Configuring object storage versioning and lifecycle policies to meet retention and cost requirements

Watch out for

Common Troubleshooting exam traps

  • ▸Blaming the application for container restarts when the real cause is a failing readiness or liveness probe threshold.
  • ▸Assuming a single VM has enough capacity during peak traffic instead of scaling out behind a load balancer.
  • ▸Confusing CPU ready time with guest CPU usage, leading to resizing the VM instead of reducing host contention.

Question index

All Troubleshooting questions (73)

Click any question to see the full explanation, or start a practice session above.

1

A company uses a cloud-based load balancer to distribute traffic to web servers. Recently, a new security policy was applied that restricts traffic to certain geographic regions. Users from an allowed region report they cannot access the website. The load balancer status shows health checks are passing. What should the administrator check?

Medium
2

Which THREE are common reasons why a cloud database instance may become unreachable?

Hard
3

Sequence the steps to set up a cloud storage bucket with versioning and lifecycle policies.

Medium
4

A cloud engineer receives an alert that the root filesystem (/) is at 93% usage. The /data volume has plenty of free space. The application stores logs in /var/log/app/ on the root filesystem. Which of the following is the BEST long-term solution?

Medium
5

A cloud database cluster is experiencing replication lag. The primary node shows high write activity, and the replicas are on different availability zones. Which of the following is the most likely cause?

Hard
6

A cloud administrator is troubleshooting why a newly launched VM did not complete its initialization. According to the exhibit, what is the most likely cause?

Hard
7

A web application is deployed across multiple availability zones behind a load balancer. The administrator notices that all traffic is being routed to instances in only one availability zone, causing performance issues. The load balancer is configured to distribute traffic across all zones evenly. What is the most likely cause?

Hard
8

A cloud administrator is troubleshooting a failed deployment of a new application version using a continuous integration/continuous deployment (CI/CD) pipeline. The pipeline fails at the 'test' stage. What is the first step the administrator should take?

Easy
9

A cloud engineer is troubleshooting a containerized application deployed on a managed Kubernetes service. Pods are repeatedly failing to start with the status 'CrashLoopBackOff'. The engineer has confirmed that the container image exists and the pod specification is valid. Which command should the engineer use to view the most recent logs from the previous instance of the crashing container?

Hard
10

A cloud administrator is troubleshooting a database performance issue in a cloud environment. The database is hosted on a virtual machine with a high-performance SSD. Users report slow query responses. Monitoring shows high disk I/O wait and low CPU utilization. Which action should the administrator take to improve performance?

Hard
11

A cloud administrator receives an alert that a virtual machine (VM) is unresponsive. The VM is hosted on a hypervisor that shows high CPU ready time. Which of the following is the most likely cause?

Medium
12

A user reports that they cannot connect to a RDS database instance from their application. The security group for the RDS instance allows inbound traffic on port 3306 from the application server's security group. What should the administrator check NEXT?

Easy
13

A cloud administrator is troubleshooting a network connectivity issue between two VPCs connected via a VPC peering connection. The administrator has verified that the route tables are correct and that the security groups allow traffic. However, instances in VPC A cannot ping instances in VPC B. Which TWO of the following could be causing the issue? (Choose TWO.)

Hard
14

A company runs a critical e-commerce application on a cloud platform. The architecture includes a load balancer in front of an auto scaling group of compute instances across two availability zones. The instances are in a private subnet and use a NAT gateway for outbound internet access. The application stores session data in a managed Redis cache cluster. During a flash sale, users report that the site is extremely slow and some requests time out. Monitoring shows the load balancer's latency metric is high, and the number of healthy hosts fluctuates. The CPU utilization on the compute instances averages 60% and memory averages 70%. The Redis cluster's CPU utilization is 90%, and its memory usage is 95%. The NAT gateway's metrics show high BytesOutToSource but no errors. Which of the following is the most likely cause of the performance issue?

Hard
15

A company has a cloud-based application that uses a relational database. The database team performs daily backups to an on-premises storage system using a VPN connection. Recently, backups have been failing with timeout errors. The network team confirms the VPN is up and stable. Which of the following is the MOST likely cause?

Easy
16

A cloud administrator is troubleshooting a performance degradation issue on a database server hosted in a public cloud. The server is experiencing high disk I/O wait times. The administrator suspects that the storage volume type is not optimized for the workload. Which two actions should the administrator take to address the issue? (Choose two.)

Medium
17

A cloud administrator notices that a virtual machine (VM) is running slowly. The hypervisor shows high CPU ready time for that VM. Which of the following is the most likely cause?

Easy
18

During a cloud migration, a database server is moved from on-premises to a cloud-managed database service. After migration, the application team reports that some queries are running slower than before. The database CPU utilization is low. What is the most likely cause?

Medium
19

A company has a three-tier application in a cloud VPC: web servers in a public subnet, application servers in a private subnet, and database servers in a private subnet. The web servers can connect to the application servers, but the application servers cannot connect to the database servers. The security groups are configured as follows: - Web SG: inbound HTTP from 0.0.0.0/0, outbound all - App SG: inbound HTTP from Web SG, outbound all - DB SG: inbound MySQL from App SG, outbound all What is the most likely cause of the connectivity issue?

Medium
20

A cloud administrator manages a three-tier application in a public cloud. After a change window, users can reach the web front end, but every request to the API tier returns HTTP 504 Gateway Timeout. The web tier and API tier are in different subnets, and the API instances report healthy in the load balancer target group. Which action should the administrator take FIRST to isolate the fault?

Medium
21

A cloud administrator is troubleshooting a database failover issue. The database is a managed service with a primary and standby replica in different availability zones. The application uses a read-write endpoint. During a recent maintenance event, the primary database failed over automatically, but the application experienced a 10-minute outage. The administrator checks the failover logs and sees that it completed within 2 minutes. What is the most likely cause of the extended outage?

Hard
22

A company migrated to a hybrid cloud and users report slow access to files stored in the cloud. The on-premises network is 100 Mbps. What troubleshooting step should be taken?

Medium
23

A cloud administrator manages a SaaS-based CRM application integrated with an on-premises Active Directory via SAML 2.0. Users report intermittent authentication failures during peak hours (09:00-11:00), with error messages indicating 'SAML assertion validation failed'. The IdP logs show successful authentications, but the SP logs show signature validation errors. The IdP's signing certificate was rotated 30 days ago, and the SP metadata was updated 45 days ago. Which of the following is the MOST likely cause?

Medium
24

A cloud administrator is troubleshooting a VM whose performance metrics show high 'CPU ready' (or 'CPU steal') time, even though the VM's own CPU utilization is only 20%. The VM runs a latency-sensitive database and is hosted on a shared hypervisor. Which action is the MOST appropriate first step to resolve the performance issue?

Hard
25

A cloud engineer is troubleshooting performance issues in a virtualized environment. Which of the following tools would BEST help identify CPU contention on a hypervisor?

Easy
26

A company is implementing a cloud governance strategy. They need to ensure that all resources are tagged with cost center and environment, and any untagged resources are automatically remediated. Which of the following best practices should be applied?

Hard
27

A cloud application returns HTTP 503 errors during high traffic. The application runs on VMs behind a load balancer. Which action is most likely to resolve the issue?

Easy
28

An automated snapshot of a cloud VM is failing with the error 'Quota exceeded for resource snapshots'. What is the most likely cause?

Medium
29

A cloud administrator is troubleshooting a web application that uses a cloud load balancer. Users report intermittent 502 Bad Gateway errors. The administrator checks the load balancer's target group and sees that some targets are marked as unhealthy. The application runs on virtual machines behind the load balancer. Which action should the administrator take to resolve the issue?

Medium
30

A global company runs a SaaS application in multiple cloud regions. They use DNS-based global load balancing to route users to the nearest region. Recently, users in Asia are experiencing high latency and timeouts. The administrator checks the health of the Asian region's resources and finds everything operational. Latency measurements from a monitoring tool show that traffic from Asian users is being routed to the European region. What should the administrator investigate first?

Hard
31

A cloud administrator is troubleshooting a web application hosted on a cloud VM that is experiencing intermittent high latency. The administrator reviews the cloud provider's monitoring metrics and sees that the VM's CPU utilization is consistently around 30%, memory usage is 40%, and network throughput is well below the instance's limit. Which factor is the most likely cause of the latency?

Medium
32

A small business hosts a web application on a single cloud server. The server has 2 vCPUs and 4 GB RAM. Recently, the application crashes when the number of concurrent users exceeds 50. The administrator checks the system logs and finds out-of-memory (OOM) errors. What is the best course of action to resolve this issue without redesigning the application?

Easy
33

A cloud administrator is troubleshooting a virtual machine (VM) in a public cloud that has become unresponsive. The administrator cannot SSH into the VM, and the cloud provider's console shows the VM is running. The administrator suspects the VM's OS has hung. Which action should the administrator take to regain access to the VM with minimal data loss?

Easy
34

A cloud administrator sees the output above when troubleshooting a virtual machine that is unresponsive. The VM is critical and must be restored quickly. What should the administrator do first?

Hard
35

A cloud administrator notices that a virtual machine is consuming excessive CPU resources with no apparent workload. Which of the following should the administrator investigate FIRST to determine the cause?

Hard
36

A cloud engineer is troubleshooting an issue where an application running in a container on a Kubernetes cluster is unable to resolve DNS names. The cluster uses CoreDNS. The engineer checks the CoreDNS pod logs and sees no errors. Which of the following should the engineer check next?

Hard
37

A cloud engineer is troubleshooting a performance issue where a web server cluster experiences high latency during peak hours. The cluster uses an auto-scaling group behind a load balancer. Which THREE steps should the engineer take to identify the root cause?

Hard
38

A cloud administrator is configuring a Linux VM as a router. The iptables rules are shown. The administrator can SSH into the VM from the network but cannot forward traffic between interfaces. What is the most likely cause?

Easy
39

A cloud administrator receives reports that a newly deployed application stack fails health checks and is repeatedly replaced by the orchestrator. Logs show the container starts, then exits after a few seconds with no error. The container image runs a process that daemonizes and returns control to the shell. Which of the following is the MOST likely cause?

Easy
40

A cloud architect is designing a multi-tier application that must be resilient to the failure of an entire availability zone. Which of the following strategies BEST meets this requirement?

Medium
41

A cloud administrator is troubleshooting a containerized application deployed on a managed Kubernetes cluster. Pods are failing to start, and the events show 'FailedMount' errors for a persistent volume claim (PVC). The PVC is bound to a persistent volume (PV) that uses a storage class with a reclaim policy of Delete. Which two actions should the administrator take to resolve the issue? (Choose two.)

Hard
42

A company is migrating on-premises workloads to the cloud. They need to ensure high availability for a stateless web application across two availability zones. Which THREE components should be configured to meet this requirement?

Hard
43

A cloud administrator is troubleshooting a connectivity issue between two VPCs in the same region. Which TWO actions should the administrator verify? (Choose two.)

Medium
44

A cloud user is unable to connect to a web server VM from the internet after a security group rule was modified. The VM is running and can be pinged from other VMs in the same subnet. What is the most likely cause?

Easy
45

A cloud administrator notices that a virtual machine is unresponsive. The VM is running on a hypervisor host that shows high CPU utilization. What should the administrator do first?

Easy
46

During a disaster recovery test, a cloud administrator discovers that the standby database in a different region is not synchronized with the primary. The primary database uses asynchronous replication. What is the MOST likely reason for the sync failure?

Medium
47

A cloud administrator is investigating why a virtual machine is running slowly. The administrator checks the hypervisor performance metrics. Which TWO of the following metrics indicate CPU contention? (Choose TWO.)

Easy
48

A cloud administrator is troubleshooting a web application that is hosted on a VM in a public cloud. Users report that the application is intermittently unavailable. The administrator checks the cloud provider's status page and sees no ongoing incidents. Which two actions should the administrator take to diagnose the issue? (Choose two.)

Medium
49

A virtual machine in a cloud environment is experiencing high disk I/O latency. The administrator checks the performance metrics and sees that the disk queue length is consistently above 100. What is the best immediate action?

Easy
50

A cloud administrator is troubleshooting a web application that is hosted on multiple virtual machines behind a load balancer. Users report that the application is occasionally slow, but the load balancer health checks show all instances as healthy. The administrator suspects that one of the virtual machines is performing poorly. Which of the following should the administrator do to confirm this suspicion?

Easy
51

A cloud administrator is troubleshooting a web application hosted on a cloud virtual machine (VM) that is experiencing intermittent high latency during peak traffic hours. The application is deployed on a single VM instance with 4 vCPUs and 8 GB RAM, running a Linux OS. The VM is connected to a virtual network with a public IP. The administrator has verified that the application code is optimized and there are no memory leaks. CPU utilization remains below 50% during peaks, but network outbound traffic shows periodic spikes up to 500 Mbps. The VM's network interface is configured with a 1 Gbps bandwidth cap. The administrator suspects that the issue is related to network throttling or packet loss. Which of the following actions should the administrator take to resolve the issue?

Hard
52

Match each disaster recovery term to its definition.

Medium
53

A cloud engineer is troubleshooting a VM that is experiencing high latency. The VM is hosted on a hypervisor with other VMs. Which TWO metrics should the engineer review to identify if resource contention is occurring?

Medium
54

A cloud engineer is troubleshooting a serverless function that is intermittently failing with timeout errors. The function is triggered by an HTTP API and processes data from an external database. The function's timeout is set to 30 seconds. Which of the following is the MOST likely cause of the timeouts?

Easy
55

A cloud load balancer is not distributing traffic evenly to backend servers. All servers pass health checks. Which of the following is the most likely cause?

Medium
56

A cloud administrator is troubleshooting slow performance on a managed relational database. Read queries against a read replica are fast, but write operations on the primary node take far longer than during testing. Monitoring shows the primary's CPU is moderate, disk queue depth is high, and provisioned IOPS are consistently saturated. Which action should the administrator take to resolve the bottleneck?

Medium
57

A cloud administrator runs a deployment script that creates multiple resources using Infrastructure as Code (IaC). The script fails with a "400 Bad Request" error when attempting to create a storage account. Which troubleshooting step should the administrator take first?

Easy
58

A cloud engineer is troubleshooting an issue where users cannot connect to a web application hosted on a cloud VM. The VM's security group allows HTTP (port 80) from 0.0.0.0/0, and the VM's OS firewall is disabled. The engineer can ping the VM's public IP from the internet. What is the most likely cause of the issue?

Medium
59

A cloud engineer notices that an application is running slower than expected. Monitoring shows that the CPU utilization is consistently below 30%, but memory usage is at 95%. Which of the following is the most likely cause of the performance issue?

Easy
60

A cloud administrator is managing a multi-tier application in a public cloud. The database tier is hosted on a VM with a persistent disk. Users report that the application is slow, and the administrator notices that disk I/O latency is high. The VM's disk is a standard network-attached storage volume. Which action should the administrator take to improve disk performance?

Medium
61

A cloud administrator is investigating a sudden increase in latency for a microservices application. Distributed traces show that a single downstream service's response time grew from 20 ms to 2 seconds, and its CPU utilization remains low at 15 percent. The service makes calls to an external third-party API. Which of the following is the MOST likely cause?

Hard
62

After deploying a new application version, users get 503 errors. The application runs on Kubernetes in a private cloud. What is the most likely cause?

Hard
63

A cloud engineer is troubleshooting a containerized application deployed on a managed Kubernetes cluster. The application pods are repeatedly restarting with 'OOMKilled' status. The engineer reviews the pod specification and sees that the memory request is 512Mi and the memory limit is 1Gi. The application is a Java-based service with a heap size set to 768Mi. Node metrics show that nodes have 8Gi of memory with 2Gi available. Which of the following is the MOST likely cause of the OOMKilled events?

Hard
64

A cloud administrator is investigating a sudden increase in egress charges for an application running on multiple Linux VMs in a public cloud. The application makes frequent calls to a third-party REST API over the public internet. Which action should the administrator take FIRST to identify the source of the unexpected egress traffic?

Medium
65

A cloud engineer is troubleshooting a storage performance issue. The storage is backed by a SAN with a mix of SSD and HDD drives. Which of the following metrics would BEST indicate that the storage subsystem is the bottleneck?

Hard
66

A cloud orchestration template fails to deploy resources with the error 'Resource limit exceeded'. The administrator has enough quota for all services. What is the most likely cause?

Hard
67

A cloud administrator manages a three-tier application in a public cloud. Users report that API calls from the web tier to the database tier fail with connection timeouts, but the database tier responds normally when queried from a bastion host on the same subnet. The web tier instances reside in a different subnet. Which of the following is the MOST likely cause?

Medium
68

A cloud administrator is troubleshooting a Linux virtual machine (VM) in a public cloud that is experiencing packet loss when communicating with another VM in the same subnet. The administrator runs 'ip -s link show eth0' and observes a high number of dropped packets on the receive side. The VM's CPU utilization is low, and the network interface is a paravirtualized driver (virtio). Which of the following is the MOST likely cause of the dropped packets?

Medium
69

A company's application is unable to connect to a managed cloud database. The database is deployed in a VPC with public accessibility disabled. The application runs on an EC2 instance in the same VPC. Which three troubleshooting steps should the administrator take? (Choose three.)

Medium
70

A cloud administrator is troubleshooting a performance issue where a web application is responding slowly. The application runs on virtual machines in a private cloud. The administrator has verified that CPU and memory utilization are within normal limits. Which TWO additional metrics should the administrator check to diagnose the issue?

Hard
71

A company uses a hybrid cloud model with an AWS Direct Connect connection to its on-premises network. Users report intermittent connectivity to cloud resources. A network engineer finds packet loss on the Direct Connect virtual interface. Which of the following should be checked FIRST to resolve the issue?

Hard
72

A cloud administrator is troubleshooting a newly deployed web application on a cloud VM. Users cannot access the application via its public IP address, but the administrator can SSH into the VM and access the application locally using curl http://localhost:8080. The VM's security group allows inbound traffic on port 22 and port 80. The application is listening on port 8080. Which of the following is the MOST likely cause of the issue?

Easy
73

A cloud administrator cannot deploy a new VM from a custom image. The deployment fails with an error stating 'Incompatible hypervisor version'. What is the most likely cause?

Easy

Frequently asked questions

What does the Troubleshooting domain cover on the CV0-004 exam?
You must read logs, metrics, and health check output to find the true fault, then choose the correct fix or sequence. The single most important thing: distinguish symptoms from root cause before changing configuration or scaling resources.
How many questions are in this domain?
This page lists all 73 Troubleshooting questions in the CV0-004 question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Troubleshooting questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
cloud-plus CLOUD-PLUS cloud troubleshooting Practice Questions