Courseiva

CCNA Troubleshooting Questions

73 questions · Troubleshooting · All types, answers revealed

1
MCQmedium

A company uses a cloud-based load balancer to distribute traffic to web servers. Recently, a new security policy was applied that restricts traffic to certain geographic regions. Users from an allowed region report they cannot access the website. The load balancer status shows health checks are passing. What should the administrator check?

A.The DNS resolution for the website
B.The SSL certificate expiration
C.The web server logs for application errors
D.The load balancer's access control lists (ACLs)
AnswerD

Geographic restrictions are enforced through load balancer ACLs, which filter client source IP ranges independently of backend health. Since health checks pass, the backend is fine; the ACL is the layer silently dropping traffic from the supposedly allowed region.

Why this answer

Geographic restrictions on a load balancer are typically implemented via access control lists (ACLs). Since health checks are passing, the web servers are functional, so the issue lies in the load balancer's ACLs blocking traffic from the allowed region. Option A is wrong: DNS resolution would affect all users similarly, not just those from a specific region.

Option B is wrong: SSL certificate issues would generate browser warnings or errors, not complete inaccessibility. Option C is wrong: web server logs are irrelevant as the traffic is not reaching the servers due to the ACL block.

2
Multi-Selecthard

Which THREE are common reasons why a cloud database instance may become unreachable?

Select 3 answers
A.Firewall rules blocking the database port
B.Incorrect connection string in the application
C.Storage volume is full on the database server
D.Database service not started
E.Hypervisor maintenance causing VM reboot
AnswersA, B, D

Security groups and network ACLs act as stateful or stateless packet filters; if the database listener port (for example 3306 or 5432) is not permitted inbound from the client subnet, connection attempts time out even though the instance itself is healthy and running.

Why this answer

Option A is correct because a cloud database instance listens on a specific port (e.g., 3306 for MySQL, 5432 for PostgreSQL, 1433 for SQL Server), and security group or firewall rules that fail to allow inbound traffic on that port will make the instance unreachable even though it is running. Option B is correct because an incorrect connection string—wrong hostname/endpoint, port, database name, or credentials—prevents the application from establishing a TCP session to the database, which is one of the most common causes of apparent unreachability. Option D is correct because if the database service/daemon (e.g., mysqld, postgresql, sqlservr) is not started or has crashed, the instance will refuse connections on its listening port.

Option C is not a typical cause of unreachability; a full storage volume usually causes write failures, errors, or degraded performance rather than making the instance completely unreachable. Option E is not a common reason either, since hypervisor maintenance typically triggers live migration or a planned reboot, and even if a VM reboots, the database would normally come back online automatically rather than remain unreachable.

Exam trap

CV0-004 often tests the distinction between 'unreachable' and 'slow' or 'failing writes'; candidates may choose storage full or hypervisor maintenance because they sound like availability issues, but the question specifically asks for reasons the instance cannot be reached at all.

3
Drag & Dropmedium

Sequence the steps to set up a cloud storage bucket with versioning and lifecycle policies.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

Create bucket, enable versioning, add lifecycle rules for transitions and deletions, then test.

4
MCQmedium

A cloud engineer receives an alert that the root filesystem (/) is at 93% usage. The /data volume has plenty of free space. The application stores logs in /var/log/app/ on the root filesystem. Which of the following is the BEST long-term solution?

A.Move the /var/log/app directory to the /data partition and create a symlink
B.Increase the size of the root filesystem
C.Delete the /data partition and merge it with root
D.Configure log rotation to delete logs more frequently
AnswerA

Relocating /var/log/app onto the /data partition and symlinking it back frees the root filesystem permanently, since logs then consume /data's capacity instead. This satisfies the long-term requirement, unlike log rotation or deletion, which only delay reoccurrence as logs regrow.

Why this answer

Moving the /var/log/app directory to the /data partition and creating a symlink is the best long-term solution because it permanently relocates the log data to a volume with ample free space without requiring application reconfiguration. The symlink (/var/log/app -> /data/app) makes the application continue to write to the same logical path, while the actual storage is on the /data filesystem. This resolves the root filesystem capacity issue without altering the application's logging behavior or risking data loss.

Exam trap

CompTIA often tests the misconception that increasing filesystem size or deleting partitions is a valid long-term fix, when in reality the correct approach is to relocate data to a separate volume using a symlink or mount bind.

How to eliminate wrong answers

Option B is wrong because increasing the size of the root filesystem only provides a temporary fix and does not address the underlying issue of log growth; it may also be impractical if the underlying disk or LVM has no free extents. Option C is wrong because deleting the /data partition and merging it with root is destructive, risks data loss on /data, and violates the principle of separating application data from the OS filesystem. Option D is wrong because configuring log rotation to delete logs more frequently reduces historical data needed for troubleshooting and compliance, and does not prevent future root filesystem exhaustion if log volume continues to grow.

5
MCQhard

A cloud database cluster is experiencing replication lag. The primary node shows high write activity, and the replicas are on different availability zones. Which of the following is the most likely cause?

A.Replication is configured as synchronous.
B.Network latency between the primary and replica zones is high.
C.The replica nodes have insufficient storage.
D.The primary node's vCPU is over-allocated.
AnswerB

High inter-zone network latency directly delays log shipping from primary to replica, so each replica applies writes later than the primary commits them. Since the replicas sit in different availability zones, cross-zone round-trip time is the bottleneck driving the observed replication lag, not write volume alone.

Why this answer

Network latency between availability zones is a common cause of replication lag in asynchronous replication setups, especially when the primary has high write activity. Option A is incorrect because synchronous replication would cause the primary to wait for acknowledgment from replicas, leading to write slowdown rather than lag on replicas. Option C is incorrect because insufficient storage on replicas would typically cause disk-full errors, not replication lag.

Option D is incorrect because vCPU over-allocation on the primary primarily affects compute performance, not the replication process which is more sensitive to network and disk I/O.

6
MCQhard

A cloud administrator is troubleshooting why a newly launched VM did not complete its initialization. According to the exhibit, what is the most likely cause?

A.Cloud-init is not installed on the VM
B.The package repository is not configured correctly
C.The cloud-init user data script contains a syntax error
D.The VM does not have internet access
AnswerB

A misconfigured package repository prevents the cloud-init script from installing required packages, so initialisation halts before completion. This directly satisfies the stem's constraint: the VM launched but never finished bootstrapping, indicating a dependency fetch failure rather than a provisioning or networking fault.

Why this answer

The error 'E: Unable to locate package python3-pip' indicates that the package repository is not configured correctly or the package does not exist in the configured sources. This prevents cloud-init from installing the specified package, so the most likely cause is an incorrect repository configuration. Therefore, option B is correct.

Option A is incorrect because cloud-init is executing commands, showing it is installed. Option C is incorrect because the error is about a missing package, not a syntax error. Option D is incorrect because the command ran successfully, indicating no network issues; the problem is repository configuration.

7
MCQhard

A web application is deployed across multiple availability zones behind a load balancer. The administrator notices that all traffic is being routed to instances in only one availability zone, causing performance issues. The load balancer is configured to distribute traffic across all zones evenly. What is the most likely cause?

A.The firewall rules for the load balancer only allow traffic from one zone.
B.The instances in the other zones are marked as unhealthy due to failing health checks.
C.The route table for the subnets in the other zones is missing a default route.
D.The listener rules are configured to forward traffic to a single backend pool.
AnswerB

When cross-zone load balancing is enabled, the load balancer should distribute traffic across all availability zones. However, if instances in other zones fail health checks, they are marked as unhealthy and removed from the target group's rotation, causing all traffic to go to healthy instances in only one zone.

Why this answer

If the load balancer is configured to distribute traffic evenly across all availability zones but traffic only reaches one zone, the most likely cause is that instances in the other zones are failing health checks and are marked unhealthy. Load balancers only route to healthy backends, so unhealthy instances in other zones are excluded, concentrating traffic in the one healthy zone. This matches the symptom of uneven distribution despite even configuration.

Exam trap

CV0-004 often tests the assumption that even load balancer configuration guarantees even traffic distribution — candidates overlook that health check status dynamically removes backends, so runtime health, not configuration, explains uneven routing.

How to eliminate wrong answers

Option A is wrong because firewall rules typically apply to the load balancer's frontend or the backend subnet as a whole, not per-availability-zone in a way that would selectively block one zone while the LB config says even distribution; also, if the firewall blocked a zone, health checks would fail there too, but the question points to health status as the differentiator. Option C is wrong because a missing default route in other zones' subnets would prevent outbound traffic but would not necessarily cause the load balancer to stop routing to them — and again, health check failures would be the observable cause. Option D is wrong because if listener rules forwarded to a single backend pool, the question's premise that the LB is configured to distribute across all zones evenly would be contradicted; the scenario states the configuration is correct, so the cause must be runtime health state.

8
MCQeasy

A cloud administrator is troubleshooting a failed deployment of a new application version using a continuous integration/continuous deployment (CI/CD) pipeline. The pipeline fails at the 'test' stage. What is the first step the administrator should take?

A.Re-run the pipeline
B.Increase the timeout of the test stage
C.Roll back to the previous version
D.Check the test logs for specific errors
AnswerD

The pipeline halts at the test stage, so the failure detail resides in that stage's output. Reviewing the test logs reveals the specific assertion, dependency or environment error before any remediation, avoiding speculative changes to code or pipeline configuration that could mask the actual fault.

Why this answer

When a CI/CD pipeline fails at the test stage, the first diagnostic step is to examine the test logs to identify the specific error — whether it is a failed assertion, a missing dependency, a timeout, or an environment issue. Logs provide the evidence needed to determine root cause before taking any corrective action. Jumping to re-runs or rollbacks without understanding the failure wastes time and may mask a recurring defect.

Exam trap

CV0-004 often tests the temptation to jump to remediation (re-run, rollback, increase timeout) before performing the fundamental troubleshooting step of reading logs to identify the actual error.

How to eliminate wrong answers

Option A is wrong because re-running the pipeline without diagnosing the failure is a blind action; if the failure is deterministic (e.g., a broken test or code defect), it will simply fail again. Option B is wrong because increasing the test stage timeout assumes the failure is due to slow execution, which is only one of many possible causes and is not supported by evidence yet. Option C is wrong because rolling back to the previous version is a remediation step, not a diagnostic step, and it discards the new version's changes without understanding why the tests failed.

9
MCQhard

A cloud engineer is troubleshooting a containerized application deployed on a managed Kubernetes service. Pods are repeatedly failing to start with the status 'CrashLoopBackOff'. The engineer has confirmed that the container image exists and the pod specification is valid. Which command should the engineer use to view the most recent logs from the previous instance of the crashing container?

A.kubectl logs <pod-name> --previous
B.kubectl exec -it <pod-name> -- /bin/sh
C.kubectl describe pod <pod-name>
D.kubectl get events --field-selector involvedObject.name=<pod-name>
AnswerA

The --previous flag retrieves logs from the previous instance of a container in a pod that has restarted. In a CrashLoopBackOff scenario, the current container may not have produced logs yet, but the previous instance likely contains the error that caused the crash. This command is essential for diagnosing why the container is failing to start.

Why this answer

In a CrashLoopBackOff situation, the container is restarting repeatedly, so the current logs may be empty or incomplete. The --previous flag allows the engineer to access logs from the last terminated instance, which typically contains the error that led to the crash. This is the most direct way to diagnose application-level failures in Kubernetes.

Exam trap

The trap here is relying on kubectl describe or events for application logs, when the previous container instance's logs are the key to understanding the crash cause.

10
MCQhard

A cloud administrator is troubleshooting a database performance issue in a cloud environment. The database is hosted on a virtual machine with a high-performance SSD. Users report slow query responses. Monitoring shows high disk I/O wait and low CPU utilization. Which action should the administrator take to improve performance?

A.Check the VM's disk queue depth and consider increasing provisioned IOPS.
B.Migrate the database to a VM with a larger memory allocation.
C.Enable read caching on the database to reduce disk reads.
D.Increase the VM's CPU count to handle more concurrent queries.
AnswerA

High disk I/O wait with low CPU indicates the storage subsystem is the bottleneck. Cloud VMs have limits on IOPS and throughput. If the disk queue is saturated, increasing provisioned IOPS (if using a cloud disk like AWS EBS or Azure Disk) can improve performance. Checking queue depth helps confirm the bottleneck before making changes.

Why this answer

High disk I/O wait with low CPU utilization points to storage as the bottleneck. Cloud disks have IOPS limits; if the workload exceeds provisioned IOPS, performance degrades. Checking queue depth confirms saturation, and increasing provisioned IOPS or upgrading the disk type resolves the issue.

CPU and memory changes do not address the storage bottleneck, and read caching may not help if the workload is write-intensive.

Exam trap

The trap here is assuming that more CPU or memory will fix slow queries, when the metrics clearly indicate a disk I/O bottleneck.

11
MCQmedium

A cloud administrator receives an alert that a virtual machine (VM) is unresponsive. The VM is hosted on a hypervisor that shows high CPU ready time. Which of the following is the most likely cause?

A.Insufficient memory allocated to the VM
B.Network latency between the VM and storage
C.Disk I/O contention from other VMs
D.Over-provisioning of vCPUs on the hypervisor
AnswerD

High CPU ready time means vCPUs wait in the hypervisor's run queue before receiving physical CPU cycles. Over-provisioning vCPUs across guests on the same host oversubscribes physical cores, so each VM waits longer, presenting as an unresponsive guest.

Why this answer

High CPU ready time indicates that the VM is ready to execute instructions but is waiting for the hypervisor to schedule physical CPU time. This is a classic symptom of over-provisioning vCPUs, where the total number of vCPUs assigned to all VMs exceeds the available physical cores, causing contention at the hypervisor scheduler level.

Exam trap

The trap here is that candidates confuse high CPU ready time with high CPU usage or memory pressure, but ready time is a hypervisor-level scheduling delay, not a guest OS metric, and is directly tied to vCPU over-provisioning.

How to eliminate wrong answers

Option A is wrong because insufficient memory would typically cause swapping or ballooning, not high CPU ready time, which is a CPU scheduling metric. Option B is wrong because network latency between the VM and storage affects storage I/O latency, not CPU scheduling, and would manifest as high disk latency or queue depth. Option C is wrong because disk I/O contention from other VMs would result in high disk queue length or latency, not CPU ready time, which is a measure of CPU starvation.

12
MCQeasy

A user reports that they cannot connect to a RDS database instance from their application. The security group for the RDS instance allows inbound traffic on port 3306 from the application server's security group. What should the administrator check NEXT?

A.IAM policy attached to the RDS instance
B.Network ACL rules for the RDS subnet
C.Route table entries for the RDS subnet
D.Outbound security group rules on the RDS instance
AnswerB

Security groups are stateful and already permit port 3306, so the next layer to verify is the stateless network ACL on the RDS subnet, which must allow inbound 3306 and outbound ephemeral return traffic. This satisfies the requirement to check the next likely blocker.

Why this answer

The security group for the RDS instance allows inbound traffic on port 3306 from the application server's security group, which is correct. However, network ACLs (NACLs) are stateless and can block traffic even if security groups allow it. The administrator should check the NACL rules for the RDS subnet to ensure that inbound and outbound rules permit traffic on port 3306 and the ephemeral ports for return traffic.

This is the next logical step after verifying security group rules.

Exam trap

CV0-004 often tests the difference between stateful security groups and stateless NACLs; candidates may forget that NACLs require explicit outbound rules for return traffic, leading them to overlook NACL checks.

How to eliminate wrong answers

Option A is wrong because IAM policies control access to RDS management APIs, not network connectivity to the database. Option C is wrong because route table entries affect routing between subnets; if the application server and RDS are in the same VPC, the default route exists, and if they are in different VPCs, peering or other connectivity would be needed, but the scenario implies they are in the same VPC or already connected. Option D is wrong because outbound security group rules on the RDS instance are not relevant for inbound connections; security groups are stateful, so return traffic is automatically allowed if inbound is allowed.

13
Multi-Selecthard

A cloud administrator is troubleshooting a network connectivity issue between two VPCs connected via a VPC peering connection. The administrator has verified that the route tables are correct and that the security groups allow traffic. However, instances in VPC A cannot ping instances in VPC B. Which TWO of the following could be causing the issue? (Choose TWO.)

Select 2 answers
A.Network ACLs in VPC B are blocking inbound ICMP
B.Security groups in VPC A are blocking inbound ICMP
C.Host-based firewall on the target instance is blocking ping
D.VPC peering connection does not support ICMP
E.Route tables are misconfigured
AnswersA, C

Network ACLs are stateless subnet-level filters separate from security groups. Even with security groups allowing traffic, an inbound rule in VPC B's NACL denying ICMP would silently drop echo requests, explaining why instances in VPC A cannot ping VPC B.

Why this answer

Option A is correct because network ACLs are stateless subnet-level filters that must explicitly allow inbound ICMP (for IPv4, protocol 1, type 8 echo request) and outbound ICMP echo reply (type 0); even though the administrator verified security groups, a restrictive NACL in VPC B would silently drop ping traffic. Option C is correct because a host-based firewall (e.g., iptables, Windows Firewall, or firewalld) running on the target instance operates above the VPC layer and can block ICMP echo requests regardless of correct route tables and security group rules. Option B is not correct because the scenario states security groups already allow traffic, and security groups are stateful, so inbound ICMP would be permitted if configured.

Option D is not correct because VPC peering fully supports ICMP traffic between peered VPCs; it is not a protocol limitation. Option E is not correct because the administrator has already verified that the route tables are correct.

Exam trap

CV0-004 often tests the layered nature of cloud networking — candidates fixate on route tables and security groups and forget that network ACLs and host-based firewalls are separate, independent layers that can block traffic.

14
MCQhard

A company runs a critical e-commerce application on a cloud platform. The architecture includes a load balancer in front of an auto scaling group of compute instances across two availability zones. The instances are in a private subnet and use a NAT gateway for outbound internet access. The application stores session data in a managed Redis cache cluster. During a flash sale, users report that the site is extremely slow and some requests time out. Monitoring shows the load balancer's latency metric is high, and the number of healthy hosts fluctuates. The CPU utilization on the compute instances averages 60% and memory averages 70%. The Redis cluster's CPU utilization is 90%, and its memory usage is 95%. The NAT gateway's metrics show high BytesOutToSource but no errors. Which of the following is the most likely cause of the performance issue?

A.The NAT gateway is throttling traffic due to bandwidth limits
B.The managed Redis cache cluster is overloaded and becoming a bottleneck for session lookups
C.The auto scaling group is not scaling quickly enough due to cooldown periods
D.The load balancer's idle timeout setting is too low, causing premature connection drops
AnswerB

Redis CPU at 90% and memory at 95% indicate the cache cluster is saturated, so session lookups slow or fail, cascading into high load balancer latency and fluctuating healthy hosts. This satisfies the observed bottleneck, since compute CPU and memory remain moderate.

Why this answer

The managed Redis cache cluster is the most likely bottleneck because its CPU utilization is at 90% and memory usage at 95%, indicating it is near capacity. Since the application stores session data in Redis, high latency and timeouts during a flash sale are consistent with an overloaded session store that cannot keep up with request volume, causing the load balancer to experience increased latency and healthy host fluctuations as sessions fail to be retrieved or written.

Exam trap

The trap here is that candidates may focus on the NAT gateway or auto scaling group because they are common bottlenecks, but the key clue is the Redis cluster's high CPU and memory metrics, which directly correlate with session store performance issues in a stateful application.

How to eliminate wrong answers

Option A is wrong because the NAT Gateway shows high BytesOutToSource but no errors, and NAT Gateway bandwidth limits are typically high (up to 10 Gbps per AZ) and would cause packet drops or errors if throttled, not just high latency. Option C is wrong because the Auto Scaling group's cooldown periods could delay scaling, but the EC2 instances are only at 60% CPU and 70% memory, which are not saturated, so scaling is not the primary issue. Option D is wrong because the ALB's idle timeout setting (default 60 seconds) controls how long the ALB keeps a connection open without data; premature connection drops would manifest as immediate disconnects, not high latency and timeouts.

15
MCQeasy

A company has a cloud-based application that uses a relational database. The database team performs daily backups to an on-premises storage system using a VPN connection. Recently, backups have been failing with timeout errors. The network team confirms the VPN is up and stable. Which of the following is the MOST likely cause?

A.The database service is not responding
B.The VPN bandwidth is insufficient for the backup data volume
C.The VPN tunnel is not properly configured
D.The on-premises firewall is blocking the backup port
AnswerB

Insufficient VPN bandwidth saturates the tunnel during large backup transfers, causing TCP retransmissions and eventual timeout errors despite the VPN remaining up and stable. A relational database's daily full or incremental backup volume can exceed the tunnel's throughput, so the constraint of a stable-but-limited VPN link is satisfied by this explanation.

Why this answer

The VPN connection is confirmed stable, so tunnel configuration and firewall issues are unlikely. Backup timeout errors with large data volumes typically indicate insufficient bandwidth, causing the transfer to exceed the timeout threshold. The database service itself is responding (backups are attempted), ruling out service unavailability.

Exam trap

The trap here is that candidates assume a stable VPN means the link has sufficient capacity, but CompTIA often tests the distinction between connectivity (layer 3) and throughput (layer 4/performance), where a stable tunnel can still be too slow for large data transfers.

How to eliminate wrong answers

Option A is wrong because if the database service were not responding, backups would fail immediately with a connection error, not a timeout after data transfer begins. Option C is wrong because the network team confirmed the VPN is up and stable, meaning the tunnel is properly configured and operational. Option D is wrong because a firewall block would cause a consistent failure (e.g., connection refused), not intermittent timeouts, and the VPN tunnel encrypts traffic, making port-specific blocking less likely.

16
Multi-Selectmedium

A cloud administrator is troubleshooting a performance degradation issue on a database server hosted in a public cloud. The server is experiencing high disk I/O wait times. The administrator suspects that the storage volume type is not optimized for the workload. Which two actions should the administrator take to address the issue? (Choose two.)

Select 2 answers
A.Increase the size of the existing volume to automatically improve IOPS.
B.Change the storage volume to a higher-performance SSD-based volume type.
C.Increase the instance type to one with more memory.
D.Move the database to a volume with a higher IOPS provisioned.
E.Enable detailed monitoring for the volume to analyze I/O patterns.
AnswersB, D

Upgrading to a higher-performance SSD volume type, such as provisioned IOPS SSD, can significantly reduce disk I/O wait times by providing more IOPS and lower latency. This is a direct solution when the current volume type is the bottleneck. The administrator should evaluate the workload's IOPS requirements and choose an appropriate volume type. This action addresses the root cause of high I/O wait.

Why this answer

High disk I/O wait times indicate that the storage volume cannot keep up with the workload's demand. The most effective solutions are to change to a higher-performance volume type or provision a volume with higher IOPS. Both actions directly increase the storage's ability to handle I/O operations, reducing wait times.

Diagnostic steps like monitoring are useful but do not resolve the issue.

Exam trap

The trap here is assuming that increasing volume size automatically scales IOPS linearly, which is not true for all volume types and may not meet performance needs.

17
MCQeasy

A cloud administrator notices that a virtual machine (VM) is running slowly. The hypervisor shows high CPU ready time for that VM. Which of the following is the most likely cause?

A.High disk I/O latency on the datastore
B.Insufficient memory allocated to the VM
C.Overcommitted physical CPU resources on the host
D.Misconfigured virtual switch
AnswerC

High CPU ready time means the VM waited for physical cores while the scheduler ran other vCPUs. That occurs when the host's physical CPU is overcommitted, so the hypervisor cannot grant cycles promptly, directly explaining the slow VM.

Why this answer

High CPU ready time indicates that the VM is ready to execute instructions but is waiting for the physical CPU to become available. This is a classic symptom of CPU overcommitment, where the host has more virtual CPUs (vCPUs) assigned to VMs than physical cores, causing contention. Option C correctly identifies this as the most likely cause.

Exam trap

CompTIA often tests the distinction between CPU ready time and other performance metrics, trapping candidates who confuse high CPU ready time with memory pressure or storage latency.

How to eliminate wrong answers

Option A is wrong because high disk I/O latency would manifest as high disk queue depth or high kernel latency, not as CPU ready time. Option B is wrong because insufficient memory would cause ballooning or swapping, not CPU ready time. Option D is wrong because a misconfigured virtual switch would cause network connectivity issues or packet loss, not CPU scheduling delays.

18
MCQmedium

During a cloud migration, a database server is moved from on-premises to a cloud-managed database service. After migration, the application team reports that some queries are running slower than before. The database CPU utilization is low. What is the most likely cause?

A.The network latency between the application and the database has increased
B.The database is not indexed properly
C.The database connection pooling is misconfigured
D.The cloud database instance type has insufficient memory
AnswerA

Low CPU indicates the database itself is not the bottleneck, so the added round-trip latency between application and cloud-managed database is the likely cause. Moving off-premises lengthens the network path, slowing query response even when processing is idle.

Why this answer

When a database is moved to a cloud-managed service and queries slow down while CPU utilization is low, the most likely cause is increased network latency between the application and the database. Low CPU indicates the database is not compute-bound, so the delay is likely in the round-trip time for queries. This is common when the application and database are in different regions or availability zones.

Exam trap

CV0-004 often tests the correlation between low CPU and slow queries; candidates may blame indexing or instance size, but low CPU points to a non-compute bottleneck like network latency, which is common in cloud migrations.

How to eliminate wrong answers

Option B is wrong because improper indexing would typically cause high CPU utilization due to full table scans, not low CPU. Option C is wrong because misconfigured connection pooling would cause connection errors or overhead, but not necessarily slower queries with low CPU; it might cause timeouts. Option D is wrong because insufficient memory would lead to swapping or disk I/O, often increasing CPU or causing out-of-memory errors, not low CPU with slow queries.

19
MCQmedium

A company has a three-tier application in a cloud VPC: web servers in a public subnet, application servers in a private subnet, and database servers in a private subnet. The web servers can connect to the application servers, but the application servers cannot connect to the database servers. The security groups are configured as follows: - Web SG: inbound HTTP from 0.0.0.0/0, outbound all - App SG: inbound HTTP from Web SG, outbound all - DB SG: inbound MySQL from App SG, outbound all What is the most likely cause of the connectivity issue?

A.The database security group is missing an inbound rule for MySQL.
B.The application security group is missing an outbound rule for MySQL.
C.The network access control list (NACL) on the database subnet is blocking inbound traffic from the application subnet.
D.The web security group is blocking traffic to the database.
AnswerC

Security groups are stateful and already permit MySQL from the App SG, so the block must come from the stateless layer: the database subnet's NACL. Its inbound rules are evaluated independently of outbound, and a missing allow for the application subnet's CIDR on port 3306 drops the traffic.

Why this answer

The most likely cause is a network ACL (NACL) on the database subnet blocking inbound traffic from the application subnet. Security groups are stateful and allow return traffic, but NACLs are stateless and must explicitly allow both inbound and outbound traffic on ephemeral ports. If the NACL denies inbound MySQL (port 3306) from the app subnet's CIDR, the connection fails even though the security groups are correctly configured.

Exam trap

The trap is focusing only on security groups because they are commonly misconfigured; the exam tests whether you remember that NACLs are stateless and can block traffic even when security groups are correct.

How to eliminate wrong answers

Option A is wrong because the DB security group already has an inbound rule for MySQL from the App SG, as stated in the scenario. Option B is wrong because the App SG has outbound 'all', which includes MySQL traffic to the DB. Option D is wrong because the Web SG is not involved in the app-to-database connection path; the web tier only talks to the app tier.

20
MCQmedium

A cloud administrator manages a three-tier application in a public cloud. After a change window, users can reach the web front end, but every request to the API tier returns HTTP 504 Gateway Timeout. The web tier and API tier are in different subnets, and the API instances report healthy in the load balancer target group. Which action should the administrator take FIRST to isolate the fault?

A.Lower the web tier's idle timeout value so slow API responses are terminated sooner.
B.Reboot all API instances because the target group reports them healthy but the application process may have hung.
C.Verify the network ACL and security group rules that govern traffic from the web subnet to the API tier's listener port.
D.Increase the API tier's autoscaling maximum because the timeout indicates the tier is saturated with requests.
AnswerC

A 504 from the front end while API instances are healthy points to connectivity between tiers rather than an application crash. A security group or network ACL that silently drops the web-to-API flow would let the target group health check pass (if it originates inside the API subnet) while user requests time out. Checking these rules first is the fastest way to confirm or eliminate a network path fault.

Why this answer

Healthy API targets plus front-end 504s isolate the fault to the path between tiers, not to the API application itself. Security groups and network ACLs are the components that can silently drop traffic on the listener port while still allowing health checks, so reviewing those rules first is the correct diagnostic step. Reboots, scaling, and timeout tuning all act on the wrong layer and would delay identification of the real cause.

Exam trap

The trap here is assuming a 504 always means the backend application is overloaded, when it can equally mean traffic never reaches the backend at all.

21
MCQhard

A cloud administrator is troubleshooting a database failover issue. The database is a managed service with a primary and standby replica in different availability zones. The application uses a read-write endpoint. During a recent maintenance event, the primary database failed over automatically, but the application experienced a 10-minute outage. The administrator checks the failover logs and sees that it completed within 2 minutes. What is the most likely cause of the extended outage?

A.The application's database connection pool does not retry DNS resolution
B.The application was not configured to use multiple availability zones
C.The standby replica was not in sync
D.The failover triggered a change in the endpoint DNS record
AnswerA

After failover the read-write endpoint's DNS record points to the new primary, but a connection pool caching the old address keeps reconnecting to the failed node until its entries expire. This explains the outage far exceeding the two-minute failover time.

Why this answer

The most likely cause is that the application's connection pool caches the DNS resolution of the read-write endpoint. After failover, the DNS record points to the new primary, but the pool continues to use the old IP until the TTL expires or the connection is refreshed. This causes a prolonged outage beyond the actual failover time.

Exam trap

The trap is blaming the failover mechanism or replication; candidates overlook application-side DNS caching in connection pools as the cause of extended downtime.

How to eliminate wrong answers

Option B is wrong because the application uses a read-write endpoint, which is inherently multi-AZ; the issue is not lack of AZ configuration. Option C is wrong because the failover logs show it completed in 2 minutes, implying the standby was in sync. Option D is wrong because a DNS change is expected during failover; the problem is the application not picking up the change, not the change itself.

22
MCQmedium

A company migrated to a hybrid cloud and users report slow access to files stored in the cloud. The on-premises network is 100 Mbps. What troubleshooting step should be taken?

A.Enable compression on the cloud storage gateway
B.Check VPN bandwidth and latency
C.Increase cloud storage performance tier
D.Move files to on-premises storage
AnswerB

Checking VPN bandwidth and latency directly addresses the hybrid cloud bottleneck: the 100 Mbps on-premises link and any tunnel overhead cap throughput for cloud file access. Measuring both reveals whether encryption, routing or congestion is throttling transfers, satisfying the stem's constraint of diagnosing slow cloud file access across the hybrid connection.

Why this answer

The most likely bottleneck in a hybrid cloud scenario where users access cloud-stored files over a 100 Mbps on-premises connection is the VPN tunnel used to connect to the cloud. VPN bandwidth and latency directly affect file access speed, and issues such as insufficient tunnel capacity, high latency, or packet loss can cause slow transfers. Checking these metrics is the logical first troubleshooting step to identify whether the network path is the limiting factor.

Exam trap

CV0-004 often tests the misconception that slow cloud file access is always due to cloud storage performance, leading candidates to choose increasing the performance tier instead of first diagnosing the network path.

How to eliminate wrong answers

Option A is wrong because enabling compression on the cloud storage gateway may improve effective throughput but does not address the underlying network bottleneck; it is a optimization step, not a troubleshooting step, and may not be supported or effective if the data is already compressed. Option C is wrong because increasing the cloud storage performance tier addresses storage IOPS or throughput limits on the cloud side, but the on-premises network at 100 Mbps is a more immediate constraint, and the scenario does not indicate storage performance is the issue. Option D is wrong because moving files to on-premises storage defeats the purpose of the hybrid cloud migration and does not troubleshoot the slow access; it is a workaround, not a diagnostic step.

23
MCQmedium

A cloud administrator manages a SaaS-based CRM application integrated with an on-premises Active Directory via SAML 2.0. Users report intermittent authentication failures during peak hours (09:00-11:00), with error messages indicating 'SAML assertion validation failed'. The IdP logs show successful authentications, but the SP logs show signature validation errors. The IdP's signing certificate was rotated 30 days ago, and the SP metadata was updated 45 days ago. Which of the following is the MOST likely cause?

A.The IdP is not including the correct NameID format in the SAML assertion.
B.The SP's SAML assertion consumer service (ACS) URL is misconfigured.
C.The SP's metadata contains an outdated IdP signing certificate.
D.The IdP's clock is skewed relative to the SP, causing timestamp validation to fail.
AnswerC

The SP metadata was updated 45 days ago, but the IdP rotated its signing certificate 30 days ago. SAML signature validation relies on the certificate in the SP's trusted metadata; if it still holds the old certificate, assertions signed with the new key fail validation. This matches the symptom of successful IdP authentication but SP-side signature errors, and the timing discrepancy confirms the metadata is stale.

Why this answer

The IdP rotated its signing certificate 30 days ago, but the SP metadata was last updated 45 days ago, meaning the SP still trusts the old certificate. SAML signature validation requires the SP to use the current IdP signing certificate. The successful IdP authentications and SP signature errors confirm the SP cannot validate assertions signed with the new key.

Updating the SP metadata with the new certificate resolves the issue.

Exam trap

The trap here is assuming that because the IdP logs show successful authentications, the problem must be on the IdP side, when in fact the SP's stale metadata is the root cause.

24
MCQhard

A cloud administrator is troubleshooting a VM whose performance metrics show high 'CPU ready' (or 'CPU steal') time, even though the VM's own CPU utilization is only 20%. The VM runs a latency-sensitive database and is hosted on a shared hypervisor. Which action is the MOST appropriate first step to resolve the performance issue?

A.Configure CPU affinity to pin the VM to a specific physical core.
B.Enable CPU hot-add and dynamically add more cores during peak hours.
C.Increase the number of vCPUs assigned to the VM.
D.Migrate the VM to a dedicated host or a host with lower oversubscription.
AnswerD

High CPU ready time indicates the VM is waiting for physical CPU resources because the hypervisor is oversubscribed. Moving the VM to a dedicated host or a host with lower oversubscription reduces contention, directly addressing the root cause. This is the most appropriate first step because it targets the resource bottleneck without unnecessary changes to the VM's configuration.

Why this answer

High CPU ready time indicates the VM is waiting for physical CPU resources due to host oversubscription. The most effective first step is to reduce contention by moving the VM to a dedicated host or a host with lower oversubscription. Increasing vCPUs or enabling hot-add does not address the host-level bottleneck and can worsen scheduling delays.

CPU affinity is not a reliable fix for oversubscription.

Exam trap

The trap here is assuming that high CPU ready time means the VM needs more vCPUs, when it actually indicates the VM is waiting for physical CPU time due to host contention.

25
MCQeasy

A cloud engineer is troubleshooting performance issues in a virtualized environment. Which of the following tools would BEST help identify CPU contention on a hypervisor?

A.iperf
B.esxtop
C.ping
D.nslookup
AnswerB

esxtop runs on the ESXi host itself, exposing per-world CPU scheduling metrics such as %RDY, which directly quantifies the time a vCPU waits for physical cores. This satisfies the stem's requirement to identify CPU contention at the hypervisor layer, unlike guest-level tools that cannot see host scheduling pressure.

Why this answer

esxtop is VMware's real-time performance monitoring tool for ESXi hosts, and it exposes CPU metrics such as %RDY (ready time), which directly indicates CPU contention when VMs are waiting for physical CPU cycles. It is the standard tool for diagnosing hypervisor-level CPU scheduling issues.

Exam trap

CV0-004 often tests tool-to-purpose mapping — candidates confuse network tools (iperf, ping, nslookup) with hypervisor performance tools, so recognizing esxtop as the VMware-specific CPU contention tool is the key discriminator.

How to eliminate wrong answers

Option A is wrong because iperf measures network throughput and bandwidth, not CPU scheduling or contention. Option C is wrong because ping tests basic network reachability and round-trip latency, which is unrelated to hypervisor CPU contention. Option D is wrong because nslookup is a DNS query tool used to resolve names to IP addresses, with no visibility into CPU performance.

26
MCQhard

A company is implementing a cloud governance strategy. They need to ensure that all resources are tagged with cost center and environment, and any untagged resources are automatically remediated. Which of the following best practices should be applied?

A.Implement role-based access control to restrict resource creation
B.Set up budget alerts to notify when costs exceed thresholds
C.Create a manual audit process to check tags weekly
D.Use policy-as-code to enforce tagging and automatically apply tags to untagged resources
AnswerD

Policy-as-code evaluates resource definitions against tagging rules and triggers automatic remediation, applying missing cost centre and environment tags. This enforces the governance requirement continuously rather than relying on manual audits, satisfying the automatic remediation constraint.

Why this answer

Policy-as-code (e.g., Azure Policy, AWS Config Rules, or Open Policy Agent) allows you to define tagging requirements declaratively and automatically remediate non-compliant resources. This approach enforces governance in real-time without manual intervention, ensuring all resources are tagged with cost center and environment as specified.

Exam trap

The trap here is that candidates often confuse manual audit processes (Option C) with automated governance, failing to recognize that policy-as-code provides the required automatic remediation in real-time.

How to eliminate wrong answers

Option A is wrong because role-based access control (RBAC) restricts who can create resources but does not automatically tag or remediate untagged resources. Option B is wrong because budget alerts notify when costs exceed thresholds but do not enforce tagging or remediate untagged resources. Option C is wrong because a manual audit process is reactive, time-consuming, and does not provide automatic remediation, which is required by the question.

27
MCQeasy

A cloud application returns HTTP 503 errors during high traffic. The application runs on VMs behind a load balancer. Which action is most likely to resolve the issue?

A.Restart the web server service on one VM.
B.Change the DNS TTL to a lower value.
C.Increase the health check interval on the load balancer.
D.Add additional VMs to the backend pool.
AnswerD

HTTP 503 signals backend capacity exhaustion, so the load balancer has healthy targets but insufficient compute. Adding VMs to the backend pool increases aggregate capacity to absorb the traffic spike, directly addressing the overload causing the errors.

Why this answer

HTTP 503 Service Unavailable during high traffic indicates that the backend servers cannot handle the current load, often because all VMs are saturated or the load balancer has no healthy backends. Adding additional VMs to the backend pool increases capacity and distributes the load, directly addressing the root cause. This is the standard horizontal scaling response to traffic-induced 503 errors.

Exam trap

CV0-004 often tests whether candidates confuse DNS or health-check tuning with actual capacity scaling — the trap is picking a 'quick fix' like restarting a server instead of addressing the root cause of insufficient backend capacity.

How to eliminate wrong answers

Option A is wrong because restarting the web server on one VM may temporarily clear the error but does not increase capacity — the remaining VMs will still be overwhelmed under high traffic. Option B is wrong because lowering DNS TTL affects how quickly DNS changes propagate, but it does not add capacity or resolve backend overload. Option C is wrong because increasing the health check interval makes the load balancer check backends less frequently, which could actually delay detection of unhealthy nodes and worsen the situation — it does not add capacity.

28
MCQmedium

An automated snapshot of a cloud VM is failing with the error 'Quota exceeded for resource snapshots'. What is the most likely cause?

A.The snapshot is being created during a backup window.
B.The maximum number of snapshots allowed has been reached.
C.The snapshot retention policy is set too high.
D.The VM's disk is too full to create a snapshot.
AnswerB

The subscription's snapshot quota has been exhausted, meaning the maximum number of snapshots permitted for that region or resource group is already allocated. The error explicitly names the quota for the snapshots resource, so the limit itself is the blocker. Deleting unused snapshots or requesting a quota increase resolves the failure.

Why this answer

The error 'Quota exceeded for resource snapshots' directly indicates that the cloud provider's limit on the number of snapshots for that resource (e.g., per volume, per region, or per account) has been reached. Cloud platforms enforce quotas to prevent resource exhaustion, and once the maximum count is hit, further snapshot creation fails until older snapshots are deleted or a quota increase is requested. This is a hard limit, not a transient condition like a backup window or disk fullness.

Exam trap

CV0-004 often tests the distinction between quota limits and other snapshot failure causes, so candidates may confuse retention policies or disk space with the actual quota error.

How to eliminate wrong answers

Option A is wrong because backup windows are scheduling constructs and do not cause quota errors; snapshots can be taken at any time unless restricted by policy, but the error explicitly mentions quota. Option C is wrong because a retention policy controls how long snapshots are kept, not how many can exist at once; a high retention policy could indirectly lead to hitting the quota, but the error itself is about the quota limit, not the policy setting. Option D is wrong because a full disk might cause snapshot failures due to lack of space for delta changes, but the error message would indicate insufficient space or I/O errors, not a quota exceeded condition.

29
MCQmedium

A cloud administrator is troubleshooting a web application that uses a cloud load balancer. Users report intermittent 502 Bad Gateway errors. The administrator checks the load balancer's target group and sees that some targets are marked as unhealthy. The application runs on virtual machines behind the load balancer. Which action should the administrator take to resolve the issue?

A.Increase the load balancer's idle timeout to match the application's response time.
B.Add more targets to the target group to distribute the load.
C.Verify that the health check path and port are correctly configured for the application.
D.Enable sticky sessions on the load balancer to maintain user sessions.
AnswerC

502 errors occur when the load balancer cannot establish a connection to a healthy target. If health checks are misconfigured (e.g., wrong path or port), targets are marked unhealthy and removed from rotation. When all targets are unhealthy, the load balancer returns 502. Correcting health check settings ensures only healthy targets receive traffic.

Why this answer

502 Bad Gateway errors from a load balancer indicate it cannot reach healthy backend targets. If targets are marked unhealthy, the health check configuration is likely incorrect. Verifying the health check path and port ensures the load balancer correctly assesses target health.

Increasing timeout, enabling sticky sessions, or adding targets do not fix misconfigured health checks.

Exam trap

The trap here is assuming that adding more targets or adjusting timeouts will fix 502 errors, when the real issue is often misconfigured health checks.

30
MCQhard

A global company runs a SaaS application in multiple cloud regions. They use DNS-based global load balancing to route users to the nearest region. Recently, users in Asia are experiencing high latency and timeouts. The administrator checks the health of the Asian region's resources and finds everything operational. Latency measurements from a monitoring tool show that traffic from Asian users is being routed to the European region. What should the administrator investigate first?

A.The latency-based routing policy
B.The DNS TTL settings
C.The geo-location records in the DNS provider
D.The load balancer configuration in the Asian region
AnswerA

Latency-based routing selects the region with the lowest measured latency, so a misconfigured or stale policy would explain Asian users being sent to Europe despite healthy Asian resources. Investigating this policy first satisfies the stem's routing anomaly, since resource health and DNS resolution are already confirmed working.

Why this answer

The symptom — Asian users being routed to the European region despite healthy Asian resources — points directly to a misconfigured or misbehaving latency-based routing policy. Latency-based routing relies on measured latency between the user's resolver and each regional endpoint, so if the policy is misconfigured or the measurements are stale, traffic will be sent to the wrong region.

Exam trap

The trap is blaming DNS TTL or geo-location records when the observed behavior — healthy target region but wrong routing — specifically implicates the latency-based routing policy's measurements or configuration.

How to eliminate wrong answers

Option B is wrong because DNS TTL settings affect how long resolvers cache records, but a TTL issue would cause stale routing to any region, not a consistent misroute of Asian users to Europe. Option C is wrong because geo-location records route based on the user's geographic location; if they were the mechanism in use, Asian users would be sent to Asia, so the observed behavior contradicts a geo-routing problem. Option D is wrong because the Asian region's resources are confirmed operational, and a load balancer misconfiguration would cause failures within the region, not redirect users to Europe.

31
MCQmedium

A cloud administrator is troubleshooting a web application hosted on a cloud VM that is experiencing intermittent high latency. The administrator reviews the cloud provider's monitoring metrics and sees that the VM's CPU utilization is consistently around 30%, memory usage is 40%, and network throughput is well below the instance's limit. Which factor is the most likely cause of the latency?

A.Network packet loss between the client and the VM
B.Disk I/O latency on the VM's volume
C.Insufficient CPU credits on a burstable instance
D.Memory swapping due to insufficient RAM
AnswerB

High latency with low CPU, memory, and network utilization often points to storage I/O bottlenecks. If the application performs frequent disk reads/writes, slow disk I/O can cause delays. Cloud monitoring may not show disk latency by default, so it is a prime suspect when other metrics are normal.

Why this answer

When CPU, memory, and network metrics are all well within normal ranges, storage I/O latency is a frequent hidden cause of application latency. Disk operations may be slow due to volume type, IOPS limits, or contention, and this may not be reflected in default monitoring dashboards.

Exam trap

The trap here is focusing on CPU or memory because they are common causes of latency, but the metrics show they are not saturated, so the bottleneck must be elsewhere, likely storage.

32
MCQeasy

A small business hosts a web application on a single cloud server. The server has 2 vCPUs and 4 GB RAM. Recently, the application crashes when the number of concurrent users exceeds 50. The administrator checks the system logs and finds out-of-memory (OOM) errors. What is the best course of action to resolve this issue without redesigning the application?

A.Add a load balancer and another server
B.Reduce the application's memory footprint by code optimization
C.Increase the server's RAM to 8 GB
D.Enable swap space on the server
AnswerC

Increasing memory directly resolves OOM errors without application changes.

Why this answer

The best course of action is to increase the server's RAM to 8 GB (Option C). The OOM errors indicate that the current 4 GB RAM is insufficient for 50+ concurrent users. Increasing RAM directly addresses the memory shortage without requiring application changes or redesign.

Option A (load balancer and another server) adds complexity and cost, and may not resolve the memory issue on the single server if the application is not stateless. Option B (code optimization) is a redesign effort that may not be feasible as a quick fix. Option D (enabling swap space) can lead to severe performance degradation because swapping is much slower than RAM, and may still cause crashes under high load.

33
MCQeasy

A cloud administrator is troubleshooting a virtual machine (VM) in a public cloud that has become unresponsive. The administrator cannot SSH into the VM, and the cloud provider's console shows the VM is running. The administrator suspects the VM's OS has hung. Which action should the administrator take to regain access to the VM with minimal data loss?

A.Delete the VM's network interface and recreate it.
B.Detach the VM's disk and attach it to another VM for repair.
C.Terminate the VM and create a new one from a snapshot.
D.Reboot the VM from the cloud provider's console.
AnswerD

A soft reboot (if supported) or hard reboot from the console can restart the OS without losing the VM's persistent disk data. This is the least disruptive action to regain access when the OS is hung but the VM is still running. It avoids data loss on attached volumes.

Why this answer

When a VM's OS is hung but the instance is still running, a reboot from the cloud console is the quickest way to restore access. It preserves the instance and its persistent storage, minimizing data loss and downtime compared to termination or disk detachment.

Exam trap

The trap here is assuming that because SSH fails, the VM must be terminated or the disk detached, when a simple reboot often resolves a hung OS without data loss.

34
MCQhard

A cloud administrator sees the output above when troubleshooting a virtual machine that is unresponsive. The VM is critical and must be restored quickly. What should the administrator do first?

A.Resume the VM using the virsh resume command.
B.Restart the libvirtd service on the host.
C.Increase the memory allocation for the host to free resources.
D.Migrate the VM to another host in the cluster.
AnswerA

Resuming with `virsh resume` restores a paused domain without rebooting, preserving memory state and avoiding downtime. The stem's constraint—a critical VM requiring rapid restoration—favours this over restarting, which would discard in-memory data and lengthen recovery. Paused VMs remain intact, so resumption is the fastest safe first action.

Why this answer

The output from `virsh list --all` shows the VM is in a 'paused' state, which means it is still resident in memory but not executing. The fastest way to restore a paused VM is to resume it with `virsh resume <vm-name>`, which immediately continues CPU execution without requiring a reboot or migration. This directly addresses the unresponsive behavior while preserving the VM's current memory state.

Exam trap

The trap here is that candidates assume a paused VM requires a full restart or host-level intervention, but the CV0-004 exam expects you to recognize that `virsh resume` is the immediate, low-risk recovery action for a paused domain.

How to eliminate wrong answers

Option B is wrong because restarting the libvirtd service would disrupt all VMs on the host and is unnecessary when only a single VM is paused; the issue is at the VM level, not the hypervisor daemon. Option C is wrong because increasing host memory allocation does not affect a paused VM—pausing is triggered by storage I/O errors, disk full conditions, or host memory overcommitment, not by insufficient host memory. Option D is wrong because migrating a paused VM requires resuming it first or using `virsh migrate --live` which cannot work on a paused domain; migration adds unnecessary complexity and downtime when a simple resume command will restore service immediately.

35
MCQhard

A cloud administrator notices that a virtual machine is consuming excessive CPU resources with no apparent workload. Which of the following should the administrator investigate FIRST to determine the cause?

A.A misconfigured load balancer sending traffic to the VM
B.CPU hotplug settings on the hypervisor
C.A runaway process inside the VM
D.Memory overcommitment ratio
AnswerC

Guest-level CPU consumption with no external workload points to something executing inside the operating system. A runaway or looping process is the most direct explanation and is checked first via Task Manager or top before examining host-level metrics or hypervisor scheduling.

Why this answer

A runaway process inside the VM is the most likely cause when a VM exhibits high CPU utilization without an apparent workload. This could be due to a background service, malware, or an application stuck in an infinite loop. Option A is incorrect because a misconfigured load balancer would direct traffic to the VM, which would result in network and CPU activity associated with processing that traffic, not idle high CPU.

Option B is incorrect; CPU hotplug settings affect the ability to add CPUs but do not themselves cause high CPU usage. Option D is incorrect; memory overcommitment affects memory availability, not CPU utilization.

36
MCQhard

A cloud engineer is troubleshooting an issue where an application running in a container on a Kubernetes cluster is unable to resolve DNS names. The cluster uses CoreDNS. The engineer checks the CoreDNS pod logs and sees no errors. Which of the following should the engineer check next?

A.The Kubernetes DNS service IP address
B.The container's /etc/resolv.conf file
C.The cloud provider's DNS resolver settings
D.The network policy for the namespace
AnswerB

With CoreDNS logging no errors, the fault likely lies in the pod's own DNS configuration. The container's /etc/resolv.conf supplies the nameserver and search domains used for lookups, so an incorrect or missing entry there would break name resolution despite healthy cluster DNS.

Why this answer

The container's /etc/resolv.conf file contains the DNS configuration, including the nameserver IP and search domains. If CoreDNS logs show no errors, the issue may be that the container is not using the correct DNS server or has a misconfigured resolv.conf. Checking this file is the next logical step.

Exam trap

CV0-004 often tests the assumption that DNS issues are always server-side, leading candidates to overlook client-side configuration like resolv.conf.

How to eliminate wrong answers

Option A is wrong because the Kubernetes DNS service IP is typically correct and automatically injected; if it were wrong, CoreDNS logs might show errors or the issue would be widespread. Option C is wrong because the cloud provider's DNS resolver settings are not directly used by pods unless configured, and CoreDNS handles DNS for the cluster. Option D is wrong because a network policy blocking DNS traffic would likely cause timeouts and could be checked, but since CoreDNS logs show no errors, the issue is more likely on the client side.

37
Multi-Selecthard

A cloud engineer is troubleshooting a performance issue where a web server cluster experiences high latency during peak hours. The cluster uses an auto-scaling group behind a load balancer. Which THREE steps should the engineer take to identify the root cause?

Select 3 answers
A.Monitor CPU and memory utilization on the web servers
B.Analyze web server access logs for slow requests
C.Check the load balancer's backend instance health status
D.Reduce the number of instances in the auto-scaling group
E.Review security group rules for the load balancer
AnswersA, B, C

High latency during peak hours may stem from CPU or memory saturation on the web servers themselves. Monitoring both metrics reveals whether instances are resource-constrained, which would prevent the auto-scaling group from serving requests promptly behind the load balancer.

Why this answer

Option A is correct because monitoring CPU and memory utilization on the web servers reveals whether the instances are resource-saturated during peak hours, which is a common cause of high latency in an auto-scaling cluster. Option B is correct because analyzing web server access logs for slow requests helps pinpoint which endpoints, queries, or upstream dependencies are contributing to the latency, providing concrete evidence of the bottleneck. Option C is correct because checking the load balancer's backend instance health status verifies whether unhealthy or failing instances are being served or whether health checks are flapping, which directly impacts response times.

Option D is incorrect because reducing the number of instances in the auto-scaling group would decrease capacity and likely worsen latency rather than help diagnose the root cause. Option E is incorrect because reviewing security group rules for the load balancer addresses connectivity and access control, not performance degradation during peak hours.

Exam trap

The trap here is that candidates may think reducing instances (Option D) is a valid troubleshooting step, but it is a remediation action that can mask the root cause and potentially crash the application under load.

38
MCQeasy

A cloud administrator is configuring a Linux VM as a router. The iptables rules are shown. The administrator can SSH into the VM from the network but cannot forward traffic between interfaces. What is the most likely cause?

A.The INPUT chain has a rule dropping invalid packets
B.The INPUT chain is missing a rule to allow forwarded traffic
C.The FORWARD chain's default policy is DROP and no rules allow forwarding
D.The NAT table is misconfigured
AnswerC

SSH works because INPUT accepts it, but forwarding relies on the FORWARD chain. A default DROP policy with no ACCEPT rules silently discards packets traversing the VM, so inter-interface traffic never passes despite routing being enabled.

Why this answer

The FORWARD chain in iptables controls traffic that passes through the VM (i.e., traffic not destined for the VM itself). If its default policy is DROP and no explicit ACCEPT rules exist for forwarding, the kernel will drop all forwarded packets, preventing the VM from acting as a router. SSH access works because it uses the INPUT chain, which is separate from FORWARD.

Exam trap

The trap here is that candidates confuse the INPUT chain (for local traffic) with the FORWARD chain (for transit traffic), assuming that allowing SSH implies forwarding is also allowed, when in fact they are handled by completely separate chains.

How to eliminate wrong answers

Option A is wrong because the INPUT chain dropping invalid packets affects only traffic destined for the VM itself, not forwarded traffic; SSH connectivity proves INPUT is functional. Option B is wrong because forwarded traffic is governed by the FORWARD chain, not the INPUT chain; the INPUT chain has no role in forwarding decisions. Option D is wrong because the NAT table is used for source/destination NAT (e.g., masquerading) and does not control basic IP forwarding; even with correct NAT, packets will be dropped if the FORWARD chain blocks them.

39
MCQeasy

A cloud administrator receives reports that a newly deployed application stack fails health checks and is repeatedly replaced by the orchestrator. Logs show the container starts, then exits after a few seconds with no error. The container image runs a process that daemonizes and returns control to the shell. Which of the following is the MOST likely cause?

A.The orchestrator's restart policy is set to Always, causing healthy containers to be restarted.
B.The container image was built for the wrong CPU architecture and the kernel is killing the process.
C.The container's main process forks into the background, so the runtime sees the foreground process exit and stops the container.
D.The container's health check endpoint is returning 500 errors because the application is not fully initialized.
AnswerC

Container runtimes tie the container lifecycle to the foreground process. When an entrypoint script launches a daemon that backgrounds itself, the foreground process exits, the runtime treats the container as complete, and the orchestrator restarts it. The fix is to run the process in the foreground or use a supervisor that keeps a foreground process alive.

Why this answer

Container lifecycle is bound to the foreground process. When the entrypoint backgrounds the application and returns, the runtime sees the main process complete and stops the container, so the orchestrator repeatedly recreates it. Running the server in the foreground or using a process supervisor keeps the container alive and allows health checks to succeed.

Exam trap

The trap here is focusing on the health check configuration when the container is exiting before health checks even matter, because the foreground process ends.

40
MCQmedium

A cloud architect is designing a multi-tier application that must be resilient to the failure of an entire availability zone. Which of the following strategies BEST meets this requirement?

A.Place all instances behind a single load balancer in one zone
B.Implement auto-scaling within the same availability zone
C.Use larger instance types to handle more load
D.Deploy application instances across three availability zones with a load balancer
AnswerD

Spreading instances across three availability zones means a single zone outage leaves two zones serving traffic, and the load balancer health checks route around the failed zone. This satisfies the requirement for resilience to the failure of an entire availability zone, which a single-zone or two-zone design cannot guarantee.

Why this answer

Deploying application instances across three availability zones behind a load balancer ensures that the failure of any single AZ does not take down the application. The load balancer distributes traffic to healthy instances in the remaining AZs, providing true AZ-level fault tolerance. Three AZs also satisfy the common cloud best practice of N+1 or 2N redundancy.

Exam trap

CV0-004 often tests whether candidates confuse 'high availability' (auto-scaling, larger instances) with 'fault tolerance' (multi-AZ) — the question specifically requires resilience to an entire AZ failure, which only multi-AZ satisfies.

How to eliminate wrong answers

Option A is wrong because placing all instances in one zone creates a single point of failure — an AZ outage takes down the entire application. Option B is wrong because auto-scaling within a single AZ does not protect against AZ failure; it only handles load fluctuations. Option C is wrong because larger instance types address capacity, not resilience — a single large instance in one AZ still fails if that AZ goes down.

41
Multi-Selecthard

A cloud administrator is troubleshooting a containerized application deployed on a managed Kubernetes cluster. Pods are failing to start, and the events show 'FailedMount' errors for a persistent volume claim (PVC). The PVC is bound to a persistent volume (PV) that uses a storage class with a reclaim policy of Delete. Which two actions should the administrator take to resolve the issue? (Choose two.)

Select 2 answers
A.Restart the kubelet service on the node to clear stale mount points.
B.Verify that the PV's storage class supports the access mode requested by the PVC.
C.Check that the node's kubelet has the necessary permissions to attach and mount the volume.
D.Change the reclaim policy to Retain to prevent data loss.
E.Increase the PVC's requested storage size to match the PV's capacity.
AnswersB, C

If the PVC requests an access mode (e.g., ReadWriteMany) that the underlying storage class does not support, the mount will fail. Checking compatibility ensures the volume can be attached to the pod. This is a common cause of FailedMount errors, especially with different storage backends like block vs. file storage.

Why this answer

FailedMount errors often stem from incompatible access modes or insufficient node permissions. The storage class must support the PVC's access mode, and the node's kubelet needs permissions to attach and mount the volume. Increasing size or changing reclaim policy does not affect mountability.

Restarting kubelet is a temporary measure, not a root-cause fix.

Exam trap

The trap here is focusing on storage size or reclaim policy, which are provisioning concerns, rather than mount-time issues like access modes and permissions.

42
Multi-Selecthard

A company is migrating on-premises workloads to the cloud. They need to ensure high availability for a stateless web application across two availability zones. Which THREE components should be configured to meet this requirement?

Select 3 answers
A.An auto scaling group spanning both availability zones
B.A load balancer in front of the web tier
C.A read replica database in a different AZ
D.A single large compute instance to handle all traffic
E.Multiple subnets, each in a different availability zone
AnswersA, B, E

An auto scaling group spanning both availability zones maintains capacity when one zone fails, replacing unhealthy instances in the surviving zone. This directly satisfies the high-availability constraint for the stateless web tier, since no session state pins users to a single instance.

Why this answer

Option A is correct because an auto scaling group spanning both availability zones ensures that instances are distributed across AZs and can automatically replace failed instances, maintaining high availability for the stateless web tier. Option B is correct because a load balancer in front of the web tier distributes incoming traffic across healthy instances in multiple AZs, providing fault tolerance and a single point of entry. Option E is correct because multiple subnets, each in a different availability zone, are required to place the load balancer and auto scaling group across distinct AZs, which is the foundation for AZ-level redundancy.

Option C is not correct because a read replica database is a data-tier concern and is not required for a stateless web application's high availability. Option D is not correct because a single large compute instance is a single point of failure and cannot provide high availability across two availability zones.

Exam trap

The trap here is that candidates often confuse database-level high availability (like read replicas or multi-AZ database replication) with application-tier high availability, leading them to select a database option (C) when the question explicitly targets the stateless web tier.

43
Multi-Selectmedium

A cloud administrator is troubleshooting a connectivity issue between two VPCs in the same region. Which TWO actions should the administrator verify? (Choose two.)

Select 2 answers
A.VPC peering connection status
B.Route table entries
C.Security group rules
D.VPN tunnel configuration
E.Internet gateway attachment
AnswersA, B

VPC peering provides a direct private route between two VPCs, so its connection status must be Active before traffic can flow. A Pending, Rejected or Expired state blocks all routing regardless of route tables or security groups, directly satisfying the stem's same-region VPC-to-VPC connectivity requirement.

Why this answer

Option A (VPC peering connection status) is correct because a peering connection must be in the 'active' state for traffic to flow between the two VPCs; if it is 'pending-acceptance', 'rejected', or 'deleted', connectivity will fail regardless of routing. Option B (Route table entries) is correct because each VPC's route tables must contain routes pointing the destination CIDR of the peer VPC to the peering connection (pcx-xxxx), and missing or incorrect routes are a common cause of peering failures. Option C (Security group rules) is not the primary check here because security groups are stateful and typically evaluated after routing works, and the question targets VPC-to-VPC connectivity rather than instance-level filtering.

Option D (VPN tunnel configuration) does not apply because VPC peering does not use VPN tunnels; VPNs are used for site-to-site or remote-access connections. Option E (Internet gateway attachment) is irrelevant because traffic between peered VPCs in the same region does not traverse an internet gateway.

Exam trap

CV0-004 often tests the misconception that security groups or IGWs are the first thing to check for VPC-to-VPC connectivity, when in fact peering status and route tables are the prerequisites that must exist before any security control matters.

44
MCQeasy

A cloud user is unable to connect to a web server VM from the internet after a security group rule was modified. The VM is running and can be pinged from other VMs in the same subnet. What is the most likely cause?

A.The VM's local firewall is blocking the traffic.
B.The VM's routing table is missing a default gateway.
C.The inbound rule for HTTP/HTTPS was removed or misconfigured.
D.The VM's DNS settings are incorrect.
AnswerC

Intra-subnet pings succeeding proves the VM, OS and network path are healthy, isolating the fault to the security group. Modifying that group most likely removed or misconfigured the inbound HTTP/HTTPS rule, blocking internet clients while internal traffic continues.

Why this answer

The most likely cause is that the inbound security group rule for HTTP/HTTPS was removed or misconfigured, as security groups act as virtual firewalls controlling traffic to the VM. Since the VM can be pinged from other VMs in the same subnet, the network path and VM's OS are functional, isolating the issue to the security group's inbound rules. Modifying the security group likely removed the rule allowing HTTP/HTTPS traffic from the internet.

Exam trap

CV0-004 often tests the confusion between security group rules and network ACLs, or between local firewall and cloud firewall; candidates may overlook that successful pings indicate the issue is specific to the port/protocol, not general connectivity.

How to eliminate wrong answers

Option A is wrong because if the VM's local firewall were blocking traffic, it would also block pings from other VMs, but pings are successful. Option B is wrong because a missing default gateway would prevent communication with other subnets, but the VM can be pinged from VMs in the same subnet, indicating local routing works. Option D is wrong because DNS settings affect name resolution, not connectivity; the user is unable to connect, which is a connectivity issue, not a name resolution issue.

45
MCQeasy

A cloud administrator notices that a virtual machine is unresponsive. The VM is running on a hypervisor host that shows high CPU utilization. What should the administrator do first?

A.Reboot the hypervisor host
B.Increase the VM's vCPU count
C.Migrate the VM to another host
D.Check the VM console for OS-level issues
AnswerD

High host CPU utilisation may be caused by the guest OS itself, so checking the VM console first distinguishes an OS-level hang or runaway process from genuine hypervisor contention. This is the least disruptive diagnostic step before migrating or restarting the VM.

Why this answer

The first step in troubleshooting an unresponsive VM is to check the VM console for OS-level issues. This allows the administrator to see if the OS is hung, has a kernel panic, or is waiting for input. Option A is wrong because rebooting the hypervisor host would affect all VMs and is a drastic measure that should only be taken after other diagnostics.

Option B is wrong because increasing the VM's vCPU count does not address the root cause of unresponsiveness and could worsen resource contention on an already overloaded host. Option C is wrong because migrating the VM to another host is premature without first determining if the issue is OS-related; also, migration might not be possible if the VM is completely unresponsive.

46
MCQmedium

During a disaster recovery test, a cloud administrator discovers that the standby database in a different region is not synchronized with the primary. The primary database uses asynchronous replication. What is the MOST likely reason for the sync failure?

A.License expiration on the standby database
B.Network latency causing replication lag
C.Firewall rules blocking port 3306 between regions
D.Incorrect replication configuration using a read replica instead of a standby
AnswerB

Asynchronous replication commits on the primary without waiting for the standby, so inter-region network latency directly manifests as replication lag. Sustained latency prevents the standby from ever catching up, which is the expected failure mode for cross-region async setups.

Why this answer

Asynchronous replication does not wait for the standby to acknowledge each write, so the standby can lag behind the primary. Cross-region replication is especially sensitive to network latency, which directly increases replication lag and can cause the standby to fall out of sync during a DR test. This is the most likely cause given the asynchronous, cross-region setup.

Exam trap

CV0-004 often tests the distinction between synchronous and asynchronous replication — candidates pick 'firewall blocking port' because it sounds like a concrete failure, but a blocked port would cause total replication failure, not a lag/sync gap, which is the hallmark of async replication under latency.

How to eliminate wrong answers

Option A is wrong because license expiration would typically stop the database service entirely or produce explicit license errors, not a gradual synchronization gap. Option C is wrong because if firewall rules blocked port 3306, replication would fail completely and immediately — there would be no partial sync, and the standby would show no recent transactions at all. Option D is wrong because a read replica is a valid replication target; the question states the primary uses asynchronous replication, and a read replica configured for async replication would still sync (just with lag), so misconfiguration is less likely than the inherent latency of async cross-region replication.

47
Multi-Selecteasy

A cloud administrator is investigating why a virtual machine is running slowly. The administrator checks the hypervisor performance metrics. Which TWO of the following metrics indicate CPU contention? (Choose TWO.)

Select 2 answers
A.High CPU ready time
B.High disk queue depth
C.High CPU co-stop time
D.High memory ballooning
E.High CPU usage percentage
AnswersA, C

High CPU ready time measures the percentage of time a virtual machine was ready to run but waited for a physical core. Sustained high values directly indicate CPU contention on the host, satisfying the stem's requirement for a metric evidencing contention.

Why this answer

CPU ready time (option A) and co-stop time (option C) are both indicators of CPU contention. Ready time is time a VM is ready to run but waiting for CPU; co-stop time is time a VM is stopped because another vCPU in the same VM is contending. Option E (High CPU usage percentage) is normal utilization, not contention.

Option B (High disk queue depth) is storage-related. Option D (High memory ballooning) is memory-related.

48
Multi-Selectmedium

A cloud administrator is troubleshooting a web application that is hosted on a VM in a public cloud. Users report that the application is intermittently unavailable. The administrator checks the cloud provider's status page and sees no ongoing incidents. Which two actions should the administrator take to diagnose the issue? (Choose two.)

Select 2 answers
A.Reboot the VM to clear any transient issues.
B.Enable detailed monitoring on the VM and wait for the next occurrence.
C.Increase the VM's CPU and memory resources to handle potential spikes.
D.Review the VM's system logs for errors or warnings around the times of unavailability.
E.Check the cloud provider's network ACLs and security group rules for the VM.
AnswersD, E

System logs on the VM can reveal application crashes, kernel panics, or resource exhaustion events that correlate with the unavailability. Checking logs around the reported times is a fundamental troubleshooting step to identify patterns or specific errors. This action is non-intrusive and can quickly point to the root cause, such as a service restarting or running out of memory.

Why this answer

To diagnose intermittent unavailability without a provider incident, the administrator should examine the VM's system logs for errors and verify network ACLs and security group rules. These actions directly investigate the VM's health and network configuration, which are common causes of intermittent accessibility. They are immediate, non-destructive steps that can reveal misconfigurations or application faults.

Exam trap

The trap here is jumping to resource scaling or rebooting before gathering diagnostic data, which can mask the root cause and lead to unnecessary changes.

49
MCQeasy

A virtual machine in a cloud environment is experiencing high disk I/O latency. The administrator checks the performance metrics and sees that the disk queue length is consistently above 100. What is the best immediate action?

A.Attach an additional disk and stripe the data
B.Upgrade the VM's network bandwidth
C.Migrate the VM to a host with faster disks
D.Increase the VM's memory
AnswerA

Stripping adds parallelism, reducing queue depth and improving latency.

Why this answer

Attach an additional disk and stripe the data. A high disk queue length indicates that the disk is overwhelmed with I/O requests. Stripping data across multiple disks (e.g., RAID 0) distributes the I/O load, reducing queue length and latency.

Option B is wrong because network bandwidth does not affect disk I/O. Option C is wrong because migrating to a host with faster disks may help but is not the immediate action; adding disks is quicker and more direct. Option D is wrong because increasing memory does not directly improve disk I/O performance.

50
MCQeasy

A cloud administrator is troubleshooting a web application that is hosted on multiple virtual machines behind a load balancer. Users report that the application is occasionally slow, but the load balancer health checks show all instances as healthy. The administrator suspects that one of the virtual machines is performing poorly. Which of the following should the administrator do to confirm this suspicion?

A.Restart each virtual machine one at a time to see if performance improves.
B.Review the load balancer's access logs for response times per backend instance.
C.Check the CPU utilization of each virtual machine using the cloud provider's console.
D.Enable detailed monitoring on all virtual machines and wait for new metrics.
AnswerB

Load balancer access logs typically include the backend instance that handled each request and the response time. By analyzing these logs, the administrator can identify if a particular instance has higher response times than others. This is a direct way to confirm which virtual machine is performing poorly without disrupting the service. Other methods may require more invasive actions or may not provide per-instance data.

Why this answer

Load balancer access logs are a valuable source of per-instance performance data. They record which backend handled each request and how long it took to respond. By examining these logs, the administrator can quickly identify if one instance is consistently slower.

This method is non-intrusive and uses existing data, making it the most efficient first step.

Exam trap

The trap here is assuming that CPU utilization is the definitive indicator of performance, when response time from logs is more directly tied to user experience.

51
MCQhard

A cloud administrator is troubleshooting a web application hosted on a cloud virtual machine (VM) that is experiencing intermittent high latency during peak traffic hours. The application is deployed on a single VM instance with 4 vCPUs and 8 GB RAM, running a Linux OS. The VM is connected to a virtual network with a public IP. The administrator has verified that the application code is optimized and there are no memory leaks. CPU utilization remains below 50% during peaks, but network outbound traffic shows periodic spikes up to 500 Mbps. The VM's network interface is configured with a 1 Gbps bandwidth cap. The administrator suspects that the issue is related to network throttling or packet loss. Which of the following actions should the administrator take to resolve the issue?

A.Increase the VM's vCPU count to 8 to improve processing capacity.
B.Upgrade the VM to a larger instance size with higher network bandwidth cap (e.g., 2 Gbps).
C.Configure the firewall to allow all traffic to reduce processing overhead.
D.Enable DDoS protection on the public IP to filter malicious traffic.
AnswerB

The 1 Gbps interface cap is the bottleneck: periodic 500 Mbps outbound spikes plus burst overhead can saturate it, causing throttling and packet loss. A larger instance size raises the network bandwidth cap to 2 Gbps, removing that ceiling so peak traffic no longer queues at the interface.

Why this answer

The VM's network bandwidth cap of 1 Gbps is being saturated during peak traffic (spikes up to 500 Mbps, but with overhead and burst behavior, the cap can cause throttling and packet loss). Upgrading to a larger instance size with a higher network bandwidth cap (e.g., 2 Gbps) directly addresses the bottleneck by providing more headroom for outbound traffic, reducing latency caused by queueing and drops. The administrator has already ruled out CPU and memory issues, so the network cap is the likely culprit.

Exam trap

The trap here is that candidates may assume CPU or memory is the bottleneck because latency is intermittent, but the question explicitly states CPU is below 50% and memory is fine, so the real issue is the network bandwidth cap, which is a common cloud-specific limitation tied to instance size.

How to eliminate wrong answers

Option A is wrong because increasing vCPUs does not increase network bandwidth capacity; the bottleneck is network throughput, not compute, and CPU utilization is already below 50%. Option C is wrong because configuring the firewall to allow all traffic would not reduce processing overhead in a meaningful way and could actually increase security risks; firewall processing overhead is negligible compared to the bandwidth cap limitation. Option D is wrong because DDoS protection is designed to filter malicious traffic, not to resolve throttling or packet loss caused by legitimate peak traffic exceeding the bandwidth cap.

52
Matchingmedium

Match each disaster recovery term to its definition.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Maximum time to restore services after outage

Maximum acceptable data loss in time

Automatic switch to standby system

Copy of data for restoration

Documented plan for disaster recovery

Why these pairings

Key disaster recovery terms: RTO focuses on downtime, RPO on data loss, failover is automatic, and cold site has minimal setup. Common confusions involve swapping RTO and RPO definitions or mischaracterizing failover as manual.

53
Multi-Selectmedium

A cloud engineer is troubleshooting a VM that is experiencing high latency. The VM is hosted on a hypervisor with other VMs. Which TWO metrics should the engineer review to identify if resource contention is occurring?

Select 2 answers
A.Memory ballooning
B.CPU ready time
C.Network packet drops
D.Swap usage
E.Disk queue length
AnswersA, B

Memory ballooning reveals host memory pressure: the hypervisor reclaims guest pages via the balloon driver, forcing the VM to swap and stall. Rising ballooning indicates the host is overcommitted on RAM, making it a direct contention signal alongside CPU ready time.

Why this answer

Memory ballooning (A) is correct because it directly indicates that the hypervisor is reclaiming guest memory under host-level memory pressure, which is a classic sign of memory contention among co-hosted VMs and can cause latency from paging or reduced cache. CPU ready time (B) is correct because it measures the time a vCPU is runnable but waiting for a physical CPU, which is the definitive metric for CPU contention on an oversubscribed hypervisor. Network packet drops (C) reflect network congestion or NIC issues rather than hypervisor resource contention.

Swap usage (D) is a guest-OS symptom that can result from memory pressure but does not itself identify contention between VMs. Disk queue length (E) indicates storage latency or an overloaded datastore, not contention for hypervisor CPU or memory resources.

Exam trap

CompTIA often tests the distinction between guest-level metrics (swap usage, disk queue length) and hypervisor-level metrics (ballooning, ready time), and the trap here is that candidates confuse swap usage (guest OS paging) with memory ballooning (hypervisor reclaim), or assume network packet drops indicate VM contention rather than network issues.

54
MCQeasy

A cloud engineer is troubleshooting a serverless function that is intermittently failing with timeout errors. The function is triggered by an HTTP API and processes data from an external database. The function's timeout is set to 30 seconds. Which of the following is the MOST likely cause of the timeouts?

A.The HTTP API gateway is throttling requests to the function.
B.The external database is experiencing high latency or connection issues.
C.The function's memory allocation is too low, causing it to run slowly.
D.The function's timeout value is set too high for the workload.
AnswerB

Serverless functions that call external databases are highly sensitive to network latency and database responsiveness. If the database is slow to respond or connections are timing out, the function will exceed its timeout. This is a common cause of intermittent timeouts in serverless architectures, especially when the database is outside the cloud provider's network or under heavy load.

Why this answer

Intermittent timeouts in a serverless function that accesses an external database are most commonly caused by database latency or connection problems. The function's execution time is dominated by the database call, so any slowness there can push it over the timeout limit. Addressing database performance or connection pooling is the key to resolving the issue.

Exam trap

The trap here is focusing on the function's configuration (memory, timeout) rather than the external dependency, which is the likely bottleneck in this scenario.

55
MCQmedium

A cloud load balancer is not distributing traffic evenly to backend servers. All servers pass health checks. Which of the following is the most likely cause?

A.The health check interval is set too long.
B.One of the backend servers has reached its connection limit.
C.Session persistence is enabled and directing traffic to specific servers.
D.The health check path is incorrect.
AnswerC

Session persistence pins each client to the backend that first served it, so subsequent requests bypass normal distribution algorithms. With all servers healthy, this stickiness explains the uneven spread described in the stem rather than any health-check or capacity issue.

Why this answer

Session persistence (sticky sessions) binds a client to a specific backend server for the duration of a session, typically via a cookie or source IP hash. When enabled, the load balancer deliberately routes repeat requests from the same client to the same backend, which skews the distribution and makes traffic appear uneven even though all servers are healthy. This is the most likely cause because the scenario explicitly states all servers pass health checks, ruling out availability-based causes.

Exam trap

CV0-004 often tests the misconception that uneven distribution always indicates a backend problem; candidates overlook that session persistence is a deliberate configuration that intentionally skews traffic, and the 'all servers pass health checks' clue is the key differentiator.

How to eliminate wrong answers

Option A is wrong because the health check interval only controls how frequently the load balancer probes backend health — it does not influence how traffic is distributed among healthy servers. Option B is wrong because if a backend server had reached its connection limit, it would typically fail health checks or be marked unhealthy, contradicting the premise that all servers pass health checks; connection limits affect capacity, not the balancing algorithm's distribution logic. Option D is wrong because an incorrect health check path would cause health checks to fail, marking servers unhealthy — again contradicting the stated condition that all servers pass health checks.

56
MCQmedium

A cloud administrator is troubleshooting slow performance on a managed relational database. Read queries against a read replica are fast, but write operations on the primary node take far longer than during testing. Monitoring shows the primary's CPU is moderate, disk queue depth is high, and provisioned IOPS are consistently saturated. Which action should the administrator take to resolve the bottleneck?

A.Move the primary node to a different availability zone to reduce storage latency.
B.Increase the provisioned IOPS and throughput of the primary's storage volume to match the write workload.
C.Increase the primary instance's memory so more of the working set can be cached.
D.Add another read replica to distribute the write load across more nodes.
AnswerB

Sustained high disk queue depth together with saturated provisioned IOPS points directly to a storage throughput ceiling, not to CPU or query logic. Write operations depend on durable storage commits, so they stall when the volume cannot absorb the write rate, while reads served from a replica with its own volume stay fast. Raising provisioned IOPS addresses the measured constraint.

Why this answer

High disk queue depth combined with saturated provisioned IOPS identifies the storage layer as the limiting factor for writes. Because write latency depends on durable commits to the volume, the fix is to raise provisioned IOPS and throughput so the primary can absorb the write rate. CPU is moderate and reads are healthy, so instance sizing, replica count, and zone placement are not the causes.

Exam trap

The trap here is focusing on CPU or instance class because the symptom is slowness, when the metrics already point to storage throughput saturation.

57
MCQeasy

A cloud administrator runs a deployment script that creates multiple resources using Infrastructure as Code (IaC). The script fails with a "400 Bad Request" error when attempting to create a storage account. Which troubleshooting step should the administrator take first?

A.Check the network connectivity to the cloud API endpoint.
B.Increase the timeout value for the API call.
C.Review the error message details for a specific validation error.
D.Verify that the script has the correct region parameter.
AnswerC

A 400 response signals client-side request rejection, and the response body carries the precise validation failure, such as an invalid name, unsupported region, or SKU mismatch. Reading those details identifies the offending property before any retry or script change is attempted.

Why this answer

A '400 Bad Request' from a cloud API indicates a client-side validation error, such as an invalid parameter, missing required property, or malformed request body. The fastest and most accurate first step is to read the error message details, which typically name the exact field or constraint that failed. This avoids guessing at network, timeout, or region issues that would produce different error codes.

Exam trap

CV0-004 often tests whether candidates jump to infrastructure causes (network, timeout, region) for any API failure — the trap is ignoring that a 400 is a client-side validation error whose details should be read first.

How to eliminate wrong answers

Option A is wrong because network connectivity problems would typically manifest as timeouts, DNS failures, or connection refused errors, not a 400 Bad Request returned by the API. Option B is wrong because increasing the timeout addresses slow responses (e.g., 408 or timeout errors), not a validation rejection that returns immediately. Option D is wrong because an incorrect region parameter would usually produce a 404 or a region-specific error, and the 400 already points to a request validation issue that the error detail will clarify.

58
MCQmedium

A cloud engineer is troubleshooting an issue where users cannot connect to a web application hosted on a cloud VM. The VM's security group allows HTTP (port 80) from 0.0.0.0/0, and the VM's OS firewall is disabled. The engineer can ping the VM's public IP from the internet. What is the most likely cause of the issue?

A.OS firewall is blocking port 80
B.Incorrect routing table on the VM
C.Security group rule is applied to the wrong subnet
D.Web server service is not running on the VM
AnswerD

Ping succeeding proves the network path, security group rule and routing are functional, so the fault lies above layer 3. A stopped or crashed web server process means nothing listens on port 80, producing connection refusals despite reachable IP connectivity.

Why this answer

Since the OS firewall is disabled and the security group allows HTTP from 0.0.0.0/0, the only remaining layer that could block connectivity is the application itself. If the web server service (e.g., Apache, Nginx, IIS) is not running on the VM, it will not listen on TCP port 80, so HTTP requests will be refused even though network-level access is permitted. The ability to ping the VM confirms IP-level reachability, isolating the issue to the application layer.

Exam trap

The trap here is that candidates assume a ping success implies all services are reachable, but ICMP (ping) operates at the network layer (Layer 3) and does not test TCP port availability, so a running web server is required for HTTP connectivity.

How to eliminate wrong answers

Option A is wrong because the OS firewall is explicitly stated as disabled, so it cannot be blocking port 80. Option B is wrong because routing tables on the VM control outbound traffic, not inbound connections to the VM; inbound traffic is handled by the cloud provider's virtual network and security groups. Option C is wrong because security groups are stateful and applied at the VM network interface level, not to subnets; even if the rule were misapplied, the VM's security group explicitly allows HTTP from 0.0.0.0/0, so this is not the cause.

59
MCQeasy

A cloud engineer notices that an application is running slower than expected. Monitoring shows that the CPU utilization is consistently below 30%, but memory usage is at 95%. Which of the following is the most likely cause of the performance issue?

A.Insufficient disk space for application logs
B.Insufficient memory causing swapping to disk
C.Network bandwidth saturation
D.CPU contention due to overprovisioning
AnswerB

Insufficient memory forces the operating system to page data from RAM to disk, and that swapping introduces severe latency because disk access is orders of magnitude slower than memory. With memory at 95% while CPU sits below 30%, the bottleneck is memory pressure, not processing capacity, so adding RAM resolves the slowdown.

Why this answer

When memory usage is at 95% and CPU utilization is low, the system is likely thrashing—the operating system is forced to page memory to disk (swap) to free RAM. Disk I/O is orders of magnitude slower than RAM, so even with idle CPU, the application stalls waiting for swap operations. This explains the performance degradation despite low CPU load.

Exam trap

The trap here is that candidates often associate performance issues solely with CPU or network bottlenecks, overlooking the severe impact of memory exhaustion and disk swapping, which can masquerade as a slow application with ample CPU headroom.

How to eliminate wrong answers

Option A is wrong because insufficient disk space for logs would cause write failures or application crashes, not a gradual slowdown with high memory and low CPU. Option C is wrong because network bandwidth saturation would manifest as high latency or packet loss, not as high memory usage with low CPU. Option D is wrong because CPU contention due to overprovisioning would show high CPU ready times or steal time, not consistently low CPU utilization; overprovisioning typically leads to CPU starvation, not memory exhaustion.

60
MCQmedium

A cloud administrator is managing a multi-tier application in a public cloud. The database tier is hosted on a VM with a persistent disk. Users report that the application is slow, and the administrator notices that disk I/O latency is high. The VM's disk is a standard network-attached storage volume. Which action should the administrator take to improve disk performance?

A.Enable read caching on the VM's operating system.
B.Change the disk type to a higher-performance SSD-based volume.
C.Increase the size of the disk volume to get higher IOPS.
D.Move the database to a VM with more memory.
AnswerB

Standard network-attached storage often has lower IOPS and higher latency compared to SSD-based volumes. Upgrading to a provisioned IOPS SSD or general-purpose SSD volume can significantly reduce latency and increase throughput. This directly addresses the performance bottleneck by providing faster storage media and potentially higher IOPS limits, which is the most effective solution for high disk I/O latency.

Why this answer

High disk I/O latency on a standard network-attached volume indicates that the storage tier is the bottleneck. Upgrading to an SSD-based volume, such as a provisioned IOPS SSD, provides lower latency and higher throughput. This is the most direct and effective way to improve disk performance for a database workload that is sensitive to I/O latency.

Exam trap

The trap here is assuming that increasing disk size or adding memory will solve I/O latency, when the real fix is to change the storage type to a higher-performance option.

61
MCQhard

A cloud administrator is investigating a sudden increase in latency for a microservices application. Distributed traces show that a single downstream service's response time grew from 20 ms to 2 seconds, and its CPU utilization remains low at 15 percent. The service makes calls to an external third-party API. Which of the following is the MOST likely cause?

A.The service's container CPU limit is throttling it during request bursts.
B.The service is blocked waiting on the third-party API, and its thread pool is saturating under the increased wait time.
C.The service's memory limit is too low, causing the kernel to swap pages to disk.
D.The service's network interface is experiencing packet loss, causing TCP retransmissions.
AnswerB

When a downstream dependency slows from milliseconds to seconds, worker threads spend most of their time waiting, so the pool fills and queued requests wait even longer. CPU stays low because threads are blocked on I/O, not computing. The growing external API latency is the trigger, and the thread pool saturation amplifies it across the service.

Why this answer

A slowdown in an external dependency pushes worker threads into long waits, so a fixed-size thread pool saturates and requests queue behind blocked workers. CPU remains low because the threads are waiting on I/O, not executing. The trace data localizing latency to the third-party call identifies the trigger, and the service's concurrency model explains the amplification.

Exam trap

The trap here is assuming low CPU utilization means the service is healthy, when blocked threads waiting on an external dependency are the classic signature of this pattern.

62
MCQhard

After deploying a new application version, users get 503 errors. The application runs on Kubernetes in a private cloud. What is the most likely cause?

A.Application health check failing
B.Incorrect ingress configuration
C.Insufficient pod resources
D.Node port exhaustion
AnswerA

A failing readiness or liveness probe causes Kubernetes to remove pods from Service endpoints or restart them, leaving no healthy backends to serve traffic, so the ingress returns 503. This directly matches the scenario: a new version deployed, private-cloud Kubernetes, and users receiving 503 errors.

Why this answer

After deploying a new application version, users get 503 errors. In Kubernetes, a 503 Service Unavailable error often indicates that the service has no healthy endpoints. The most likely cause is that the application's health check (readiness probe) is failing, causing Kubernetes to remove the pod from the service's endpoints.

This can happen if the new version has a bug, misconfiguration, or takes longer to start than the probe's thresholds allow.

Exam trap

CV0-004 often tests Kubernetes troubleshooting; candidates may jump to ingress or resource issues, but 503 errors specifically point to health check failures or no available endpoints.

How to eliminate wrong answers

Option B is wrong because an incorrect ingress configuration would typically result in 404 errors or routing issues, not 503, unless the ingress cannot reach the service, but that is less direct. Option C is wrong because insufficient pod resources might cause pods to crash or be evicted, leading to 503 if no pods are available, but health check failure is more specific to a new deployment. Option D is wrong because node port exhaustion would affect all services on the node, not just the new application, and is less likely in a private cloud with proper planning.

63
MCQhard

A cloud engineer is troubleshooting a containerized application deployed on a managed Kubernetes cluster. The application pods are repeatedly restarting with 'OOMKilled' status. The engineer reviews the pod specification and sees that the memory request is 512Mi and the memory limit is 1Gi. The application is a Java-based service with a heap size set to 768Mi. Node metrics show that nodes have 8Gi of memory with 2Gi available. Which of the following is the MOST likely cause of the OOMKilled events?

A.The Java heap size exceeds the container memory limit when accounting for JVM overhead.
B.The container is being killed by the Kubernetes liveness probe due to high memory usage.
C.The Java application has a memory leak that causes it to exceed the heap size.
D.The memory request is too low, causing the pod to be scheduled on a node with insufficient memory.
AnswerA

The container memory limit is 1Gi, but the Java heap is set to 768Mi. The JVM requires additional memory beyond the heap for metaspace, thread stacks, code cache, and native memory. This overhead can easily exceed 256Mi, pushing total memory usage above the 1Gi limit. When the container exceeds its limit, the kernel OOM killer terminates the process, resulting in 'OOMKilled' status. The node has available memory, so the issue is the container limit, not node pressure.

Why this answer

The container memory limit is 1Gi, but the Java heap is 768Mi. The JVM requires additional native memory for metaspace, thread stacks, and other structures. When the total memory usage exceeds the 1Gi limit, the kernel OOM killer terminates the container, resulting in 'OOMKilled'.

Although the node has available memory, the container's cgroup limit is the constraint. Increasing the memory limit or reducing the heap size would resolve the issue.

Exam trap

The trap here is focusing on node-level memory availability and overlooking the container's cgroup limit, which is the actual constraint triggering the OOM killer.

64
MCQmedium

A cloud administrator is investigating a sudden increase in egress charges for an application running on multiple Linux VMs in a public cloud. The application makes frequent calls to a third-party REST API over the public internet. Which action should the administrator take FIRST to identify the source of the unexpected egress traffic?

A.Configure a NAT gateway and route all outbound traffic through it to centralize logging.
B.Install a host-based intrusion detection system (HIDS) on each VM to monitor outbound connections.
C.Enable VPC flow logs and analyze the destination IP addresses and byte counts for outbound traffic.
D.Review the cloud provider's billing dashboard for the previous month to compare costs.
AnswerC

VPC flow logs capture metadata about IP traffic to and from network interfaces in a VPC, including source/destination IPs and bytes transferred. Filtering for outbound traffic to the third-party API's public IPs will reveal which instances are generating the most egress and confirm whether the increase is due to API calls or other traffic. This is the most direct first step to localize the source before deeper packet analysis.

Why this answer

VPC flow logs provide the necessary network-level visibility to identify the source of increased egress traffic. By analyzing flow logs, the administrator can see which instances are communicating with external IPs and the volume of data transferred. This targeted approach quickly narrows down the cause, whether it is a misconfigured application, a compromised instance, or a change in API usage patterns.

Exam trap

The trap here is assuming that billing or host-based tools can provide the granular network traffic details needed to pinpoint egress sources, when in fact flow logs are the correct first step.

65
MCQhard

A cloud engineer is troubleshooting a storage performance issue. The storage is backed by a SAN with a mix of SSD and HDD drives. Which of the following metrics would BEST indicate that the storage subsystem is the bottleneck?

A.Low memory usage on the hypervisor
B.High network utilization on storage network links
C.High disk queue depth and latency
D.High CPU utilization on all application servers
AnswerC

Queue depth measures outstanding I/O requests awaiting service, and latency measures response time. Both rising together shows requests are queuing faster than the SSD/HDD mix can serve them, isolating the storage subsystem rather than CPU, network or application as the bottleneck.

Why this answer

High disk queue depth and latency directly indicate that I/O requests are waiting, which is a classic sign of a storage bottleneck. Low memory usage (A) does not indicate a storage issue. High network utilization (B) could be caused by storage traffic but does not confirm the storage subsystem is the bottleneck; it could be normal.

High CPU utilization (D) points to compute, not storage.

66
MCQhard

A cloud orchestration template fails to deploy resources with the error 'Resource limit exceeded'. The administrator has enough quota for all services. What is the most likely cause?

A.The template has a syntax error in the JSON.
B.A specific resource type has reached its service limit.
C.The custom image used is corrupted.
D.The IAM role used does not have permission to create resources.
AnswerB

Quota is per resource type, not aggregate, so overall headroom can exist while one service limit is exhausted. The deployment fails because that specific resource type has hit its ceiling, satisfying the 'Resource limit exceeded' error.

Why this answer

Certain resource types have service-specific limits that are separate from the overall account quota. Even if the administrator has enough total quota, a specific resource type (e.g., virtual machines, storage accounts) may have reached its maximum allowed count. Option A is incorrect because a syntax error would cause a different error, such as parsing failure.

Option C is incorrect because a corrupted image would typically result in image-related errors. Option D is incorrect because permission issues would generate an access denied error, not 'Resource limit exceeded'.

67
MCQmedium

A cloud administrator manages a three-tier application in a public cloud. Users report that API calls from the web tier to the database tier fail with connection timeouts, but the database tier responds normally when queried from a bastion host on the same subnet. The web tier instances reside in a different subnet. Which of the following is the MOST likely cause?

A.The web tier instances are using an outdated database client library that cannot negotiate the TLS version required by the database.
B.The database's security group inbound rules do not allow traffic from the web tier's subnet CIDR range on the database port.
C.The database engine's max_connections parameter is set too low for the incoming web tier connections.
D.The web tier's route table lacks a route to the database subnet's CIDR range.
AnswerB

Security groups are stateful and evaluated per source. Because the bastion host on the same subnet succeeds while web tier instances in a different subnet time out, the database's inbound rule likely scopes the permitted source to the bastion's subnet or IP rather than the web tier CIDR. Adding the web tier subnet on the correct database port resolves the timeout.

Why this answer

The bastion host succeeds from its own subnet while web tier instances in a separate subnet time out, which isolates the problem to source-based filtering rather than routing or the database engine itself. Security groups and network ACLs evaluate the source address, so a rule scoped to the bastion's range blocks the web tier. Allowing the web tier subnet CIDR on the database port restores connectivity.

Exam trap

The trap here is assuming that because the bastion can reach the database, the database is healthy and the problem must be in the web tier, when in fact source-scoped security group rules commonly differ by subnet.

68
MCQmedium

A cloud administrator is troubleshooting a Linux virtual machine (VM) in a public cloud that is experiencing packet loss when communicating with another VM in the same subnet. The administrator runs 'ip -s link show eth0' and observes a high number of dropped packets on the receive side. The VM's CPU utilization is low, and the network interface is a paravirtualized driver (virtio). Which of the following is the MOST likely cause of the dropped packets?

A.The receive ring buffer on the network interface is too small, causing packets to be dropped when the buffer overflows.
B.The VM's firewall is dropping packets due to a misconfigured rule.
C.There is a physical network issue, such as a faulty cable or switch port.
D.The VM's CPU is not allocating enough cycles to process network interrupts.
AnswerA

A high number of receive drops on a virtio interface often indicates that the receive ring buffer is full. When packets arrive faster than the guest can process them, the buffer overflows and packets are dropped. This is common in virtualized environments where the virtio driver's ring size may be default and insufficient for high throughput. Increasing the ring buffer size using ethtool -G can mitigate the issue, and low CPU utilization suggests the guest is not overwhelmed, pointing to buffer capacity.

Why this answer

The high number of receive drops on the virtio interface, combined with low CPU utilization, points to receive ring buffer exhaustion. The virtio driver uses a ring buffer to hold incoming packets; if the buffer is too small for the traffic rate, packets are dropped. Increasing the ring buffer size with ethtool -G eth0 rx <size> can resolve the issue.

Physical network issues and firewall drops are less likely in a cloud environment and would not manifest as interface-level receive drops.

Exam trap

The trap here is assuming that packet loss in the cloud is due to physical network problems, when in virtualized environments it is often a driver buffer or configuration issue.

69
Multi-Selectmedium

A company's application is unable to connect to a managed cloud database. The database is deployed in a VPC with public accessibility disabled. The application runs on an EC2 instance in the same VPC. Which three troubleshooting steps should the administrator take? (Choose three.)

Select 3 answers
A.Ensure the VPC has an internet gateway attached.
B.Check the network ACL associated with the database subnet for appropriate rules.
C.Verify that the database endpoint is correctly configured in the application.
D.Verify that the EC2 instance has a public IP address.
E.Check the security group for the database to ensure it allows inbound traffic from the EC2 instance's security group.
AnswersB, C, E

Network ACLs are stateless and evaluated per subnet, unlike security groups. The database subnet's ACL must permit inbound traffic on the database port and outbound return traffic to the EC2 subnet's ephemeral range, or packets are dropped before reaching the database.

Why this answer

Option B is correct because network ACLs are stateless subnet-level firewalls; if the database subnet's NACL lacks an inbound rule allowing the database port (e.g., 3306 for MySQL or 5432 for PostgreSQL) from the EC2 instance's subnet CIDR, and a corresponding outbound rule for the return traffic, connectivity will fail even if security groups are correct. Option C is correct because a misconfigured endpoint (wrong hostname, port, or database name) in the application's connection string is a common cause of connection failures and must be verified before deeper network troubleshooting. Option E is correct because the database's security group must have an inbound rule referencing the EC2 instance's security group (or its CIDR) on the database listener port; since the database is not publicly accessible, this security group reference is the primary stateful access control.

Option A is not needed because an internet gateway only enables internet connectivity for public subnets and is irrelevant for instance-to-database traffic within the same VPC. Option D is not needed because a public IP is only required for internet-facing communication, not for private communication between an EC2 instance and a database in the same VPC.

Exam trap

CV0-004 often tests the misconception that private intra-VPC connectivity requires an internet gateway or public IP, when in fact security groups, NACLs, and endpoint configuration are the real determinants.

70
Multi-Selecthard

A cloud administrator is troubleshooting a performance issue where a web application is responding slowly. The application runs on virtual machines in a private cloud. The administrator has verified that CPU and memory utilization are within normal limits. Which TWO additional metrics should the administrator check to diagnose the issue?

Select 2 answers
A.Number of running processes
B.Network latency between the application and database servers
C.Disk I/O wait time on the hypervisor
D.Virtual machine snapshot size
E.Hypervisor version
AnswersB, C

Network latency directly affects request round-trip time between application and database tiers, a bottleneck CPU and memory metrics cannot reveal. Since the stem confirms compute resources are normal, measuring inter-tier latency isolates transport delay or congestion as the cause of slow responses, satisfying the need to identify non-compute performance constraints.

Why this answer

Network latency between the application and database servers is a critical metric because slow database queries or network congestion can cause the web application to respond slowly even when CPU and memory on the VMs are normal. High latency increases round-trip time for SQL queries, directly impacting page load times. Disk I/O wait time on the hypervisor is also essential because excessive I/O wait indicates storage contention, which can throttle read/write operations for the VMs, leading to application sluggishness.

Exam trap

CompTIA often tests the distinction between VM-level metrics (CPU/memory) and infrastructure-level metrics (network/storage), trapping candidates who overlook that application performance can degrade due to external dependencies even when the VM itself appears healthy.

71
MCQhard

A company uses a hybrid cloud model with an AWS Direct Connect connection to its on-premises network. Users report intermittent connectivity to cloud resources. A network engineer finds packet loss on the Direct Connect virtual interface. Which of the following should be checked FIRST to resolve the issue?

A.The physical port status of the Direct Connect router
B.The MTU setting on the on-premises firewall
C.The BGP session status between the on-premises router and the AWS Direct Connect endpoint
D.The VPN tunnel status for the Direct Connect link
AnswerC

Packet loss on a Direct Connect virtual interface most commonly stems from a flapping or down BGP session, which withdraws routes and drops traffic. Verifying the BGP session status between the on-premises router and the AWS endpoint is the fastest first diagnostic step.

Why this answer

Intermittent packet loss on a Direct Connect virtual interface is most commonly caused by BGP session flapping or misconfiguration, as BGP is the routing protocol that establishes and maintains connectivity between the on-premises router and the AWS Direct Connect endpoint. Checking the BGP session status first allows the engineer to quickly identify if the issue is due to route advertisement problems, hold timer mismatches, or session resets, which are frequent root causes of intermittent packet loss.

Exam trap

The trap here is that candidates often confuse Direct Connect with VPN-based connections and assume a VPN tunnel is involved, leading them to check VPN status (Option D) instead of the BGP session that actually governs the virtual interface routing.

How to eliminate wrong answers

Option A is wrong because the physical port status of the Direct Connect router would show a hard failure (e.g., link down) rather than intermittent packet loss; intermittent issues are rarely caused by physical port problems unless there is a duplex mismatch or cable fault, but these are less likely to be the first check. Option B is wrong because MTU settings on the on-premises firewall typically cause fragmentation or black-hole issues for large packets, not intermittent packet loss across all traffic; MTU mismatches usually result in consistent packet drops for packets exceeding the MTU, not sporadic loss. Option D is wrong because Direct Connect does not use a VPN tunnel; it is a dedicated physical connection, and VPN tunnels are used for AWS Site-to-Site VPN, not Direct Connect virtual interfaces.

72
MCQeasy

A cloud administrator is troubleshooting a newly deployed web application on a cloud VM. Users cannot access the application via its public IP address, but the administrator can SSH into the VM and access the application locally using curl http://localhost:8080. The VM's security group allows inbound traffic on port 22 and port 80. The application is listening on port 8080. Which of the following is the MOST likely cause of the issue?

A.The application is configured to use HTTPS instead of HTTP.
B.The application is bound to the loopback interface instead of all interfaces.
C.The security group does not allow inbound traffic on port 8080.
D.The VM's operating system firewall is blocking port 8080.
AnswerC

The security group allows inbound traffic on ports 22 and 80, but the application is listening on port 8080. Since users are trying to access the application via its public IP, they are likely using a URL without specifying port 8080 (e.g., http://public-ip), which defaults to port 80. However, even if they specified port 8080, the security group would block it because only ports 22 and 80 are allowed. The local curl to localhost:8080 works because it bypasses the security group. Thus, the security group is the most likely cause.

Why this answer

The security group allows inbound traffic only on ports 22 and 80, but the application listens on port 8080. External users cannot reach port 8080 because the security group blocks it. Local access via localhost:8080 works because it does not traverse the security group.

To resolve the issue, the administrator should either add an inbound rule for port 8080 or reconfigure the application to listen on port 80.

Exam trap

The trap here is assuming that because local access works, the application is correctly configured, and overlooking that the security group only permits specific ports.

73
MCQeasy

A cloud administrator cannot deploy a new VM from a custom image. The deployment fails with an error stating 'Incompatible hypervisor version'. What is the most likely cause?

A.The image was created on a newer hypervisor than the current host.
B.The VM's virtual hardware version is too old.
C.The image file is corrupt.
D.The storage backend does not support the image format.
AnswerA

A custom image carries a hypervisor compatibility level set by the host that created it. If that level exceeds the destination host's hypervisor version, the platform rejects deployment, producing exactly this incompatibility error. The image must be recreated on an equal or older hypervisor.

Why this answer

The error 'Incompatible hypervisor version' indicates that the custom image was created on a hypervisor version that is newer than the one on the host where deployment is attempted. Hypervisors like VMware ESXi, Hyper-V, or KVM have version-specific features and virtual hardware compatibility; an image from a newer version may not be supported on an older host.

Exam trap

CV0-004 often tests the direction of compatibility: candidates might think an older image is incompatible with a newer hypervisor, but the error specifically indicates the image is from a newer hypervisor than the host.

How to eliminate wrong answers

Option B is wrong because if the VM's virtual hardware version is too old, it would typically still be compatible (newer hypervisors support older versions), and the error would not mention hypervisor version. Option C is wrong because a corrupt image would produce a different error, such as checksum failure or invalid format. Option D is wrong because an unsupported image format would yield an error about format, not hypervisor version.

Ready to test yourself?

Try a timed practice session using only Troubleshooting questions.