An administrator is troubleshooting a server that runs a critical application. The server has 16 GB of RAM and 8 CPU cores. The administrator notices that the server becomes very slow during peak hours. Analysis of 'iostat -x 1' shows that the average wait time (await) for the main disk (sda) is consistently above 1000 ms, while the average service time (svctm) is around 5 ms. What is the most likely cause?
Trap 1: The CPU is overloaded, causing processes to wait for CPU time.
High await with low svctm points to I/O queueing, not CPU contention; CPU saturation would show high load or run-queue delays, not disk wait. It is tempting because slow peak-hour performance often stems from CPU, and would be correct if vmstat showed runnable processes waiting.
Trap 2: The system is using swap space heavily, causing disk I/O.
Heavy swap usage would cause high disk utilization, but the svctm would likely be higher because swapping involves random I/O, and the await would be high as well. However, the low svctm suggests the disk is serving requests quickly once they start, which is not typical of swap.
Trap 3: The disk is experiencing hardware errors.
Hardware errors would raise svctm, not leave it at 5 ms while await exceeds 1000 ms; the gap indicates queueing. It is tempting because disk faults cause slowness, and would be correct if iostat showed elevated service times or errors.
- A
The CPU is overloaded, causing processes to wait for CPU time.
Why it fails: High await with low svctm points to I/O queueing, not CPU contention; CPU saturation would show high load or run-queue delays, not disk wait. It is tempting because slow peak-hour performance often stems from CPU, and would be correct if vmstat showed runnable processes waiting.
- B
The system is using swap space heavily, causing disk I/O.
Why it fails: Heavy swap usage would cause high disk utilization, but the svctm would likely be higher because swapping involves random I/O, and the await would be high as well. However, the low svctm suggests the disk is serving requests quickly once they start, which is not typical of swap.
- C
The disk is experiencing hardware errors.
Why it fails: Hardware errors would raise svctm, not leave it at 5 ms while await exceeds 1000 ms; the gap indicates queueing. It is tempting because disk faults cause slowness, and would be correct if iostat showed elevated service times or errors.
- D
There is a large queue of I/O requests waiting to be serviced.
Await of 1000ms against svctm of 5ms means requests spend almost all their time queued rather than being serviced. The disk itself is fast; the bottleneck is the backlog of pending I/O requests exceeding what sda can dispatch concurrently.